Site Reliability Engineer responsible for infrastructure supporting AI platform. Safeguarding US customer data and ensuring compliance in the Aerospace and Defense sector.
Responsibilities
Design, implement, and operate highly available, scalable, and fault-tolerant infrastructure primarily on GCP, but to include multi-cloud deployments. Optimize system performance, manage disaster recovery, and ensure cost-effectiveness.
Lead Terraform-based infrastructure development with security best practices, encrypted state management, and governance tools.
Build robust pipelines supporting hundreds of developers and AI engineers. Integrate automated security testing, vulnerability scanning, and compliance checks throughout the development lifecycle.
Implement comprehensive observability strategies using Prometheus, Grafana, and ELK. Define SLOs/SLIs, manage error budgets, and lead incident response with blameless post-mortems.
Navigate complex regulatory requirements for U.S. Aerospace and Defense Industrial Base. Collaborate with security and legal teams on expanding compliance standards.
Reduce operational toil through Python, Go, or Bash automation. Work in a follow-the-sun model with global teams while taking primary responsibility for US platform partition incidents and operations.
Requirements
Bachelor's degree in Computer Science, Engineering, or equivalent experience
2+ years in Site Reliability Engineering, DevOps, or Systems Engineering with cloud-based SaaS platforms
Deep Terraform and Infrastructure as Code expertise with security best practices
Proficiency in Python and other scripting/programming languages
Modern CI/CD experience (Github Actions, GitLab CI, Jenkins, ArgoCD, Spinnaker) including AI/ML workloads
Strong cloud platform experience, preferably GCP (AWS, Azure experience also valuable for future multi-cloud deployments)
Experience building and optimizing containers (Docker) and configuring orchestration (Kubernetes)
Regulated industry experience (Aerospace & Defense, Finance, Healthcare) with experience building secure platforms
DevSecOps principles and security integration experience
Security-first development mindset with understanding of secure infrastructure practices
Strong problem-solving and communication skills for distributed team environments
What would have us dialing your number immediately
Hyper-growth startup experience
AI Safety experience
MLOps and AI/ML infrastructure security experience
Benefits
Comprehensive Health Benefits: We provide 100% company-covered employee comprehensive health insurance, including medical (UnitedHealth), dental (Principal), and vision (VSP) to keep you and your family healthy.
Compensation: Base salary is $110,000 to $150,000 annually based off of experience and skills of the candidate.
Ownership & Rewards: Be a part of our success story with a competitive stock options plan.
Financial Security: Start saving for your future with our 401k plan, featuring a generous 4% company match starting on day one.
Generous Time Off: Maintain a healthy work-life balance with 15 days of paid time off, five dedicated sick days, and ten company holidays to celebrate throughout the year.
Thriving Culture: We foster a vibrant work environment with delicious company lunches, engaging events, and healthy drinks and snacks to keep you fueled. Celebrate your achievements with us at quarterly events and holiday gatherings.
Learning & Development: We invest in your growth by providing opportunities to join professional organizations, attend industry conferences, and participate in various learning initiatives.
Financial Incentives: Benefit from commuter and parking benefits to simplify your daily commute. We also offer referral bonuses to help you spread the word about exciting opportunities at CADDi.
Secure DevOps Engineer responsible for integrating security into CI/CD pipelines and strengthening AWS infrastructure. Key expertise in AWS security and container management.
DevOps Engineer responsible for CI/CD pipeline development and automation for urban software solutions. Collaborating with teams to enhance efficiency in software deployment and infrastructure.
DevOps Engineer managing cloud and on - premise platforms for a public sector infrastructure project. Collaboration primarily remote, with occasional on - site meetings.
DevSecOps Engineer architecting CI/CD framework services for Truist, enhancing the flow of business value through DevSecOps practices. Building and maintaining automation for software delivery and operations.
Application Security Manager at Evertec, handling security strategy and implementation in financial tech. Leading efforts in Application Security, DevSecOps, and compliance with financial regulations.
Databricks Senior DevOps Engineer designing and operating platforms on AWS and Databricks for Financial Crime. Focused on platform infrastructure, governance, security, and operations.
Site Reliability Engineer at Assecor, focusing on SLIs, SLOs, and incident management. Enhancing performance and reliability through observability and automation in a hybrid work environment.
DevOps Architect at Ascensus, responsible for technical direction and oversight for application engineering practices across scrum teams. Promotes DevOps culture and innovative solutions.
Cloud Site Reliability Engineer ensuring scalability, performance, and reliability of cloud infrastructure deployed in Woven City. Working with product owners and teams for innovative solutions.
Senior DevOps Engineer supporting enterprise - grade Kubernetes infrastructure and CI/CD automation for U.S. Army projects. Engaging in critical system designs and automation processes with a focus on cloud - based platforms.