Staff SRE at Insulet managing teams to ensure system reliability and scalability. Driving best practices in Site Reliability Engineering with a focus on automation and modern technologies.
Responsibilities
Provide technical guidance and mentorship to the SRE team.
Drive the implementation of best practices in reliability, scalability, and performance.
Lead by example, demonstrating excellence in technical skills and problem-solving.
Collaborate with cross-functional teams to design scalable, resilient, and efficient systems.
Architect and implement infrastructure solutions that meet the requirements of high availability and performance.
Drive the adoption of modern technologies and tools to improve system reliability and efficiency.
Develop and maintain automation tools for provisioning, deployment, and monitoring.
Automate routine tasks to improve operational efficiency and reduce manual intervention.
Design and implement monitoring solutions to proactively identify issues and prevent service disruptions.
Lead incident response efforts, conducting post-mortem analysis, and implementing measures to prevent recurrence.
Develop & Automate runbooks and playbooks to streamline incident resolution processes.
Conduct capacity planning exercises to ensure systems can handle current and future loads.
Identify performance bottlenecks and optimize system performance through tuning and optimization efforts.
Collaborate with development teams to design and implement scalable architectures.
Document system architectures, configurations, and procedures.
Promote knowledge sharing within the team through technical presentations, workshops, and documentation.
Requirements
Bachelor’s in computer science, Engineering, or a related field.
9+ years of experience in the field including 5+ Site Reliability Engineering, DevOps, or a similar role.
Proven experience architecting and managing highly available, scalable, and fault-tolerant systems.
Strong understanding of cloud computing platforms (e.g., AWS, Azure, GCP) and container orchestration technologies (e.g., Kubernetes).
In-Depth knowledge of AWS services including VPC, Lambda, IAM, ELB, EC2, ECS, CloudWatch, API Gateway, S3, SQS, SNS, WAF, X-Ray, and Route53 or GCP services including VPC, Cloud Functions, IAM, Cloud Load Balancing, Compute Engine, Google Kubernetes Engine (GKE), Stackdriver, API Gateway, Cloud Storage, Pub/Sub, Firebase Cloud Messaging, Cloud Armor, Cloud Trace, Cloud DNS
Experience with infrastructure as code tools such as Terraform, Ansible, or similar.
Excellent troubleshooting and problem-solving skills.
Strong communication and leadership skills, with the ability to collaborate effectively with cross-functional teams.
Experience leading and mentoring engineering teams is highly desirable.
Senior DevOps Engineer responsible for cloud infrastructure and deployments. Optimizing AWS services and ensuring system security and reliability for Verizon.
Senior DevOps Engineer responsible for automating infrastructure and building CI/CD pipelines for collaborative robotics company. Collaborating with global engineering teams from the Bangalore office.
Site Reliability Engineer Intern at Tencent working on gaming services and cloud native solutions. Collaborating with global teams to eliminate toil and enhance reliability.
Cloud/DevOps Specialist at N5X managing and optimizing critical cloud infrastructures for Brazilian energy trading. Collaborating with a multidisciplinary team to ensure high availability and performance.
Cloud/Devops Specialist responsible for designing a hybrid architecture combining cloud and on - premises infrastructure for energy trading systems. Collaborating with a multidisciplinary team in a dynamic environment.
Reliability Engineering Specialist utilizing reliability tools and models to improve asset performance at Enbridge. Collaborating across teams to guide investment decisions for safe operations.
DevOps Engineer responsible for structuring and supporting cloud DevOps architecture in Brazil. Working strategically on automation and CI/CD practices with development teams in Pernambuco.
DevSecOps Software Engineer developing secure CI/CD pipelines for Boeing's military software systems. Collaborate with cross - functional teams and implement automation and security best practices.
DevOps Manager responsible for managing a team for multi - cloud solutions supporting the USAF Cloud One project. Focus on scalable cloud - native solutions and CI/CD practices.
Lead Site Reliability Engineer overseeing SRE practices across Azure and GCP platforms. Driving reliability improvements and leading a team at Lloyds Banking Group.