Senior Site Reliability Engineer focused on building reliable, scalable infrastructure at a tech company. Driving best practices in observability, incident response, and engineering collaboration.
Responsibilities
Design, build, and maintain highly available, scalable, and fault-tolerant systems
Lead reliability improvements across production and non-production environments
Own and improve monitoring, alerting, and observability platforms
Drive incident response, root cause analysis, and post-incident reviews
Implement automation to reduce manual operational work
Partner with Engineering, Security, and Product to support platform needs
Establish and track SLIs, SLOs, and error budgets
Lead capacity planning and performance tuning efforts
Improve deployment, CI/CD, and infrastructure-as-code practices
Identify and mitigate reliability and scalability risks before they impact customers
Mentor and guide junior engineers and contribute to team technical standards
Participate in on-call rotation and help mature on-call processes
Requirements
6+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related roles
Strong experience with cloud platforms (AWS, Azure, or GCP)
Proficiency with infrastructure as code (Terraform, CloudFormation, Pulumi, etc.)
Experience with containerization and orchestration (Docker, Kubernetes)
Strong Linux systems administration and networking fundamentals
Experience building and maintaining CI/CD pipelines
Hands-on experience with monitoring and observability tools (Datadog, Prometheus, Grafana, New Relic, etc.)
Strong troubleshooting and incident management skills
Experience with scripting and automation (Python, Bash, Go, or similar)
Lead DevOps Engineer focused on AWS and Azure data platform solutions. Collaborating with teams to deliver scalable, secure, and highly available solutions.
DevOps Engineer working at GRÜN Software Group to automate and maintain stable infrastructures. Collaborating with teams to improve deployments and processes for better performance.
Linux System Administrator managing IT infrastructures for educational institutions and research. Collaborating on DevOps and HPC projects while ensuring system security and performance.
Azure SRE Engineer responsible for designing and maintaining secure, scalable Azure cloud infrastructure. Driving automation and operational excellence for leading organizations in technology transformation.
Senior Manager of Site Reliability Engineering overseeing Workday Kubernetes based platform. Leading teams while ensuring high availability and collaborating with federal agencies.
Site Reliability Engineer focusing on AWS cloud environments, SRE practices, and system reliability within GFT's team. Collaborating on cloud migrations and observability initiatives.
Senior DevOps Analyst enhancing infrastructure automation in a transformative technology firm. Collaborating on innovative projects in sectors like healthcare, finance, and utilities in Brazil.
Consultant at Minsait supporting technical decisions in infrastructure automation and developing solutions. Collaborating with teams for maintaining and evolving automation platforms.
Practical Trainee focusing on hardware reliability engineering at Sonova. Support reliability improvement initiatives and work closely with experienced engineers on real - life product challenges.