Senior Site Reliability Engineer designing and implementing high-reliability platforms for Broadridge. Collaborating with teams across hybrid environments and driving automation and efficiency in service delivery.
Responsibilities
Design and implement high-availability, fault-tolerant architectures across on-prem and cloud platforms (AWS)
Lead multi-region DR planning, implementation, and testing, including RTO/RPO definition and validation
Define and enforce SLOs, SLIs, and error budgets to balance reliability with delivery velocity
Drive self-healing automation and proactive remediation strategies
Build and maintain infrastructure using Terraform and configuration management tools (e.g., Chef)
Develop automation to eliminate manual operational tasks (TOIL reduction)
Create reusable modules, pipelines, and guardrails for standardized deployments
Automate certificate lifecycle management, key rotation, and security updates
Design and implement end-to-end observability (metrics, logs, traces, synthetic monitoring)
Build dashboards, alerts, and runbooks to enable fast detection and resolution of incidents
Perform root cause analysis (RCA) and lead post-incident reviews with actionable follow-ups
Engineer and operate platforms on AWS, including services such as EKS, EC2, RDS/Aurora, Lambda, API Gateway, CloudFront, WAF, ALB/NLB, CloudWatch, X-Ray, IAM, Secrets Manager
Lead cloud migrations and modernization initiatives, including legacy system refactoring
Identify and resolve performance bottlenecks through testing and analysis
Design and support CI/CD pipelines enabling safe, repeatable deployments
Partner with security and legal teams to meet regulatory and compliance requirements (e.g., data residency, GDPR-related controls)
Requirements
8+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or Systems Engineering
Strong programming experience in Python, Java, or similar languages
Deep experience with Linux/Unix systems
Hands-on expertise with AWS and cloud-native architectures
Proven experience with Terraform and Infrastructure as Code
Strong understanding of networking, security, and distributed systems
Lead DevOps Engineer focused on AWS and Azure data platform solutions. Collaborating with teams to deliver scalable, secure, and highly available solutions.
DevOps Engineer working at GRÜN Software Group to automate and maintain stable infrastructures. Collaborating with teams to improve deployments and processes for better performance.
Linux System Administrator managing IT infrastructures for educational institutions and research. Collaborating on DevOps and HPC projects while ensuring system security and performance.
Azure SRE Engineer responsible for designing and maintaining secure, scalable Azure cloud infrastructure. Driving automation and operational excellence for leading organizations in technology transformation.
Senior Manager of Site Reliability Engineering overseeing Workday Kubernetes based platform. Leading teams while ensuring high availability and collaborating with federal agencies.
Site Reliability Engineer focusing on AWS cloud environments, SRE practices, and system reliability within GFT's team. Collaborating on cloud migrations and observability initiatives.
Senior DevOps Analyst enhancing infrastructure automation in a transformative technology firm. Collaborating on innovative projects in sectors like healthcare, finance, and utilities in Brazil.
Consultant at Minsait supporting technical decisions in infrastructure automation and developing solutions. Collaborating with teams for maintaining and evolving automation platforms.
Practical Trainee focusing on hardware reliability engineering at Sonova. Support reliability improvement initiatives and work closely with experienced engineers on real - life product challenges.