Senior Site Reliability Engineer improving the reliability of Acuity’s cloud services. Collaborating across teams to define observability standards and incident response in Cork Digital Centre of Excellence.
Responsibilities
Own the availability, reliability, and performance of Reflect’s production environments.
Define, track, and report on service health metrics including uptime, availability, and reliability indicators.
Drive root cause analysis (RCA), analyze system logs and ensure corrective and preventative actions are implemented.
Part of a global team providing operational & escalation coverage, leading incident response and recovery for critical services.
Automate operational workflows to reduce manual toil and improve consistency.
Support and improve deployment processes for features, patches, and hotfixes while maintaining a strong security posture.
Create, maintain, and continuously improve runbooks and standard operating procedures (SOPs).
Design and evolve monitoring, alerting, and observability standards across the platform.
Build and maintain dashboards and alerts that provide clear, actionable insight into system health.
Enable engineers to embed reliability best practices into system design and delivery.
Requirements
5+ years of professional experience in software engineering, SRE, or a related role.
Strong hands‑on experience with Microsoft Azure, including services such as: AKS, Azure Monitor / Log Analytics, Key Vault, ACR, VNets, Managed Identity.
Deep experience with containerized and orchestrated environments (Docker, Kubernetes).
Proven experience operating and supporting production SaaS systems at scale.
A keen eye for detail and a knack for troubleshooting complex issues.
Excellent communication and collaboration skills, with the ability to work effectively across teams.
A passion for learning and a drive to stay up-to-date with emerging technologies.
Bachelor's degree in Computer Science or a related field.
DevOps/IT Apprentice supporting cloud infrastructure and CI/CD pipelines at tech startup. Involves learning, taking ownership, and growing within the engineering team.
DevOps Engineer at Cloud++ collaborating on infrastructure and CI/CD pipelines across multi - cloud environments. Engaging with development teams to ensure reliable and secure releases.
Fullstack Developer at Zenika engaging in impactful tech projects like B2B platforms and architecture modernization. Collaborating with senior consultants in a quality - focused environment.
Production Engineer in a hybrid role ensuring operational performance of applications for a strategic international project. Focusing on automation and optimization within a technical environment at EOLEN.
Senior/Expert DevOps Engineer for AI project in pharmaceutical sector at GECI International. Involves designing, deploying, and operating autonomous AI agents.
DevOps Engineer Intern at Emeria Technologies focusing on Cloud infrastructure design and support. Involves maintaining CI/CD platforms and collaborating with DevOps teams for optimization.
Intern supporting software development infrastructure including CI/CD and cloud integration at Intel. Collaborating with teams to optimize development and release processes.
DevOps Engineer at NetBrain responsible for AWS cloud infrastructure and automating processes. Collaborate with development and security teams to deliver secure solutions efficiently.
Site Reliability Engineer at BlaBlaCar improving CI/CD and tooling for developer efficiency and autonomy. Collaborating with engineering teams to enhance service reliability and facilitate software development.