Cloud Site Reliability Engineer managing Solace Cloud services across leading cloud providers. Ensuring reliability, handling incidents, and collaborating with customers for operational excellence.
Responsibilities
Ensuring that the Solace Cloud Services are healthy and reliable, and that SLAs are being met
Assist in implementing the infrastructure tooling, observability, and automation
Contribute to making the production operations more efficient, less error-prone
Handle production incidents in enterprise-grade multi-cloud environments according to industry-standard Incident management process
Process handling service requests and provisioning by the customers
Manage customer escalations and drive resolution in mission-critical, high-impact production environments
Work directly with customers to identify, troubleshoot, and resolve operational issues
Expert debugging knowledge in Linux and Kubernetes to detect operational issues
Be on-call rotation and provide 24x7 off-hours support
Requirements
Proven expertise with public cloud providers (AWS, Azure, GCP) services & features
Proven expertise with cloud Kubernetes infrastructure platforms such as AWS Elastic Kubernetes Service, Azure Kubernetes Service, Google Kubernetes Service
Hands-on experience with Monitoring tools like Datadog, Kibana, Prometheus etc.
Hands-on experience with Infrastructure Automation using Terraform, Cloud Formation
Hands-on expertise in debugging production alerts
Strong understanding of Linux Operating Systems
Programmer in languages such as Groovy, Python, and Go
Hands-on experience with AI tools and a strong interest in advancing AI capabilities.
Benefits
Balance matters – We believe work should fit into your life, not the other way.
Hybrid-first – Flexibility is built into how we work, so everyone feels included and empowered.
Values-driven – We live and breathe our core values: craftsmanship, trust, courage, freedom, momentum, humility, and human experience.
Growth mindset – Our training programs are designed to help you level up, fast.
Customer Obsessed – We’re proud of our world-class customer lineup.
Senior Cloud Site Reliability Engineer ensuring reliability and health of Solace Cloud Services with hands - on cloud operations expertise. Lead incident management and customer support for high - impact environments.
DevOps Engineer designing and operating AWS infrastructure within industrial IoT environments. Working on systems that ensure security, resilience, and end - to - end observability.
Sr. Site Reliability Engineer (SRE) III providing technical solutions for the federal government. Collaborating in a high - performing team focused on reliability and application scalability.
Senior Linux System Engineer developing and maintaining Linux server infrastructure for Th. Geyer GmbH. Collaborating on ERP systems and CI/CD processes while ensuring system performance and security.
Cloud Platform Engineer (ML DevOps) developing and managing CI/CD pipelines for ML workflows in a leading insurance company. Collaborating with data scientists and ensuring infrastructure security and compliance.
Platform Engineer leading the development of cloud application platforms for Allstate. Responsible for cloud infrastructure for ML experimentation and production deployments.
DevOps Engineer developing and managing container platforms for client solutions at Booz Allen Hamilton. Utilizing cloud technologies to enhance capabilities and secure deployments.
Senior DevOps/Platform Engineer automating cloud infrastructure and optimizing delivery pipelines at S&P Global Mobility. Collaborating with teams to enhance product reliability and security.
DevOps Engineer responsible for maintaining and enhancing AWS/EKS platform for energy transition products. Ensuring platform stability, security compliance, and streamlined deployment processes.