SRE Observability SLO Engineer for GE Vernova’s GridOS Platform Engineering team. Building telemetry stack in SaaS reliability for critical energy infrastructure.
Responsibilities
Implement organization-wide telemetry standards covering metrics, logs, and distributed traces across all GridOS SaaS services.
Implement metrics collection for Kubernetes-hosted services (EKS/Rancher) including pod-level, namespace-level, and cluster-level metrics.
Publish and maintain an Observability Runbook library covering onboarding, alert tuning, and dashboard standards for Platform SRE and Production DevOps teams.
Partner with product engineering, Platform SRE, and customer stakeholders to define meaningful Service Level Indicators (SLIs) and Service Level Objectives (SLOs) per product and customer tier.
Build and maintain SLO tooling — error budget burn-rate alerts, burn-rate dashboards, and automated SLO compliance reports.
Design and build operational dashboards covering availability, latency, error rates, and saturation (the 'Golden Signals') for every GridOS SaaS product.
Create executive-level dashboards for SRE leadership and customer-facing uptime/availability reports aligned to contractual SLAs.
Conduct periodic observability health reviews to identify gaps in coverage, reduce MTTD (Mean Time to Detect), and improve MTTR (Mean Time to Resolve).
Requirements
2–3 years in SRE, observability engineering, or infrastructure reliability roles.
Deep expertise with at least one major observability platform — Datadog, Grafana + Prometheus, AWS CloudWatch, Dynatrace, or New Relic.
Hands-on experience implementing SLIs, SLOs, and error budget burn-rate alerting in a production SaaS environment.
Strong understanding of distributed systems telemetry: metrics (Prometheus/CloudWatch), structured logging (CloudWatch Logs Insights, ELK), and distributed tracing (OpenTelemetry, AWS X-Ray).
Experience with Kubernetes observability — kube-state-metrics, node exporters, Helm-deployed monitoring stacks, and namespace-level resource metrics.
Proficiency in at least one query/visualization language: PromQL, Splunk SPL, Datadog Query Language, or CloudWatch Logs Insights query syntax.
Experience designing alerting strategies that minimize alert fatigue through symptom-based and burn-rate approaches.
Scripting skills in Python and/or Bash for automation of monitoring configuration and report generation.
As Learning Content Engineer, developing and enhancing training content for Cloud and DevOps. Engaging in creating practical learning materials from basics to advanced topics.
AWS DevOps Microservices Engineer at Solventum designing secure and scalable AWS infrastructures. Collaborating with diverse teams for innovative healthcare solutions using cloud technology.
DevOps Engineer building and maintaining Catena’s scalable platform infrastructure. Collaborating with engineers to enhance CI/CD pipelines and support cloud - native workloads on AWS.
Platform System Reliability Engineer focused on operations of EKS Kubernetes environment for GE Vernova's SaaS grid products. Responsible for the full lifecycle of production clusters from performance tuning to securing infrastructure.
DevOps Engineer responsible for building and operating automation services using Ansible for Rabobank. Collaborating with teams to ensure stable, secure, and auditable infrastructure across multiple servers.
Engineer collaborating with AI startups to enhance their systems and contribute to OpenAI's products. Engaging in technical problem - solving and building relationships within the startup ecosystem.
Senior Software Engineer designing and developing software applications for space technologies. Leading technical decisions and collaborating on innovative solutions to enhance national security.
DevOps Engineer responsible for web application operations and developer experience at Nitrado, a global game server hosting provider. Collaborating with developers on automation, Kubernetes, and Docker management.