Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
GE Vernova

System Reliability Engineering Lead

GE Vernova

. Serve as the hands-on technical authority for production stability across the GridOS SaaS portfolio .

Posted 9/25/2026full-timeRemote • United StatesSenior💰 $151,800 - $227,700 per yearWebsite

Tech Stack

Tools & technologies
AnsibleAWSCloudEC2GrafanaKubernetesPrometheusSplunkTerraform

About the role

Key responsibilities & impact
  • Serve as the hands-on technical authority for production stability across the GridOS SaaS portfolio
  • Own Change Management and approve or halt production deployments based on system health
  • Drive high reliability and engineering excellence across a distributed team
  • Architect and implement standardized, secure cloud infrastructure provisioning
  • Automate account provisioning to accelerate customer onboarding
  • Define and build the standardized Middle-Mile software delivery platform using Backstage, ArgoCD, and GitHub Actions
  • Establish global handover protocols and 24/7 operational coverage across US, India, and Mexico time zones
  • Establish and own enterprise-wide SLOs, SLIs, and error budgets
  • Serve as final technical authority for production releases and enforce security and performance quality gates
  • Implement Canary and Blue/Green deployments with automated rollback capabilities
  • Build and mature the SRE Center for Enablement with coaching, templates, and reliability patterns
  • Lead incident response for Sev1/Sev2 events and P1 escalations
  • Facilitate blameless Root Cause Analysis and own the post-incident lifecycle
  • Architect and validate backup and disaster recovery strategies, including cross-region failover and automated recovery testing
  • Own FinOps, cloud cost optimization, and long-term capacity planning
  • Serve as primary SRE point of contact for North American utility customers
  • Participate in customer reviews, incident communications, and service health reporting
  • Lead a distributed team of 8 SRE engineers across Hyderabad and Querétaro
  • Set technical direction, assign tasks, own deliverables, mentor engineers, and provide performance feedback to the people leader of record
  • Travel up to 10% to customer sites and team locations as needed

Requirements

What you’ll need
  • Deep expertise in AWS core services: EC2, EKS, RDS, S3, and IAM
  • Experience with AWS management tools including CloudTrail and CloudWatch
  • Advanced mastery of Kubernetes internals and EKS cluster operations across multi-region architectures
  • Expert knowledge of ArgoCD, GitHub Actions, and GitOps-first workflows
  • Proficiency in Infrastructure as Code using Terraform
  • Proficiency in configuration management via Ansible
  • Hands-on experience with Prometheus, Grafana, Splunk or Datadog, and OpenTelemetry
  • Experience with cloud cost optimization, reserved instance management, right-sizing, and long-term capacity planning for multi-tenant SaaS platforms
  • 12+ years in software engineering, cloud operations, or infrastructure roles
  • 8–10 years of hands-on experience in SRE, Platform Engineering, Cloud Operations, or Production Support for large-scale, distributed SaaS applications
  • Proven track record leading distributed engineering teams as a player-coach while remaining hands-on with architecture, automation, and incident response
  • Exceptional troubleshooting skills under pressure and a “Fire Marshal” mindset toward investigation and proactive inspection
  • Experience working directly with enterprise customers on production reliability, incident communication, and service-level reporting
  • Must pass customer-mandated background screening for access to critical infrastructure environments
  • Must be legally authorized to work in the United States
  • Must complete a drug screen, as applicable
  • General shift during US business hours and on-call availability for P1/Sev1 incidents
  • Up to 10% travel to customer sites and team locations
  • Desired: NERC CIP, SOC2, ISO 27001, or IEC 62443 knowledge/experience
  • Desired: experience in highly regulated industries such as utilities, financial services, or critical national infrastructure
  • Desired certifications: AWS DevOps Engineer—Professional or Solutions Architect—Associate/Professional, CKA, SRE Practitioner, and AWS FinOps Practitioner or equivalent

Benefits

Comp & perks
  • Discretionary annual bonus
  • Medical, dental, vision, and prescription drug coverage
  • Health Coach from GE Vernova, a 24/7 nurse-based resource
  • Employee Assistance Program with 24/7 confidential assessment, counseling, and referral services
  • GE Vernova Retirement Savings Plan
  • Tax-advantaged 401(k) savings opportunity with company matching contributions and company retirement contributions
  • Fidelity resources and financial planning consultants
  • Tuition assistance
  • Adoption assistance
  • Paid parental leave
  • Disability benefits
  • Life insurance
  • 12 paid holidays
  • Permissive time off
  • Professional development opportunities
  • Relocation assistance not provided