FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

System Reliability Engineering Lead
GE Vernova. Serve as the hands-on technical authority for production stability across the GridOS SaaS portfolio .
Tech Stack
Tools & technologiesAnsibleAWSCloudEC2GrafanaKubernetesPrometheusSplunkTerraform
About the role
Key responsibilities & impact- Serve as the hands-on technical authority for production stability across the GridOS SaaS portfolio
- Own Change Management and approve or halt production deployments based on system health
- Drive high reliability and engineering excellence across a distributed team
- Architect and implement standardized, secure cloud infrastructure provisioning
- Automate account provisioning to accelerate customer onboarding
- Define and build the standardized Middle-Mile software delivery platform using Backstage, ArgoCD, and GitHub Actions
- Establish global handover protocols and 24/7 operational coverage across US, India, and Mexico time zones
- Establish and own enterprise-wide SLOs, SLIs, and error budgets
- Serve as final technical authority for production releases and enforce security and performance quality gates
- Implement Canary and Blue/Green deployments with automated rollback capabilities
- Build and mature the SRE Center for Enablement with coaching, templates, and reliability patterns
- Lead incident response for Sev1/Sev2 events and P1 escalations
- Facilitate blameless Root Cause Analysis and own the post-incident lifecycle
- Architect and validate backup and disaster recovery strategies, including cross-region failover and automated recovery testing
- Own FinOps, cloud cost optimization, and long-term capacity planning
- Serve as primary SRE point of contact for North American utility customers
- Participate in customer reviews, incident communications, and service health reporting
- Lead a distributed team of 8 SRE engineers across Hyderabad and Querétaro
- Set technical direction, assign tasks, own deliverables, mentor engineers, and provide performance feedback to the people leader of record
- Travel up to 10% to customer sites and team locations as needed
Requirements
What you’ll need- Deep expertise in AWS core services: EC2, EKS, RDS, S3, and IAM
- Experience with AWS management tools including CloudTrail and CloudWatch
- Advanced mastery of Kubernetes internals and EKS cluster operations across multi-region architectures
- Expert knowledge of ArgoCD, GitHub Actions, and GitOps-first workflows
- Proficiency in Infrastructure as Code using Terraform
- Proficiency in configuration management via Ansible
- Hands-on experience with Prometheus, Grafana, Splunk or Datadog, and OpenTelemetry
- Experience with cloud cost optimization, reserved instance management, right-sizing, and long-term capacity planning for multi-tenant SaaS platforms
- 12+ years in software engineering, cloud operations, or infrastructure roles
- 8–10 years of hands-on experience in SRE, Platform Engineering, Cloud Operations, or Production Support for large-scale, distributed SaaS applications
- Proven track record leading distributed engineering teams as a player-coach while remaining hands-on with architecture, automation, and incident response
- Exceptional troubleshooting skills under pressure and a “Fire Marshal” mindset toward investigation and proactive inspection
- Experience working directly with enterprise customers on production reliability, incident communication, and service-level reporting
- Must pass customer-mandated background screening for access to critical infrastructure environments
- Must be legally authorized to work in the United States
- Must complete a drug screen, as applicable
- General shift during US business hours and on-call availability for P1/Sev1 incidents
- Up to 10% travel to customer sites and team locations
- Desired: NERC CIP, SOC2, ISO 27001, or IEC 62443 knowledge/experience
- Desired: experience in highly regulated industries such as utilities, financial services, or critical national infrastructure
- Desired certifications: AWS DevOps Engineer—Professional or Solutions Architect—Associate/Professional, CKA, SRE Practitioner, and AWS FinOps Practitioner or equivalent
Benefits
Comp & perks- Discretionary annual bonus
- Medical, dental, vision, and prescription drug coverage
- Health Coach from GE Vernova, a 24/7 nurse-based resource
- Employee Assistance Program with 24/7 confidential assessment, counseling, and referral services
- GE Vernova Retirement Savings Plan
- Tax-advantaged 401(k) savings opportunity with company matching contributions and company retirement contributions
- Fidelity resources and financial planning consultants
- Tuition assistance
- Adoption assistance
- Paid parental leave
- Disability benefits
- Life insurance
- 12 paid holidays
- Permissive time off
- Professional development opportunities
- Relocation assistance not provided