Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Gen

Principal Site Reliability Engineer

Gen

. Support Gen Digital products and services, including Gen Insurance and Engine by Gen, delivering trusted digital and financial experiences at scale .

Posted 9/21/2026full-timeTempe • Arizona • United StatesLead💰 $150,000 - $160,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates extensive expertise in Site Reliability Engineering (SRE) practices, cloud-native architecture, and automation, with a strong focus on AWS and Kubernetes. Proven ability to lead cross-functional teams in implementing operational excellence and reliability strategies while ensuring compliance with security standards.

Highest-signal resume keywords
Site Reliability Engineering (SRE)AWS ExpertiseKubernetes ProficiencyInfrastructure as Code (Terraform)CI/CD Systems

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Cloud-Native ArchitectureDistributed Systems ReliabilityIncident ManagementObservabilityAutomationMicroservices DesignCapacity PlanningError BudgetsProduction ReadinessProgramming (Python or Go)
Soft Skills
Excellent Communication SkillsMentoringCross-Team CollaborationStrategic DirectionHands-On Execution
Tools & Technologies
TerraformJenkinsCircleCIGitHub ActionsEKSECSGKEAI-Assisted Engineering Tools
Industry Keywords
PCI ComplianceSOC 2 ComplianceNIST StandardsOperational ExcellenceDisaster Recovery

Tech Stack

Tools & technologies
AWSCloudDistributed SystemsGoogle Cloud PlatformJenkinsKubernetesPythonTerraformGo

About the role

Key responsibilities & impact
  • Support Gen Digital products and services, including Gen Insurance and Engine by Gen, delivering trusted digital and financial experiences at scale
  • Define and lead the platform's reliability, scalability, and operational excellence strategy
  • Partner with application, platform, security, and infrastructure teams to build resilient systems and modernize cloud-native architecture
  • Shape SRE practices including service reliability frameworks, production readiness, observability, incident management, automation, and capacity planning
  • Collaborate with SRE teams on global standards, shared practices, and strategic platform direction
  • Define long-term reliability strategy and operational standards across services and teams
  • Lead architecture decisions for highly available, fault-tolerant distributed systems on AWS, GCP, Kubernetes, and GKE
  • Guide platform architecture, deployment patterns, and infrastructure automation using Terraform and infrastructure-as-code tooling
  • Design and implement observability capabilities covering metrics, logging, tracing, alerting, and dashboards
  • Lead major incident response and improve escalation, response, and post-incident review processes
  • Review RFCs, set engineering guardrails, and provide technical leadership on high-impact initiatives
  • Champion resilience engineering, disaster recovery, failover design, capacity forecasting, and business continuity planning
  • Partner with security and compliance stakeholders on infrastructure security and operational controls for PCI and SOC 2 environments
  • Improve CI/CD systems and software delivery workflows
  • Reduce operational toil through automation, self-service platform capabilities, and better engineering abstractions
  • Mentor senior engineers and technical leads
  • Align with global SRE counterparts to standardize practices and share learnings
  • Advise engineering and product leadership on reliability tradeoffs, infrastructure investments, and operational risk

Requirements

What you’ll need
  • 8+ years of experience in SRE, DevOps, platform engineering, or infrastructure engineering
  • Strong record of leading large-scale cloud initiatives in production environments
  • Deep expertise in AWS and Kubernetes
  • Hands-on experience designing, operating, and evolving large-scale containerized microservice-based systems
  • Strong experience with infrastructure as code, especially Terraform
  • Experience with AWS services supporting distributed systems and container platforms such as EKS, ECS, or GKE
  • Proven success defining and implementing SRE practices such as SLIs, SLOs, error budgets, incident management, observability, and production readiness standards
  • Strong understanding of distributed systems reliability, performance, scaling, availability engineering, and failure-mode analysis
  • Experience leading complex cross-team technical initiatives and influencing architecture, standards, and engineering practices beyond direct reporting lines
  • Strong background in CI/CD and delivery engineering
  • Experience with Jenkins, CircleCI, and GitHub Actions
  • Proficiency in at least one programming language such as Python or Go
  • Ability to build automation, tooling, and maintainable production-quality code
  • Experience operating in security-conscious and compliant environments
  • Familiarity with PCI, SOC 2, and NIST
  • Excellent written and verbal communication skills
  • Experience using AI-assisted and agentic engineering tools
  • Ability to balance strategic direction with hands-on execution
  • Ability to work from one of the company's offices at least three days per week

Benefits

Comp & perks
  • Bonus offered
  • Hybrid work arrangement
  • Interview process includes Zoom interview