FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates extensive expertise in Site Reliability Engineering (SRE) practices, cloud-native architecture, and automation, with a strong focus on AWS and Kubernetes. Proven ability to lead cross-functional teams in implementing operational excellence and reliability strategies while ensuring compliance with security standards.
Highest-signal resume keywords
Site Reliability Engineering (SRE)AWS ExpertiseKubernetes ProficiencyInfrastructure as Code (Terraform)CI/CD Systems
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Cloud-Native ArchitectureDistributed Systems ReliabilityIncident ManagementObservabilityAutomationMicroservices DesignCapacity PlanningError BudgetsProduction ReadinessProgramming (Python or Go)
Soft Skills
Excellent Communication SkillsMentoringCross-Team CollaborationStrategic DirectionHands-On Execution
Tools & Technologies
TerraformJenkinsCircleCIGitHub ActionsEKSECSGKEAI-Assisted Engineering Tools
Industry Keywords
PCI ComplianceSOC 2 ComplianceNIST StandardsOperational ExcellenceDisaster Recovery
Tech Stack
Tools & technologiesAWSCloudDistributed SystemsGoogle Cloud PlatformJenkinsKubernetesPythonTerraformGo
About the role
Key responsibilities & impact- Support Gen Digital products and services, including Gen Insurance and Engine by Gen, delivering trusted digital and financial experiences at scale
- Define and lead the platform's reliability, scalability, and operational excellence strategy
- Partner with application, platform, security, and infrastructure teams to build resilient systems and modernize cloud-native architecture
- Shape SRE practices including service reliability frameworks, production readiness, observability, incident management, automation, and capacity planning
- Collaborate with SRE teams on global standards, shared practices, and strategic platform direction
- Define long-term reliability strategy and operational standards across services and teams
- Lead architecture decisions for highly available, fault-tolerant distributed systems on AWS, GCP, Kubernetes, and GKE
- Guide platform architecture, deployment patterns, and infrastructure automation using Terraform and infrastructure-as-code tooling
- Design and implement observability capabilities covering metrics, logging, tracing, alerting, and dashboards
- Lead major incident response and improve escalation, response, and post-incident review processes
- Review RFCs, set engineering guardrails, and provide technical leadership on high-impact initiatives
- Champion resilience engineering, disaster recovery, failover design, capacity forecasting, and business continuity planning
- Partner with security and compliance stakeholders on infrastructure security and operational controls for PCI and SOC 2 environments
- Improve CI/CD systems and software delivery workflows
- Reduce operational toil through automation, self-service platform capabilities, and better engineering abstractions
- Mentor senior engineers and technical leads
- Align with global SRE counterparts to standardize practices and share learnings
- Advise engineering and product leadership on reliability tradeoffs, infrastructure investments, and operational risk
Requirements
What you’ll need- 8+ years of experience in SRE, DevOps, platform engineering, or infrastructure engineering
- Strong record of leading large-scale cloud initiatives in production environments
- Deep expertise in AWS and Kubernetes
- Hands-on experience designing, operating, and evolving large-scale containerized microservice-based systems
- Strong experience with infrastructure as code, especially Terraform
- Experience with AWS services supporting distributed systems and container platforms such as EKS, ECS, or GKE
- Proven success defining and implementing SRE practices such as SLIs, SLOs, error budgets, incident management, observability, and production readiness standards
- Strong understanding of distributed systems reliability, performance, scaling, availability engineering, and failure-mode analysis
- Experience leading complex cross-team technical initiatives and influencing architecture, standards, and engineering practices beyond direct reporting lines
- Strong background in CI/CD and delivery engineering
- Experience with Jenkins, CircleCI, and GitHub Actions
- Proficiency in at least one programming language such as Python or Go
- Ability to build automation, tooling, and maintainable production-quality code
- Experience operating in security-conscious and compliant environments
- Familiarity with PCI, SOC 2, and NIST
- Excellent written and verbal communication skills
- Experience using AI-assisted and agentic engineering tools
- Ability to balance strategic direction with hands-on execution
- Ability to work from one of the company's offices at least three days per week
Benefits
Comp & perks- Bonus offered
- Hybrid work arrangement
- Interview process includes Zoom interview
