FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in Site Reliability Engineering, with a strong focus on Kubernetes, AWS, and observability practices. Proven ability to lead teams, manage performance, and drive operational excellence while ensuring compliance in regulated environments.
Highest-signal resume keywords
Site Reliability Engineering LeadershipKubernetes in ProductionAWS at Meaningful ScaleInfrastructure as Code with TerraformIncident Response and SLO Management
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Site Reliability EngineeringCloud Infrastructure EngineeringDevOps PracticesObservability PrinciplesAI Infrastructure BuildingTerraformGitLab CIGitHub ActionsJenkinsDocker
Soft Skills
Coaching Technical CraftClear Written CommunicationRelationship BuildingPerformance AssessmentHandling Underperformance
Tools & Technologies
KubernetesAWSPostgreSQLDNSTLSCI Infrastructure
Industry Keywords
Regulated EnvironmentsIncident ResponseSLOsError BudgetsOperational Load Management
Tech Stack
Tools & technologiesAWSCloudDNSDockerJenkinsKubernetesPostgresShell ScriptingTerraform
About the role
Key responsibilities & impact- Lead the Site Reliability Engineering team as a 60% individual contributor and 40% leadership role
- Own direct reports' onboarding, feedback, performance assessment, progression and hiring
- Steer the team's focus against company goals and represent the team across engineering and senior leadership
- Set technical direction and review the team's work
- Own SRE goals, prioritization, support rotation and on-call model
- Oversee Kubernetes, AWS, PostgreSQL, DNS and TLS, and CI infrastructure
- Develop the reliability practice, including SLOs, error budgets, incident response and observability
- Partner with Security on threats, patching, infrastructure controls, audit and compliance obligations
- Manage platform vendor relationships, renewals and commercial conversations
- Improve operational load versus project delivery balance and extend the SLO framework across teams
Requirements
What you’ll need- Experience leading an SRE, infrastructure or platform engineering team
- Ownership of reports' growth, performance and career progression
- Experience coaching technical craft and soft skills
- Experience handling underperformance directly and early
- Experience hiring engineers and assessing engineering quality
- Hands-on background in site reliability, DevOps or cloud infrastructure engineering
- Kubernetes in production
- AWS at meaningful scale
- Hands-on AI building, enablement, and scaling AI infrastructure
- Solid observability practices and principles
- Infrastructure as code with Terraform
- CI/CD systems such as GitLab CI, GitHub Actions or Jenkins
- Docker and shell scripting
- Experience running a reliability practice: incident response, on-call, SLOs and error budgets
- Understanding and history of working in regulated environments
- Ability to prioritize operational load and project work
- Clear written communication for asynchronous work
- Ability to build relationships across teams
- English application materials required
Benefits
Comp & perks- work from anywhere
- flexible paid time off
- flexible working hours (we are async)
- 16 weeks paid parental leave
- budget towards co-working spaces, learning and wellness (including gym memberships)
- mental health support services
- stock options
- home office budget & IT equipment
- life-work balance and schedule flexibility
- employee resource groups (Women, Disability, Queer, Minorities in Tech)
- accommodation support during interviews and beyond
