FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in Site Reliability Engineering (SRE) practices, cloud architecture, and observability tooling, with a strong focus on improving system reliability and performance. Proven ability to mentor teams, implement Infrastructure as Code (IaC), and collaborate across multiple stakeholders to enhance incident management and operational efficiency.
Highest-signal resume keywords
Site Reliability Engineering (SRE)Cloud Architecture (AWS, Azure, GCP)Infrastructure as Code (IaC)Observability Tooling (Grafana, Datadog, ELK, Prometheus)Incident Management Practices
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Cloud ArchitectureInfrastructure as Code (IaC)Automation ScriptingDebugging Distributed SystemsIncident ManagementPerformance MonitoringTechnical Debt RemediationSLO EstablishmentArchitecture ReviewProduction Readiness
Soft Skills
Excellent CommunicationFacilitation SkillsMentoring AbilityPsychological SafetyCuriosity
Tools & Technologies
TerraformGitHub ActionsAnsiblePackerKubernetesGrafanaDatadogELKPrometheus
Certifications & Qualifications
SOC2ISO27001PCI Compliance
Industry Keywords
SaaSMulti-TenantCloud MigrationITILIncident Command Frameworks
Tech Stack
Tools & technologiesAnsibleAWSAzureCloudDistributed SystemsGoogle Cloud PlatformGrafanaKubernetesPackerPrometheusTerraform
About the role
Key responsibilities & impact- Implement and iterate on SRE practices across the team
- Deliver measurable improvements in availability, performance, and incident response
- Partner with Platform and LiveOps to build and run secure, efficient, scalable solutions
- Instrument services with metrics, tracing, and logging through observability stacks
- Support on-call readiness and run blameless post-mortems
- Feed incident-response lessons into engineering and product planning
- Implement and improve IaC and CI/CD practices using Terraform, GitHub Actions, Ansible, Packer, and Kubernetes
- Establish and evolve SLOs, on-call structures, toil reduction, and production readiness practices
- Review architecture proposals and prototype reliability solutions
- Debug complex production issues and make trade-off decisions involving speed, resilience, and cost
- Identify systemic reliability risks and prioritise technical debt remediation
- Mentor engineers and foster psychological safety, curiosity, and accountability
- Collaborate with Security and Compliance teams to integrate SOC2, ISO27001, and PCI requirements
Requirements
What you’ll need- Proven experience leading SRE, DevOps, or Cloud/Software engineering teams, typically 5+ engineers, within a SaaS or complex multi-product environment
- Strong technical foundation in cloud architecture (AWS, Azure, GCP, or equivalent) and hybrid or cloud migration patterns
- Ability to read and review application and infrastructure code
- Comfortable writing automation, debugging distributed systems issues, and making architecture decisions grounded in hands-on experience
- Solid background in observability tooling and incident management practices, including monitoring, tracing, and alerting
- Experience using observability data to diagnose real production problems
- Excellent communication and facilitation skills across Product, Platform, Security, Support and LiveOps stakeholders
- Skilled mentoring ability
- Passion for reliability and user experience as a culture
- Experience with large-scale multi-tenant SaaS or hybrid cloud migrations (nice to have)
- Awareness of ITIL or incident command frameworks adapted for modern SRE cultures (nice to have)
- Familiarity with Grafana, Datadog, ELK, and Prometheus (nice to have)
- Experience working with distributed global teams (nice to have)
- Experience working in software development and strong familiarity with software lifecycle (nice to have)
Benefits
Comp & perks- 25 Days Annual Leave + bank holidays
- Option to buy up to 10 extra days
- Days of Difference – Up to 3 extra days off for volunteering
- Pension Contributions – 5% employer match
- Income Protection – Up to 75% salary cover for long-term illness
- Life Assurance – 4x salary tax-free lump sum
- Critical Illness Cover – £25,000 lump sum, extendable to dependents
- Private Medical Insurance
- Health Cash Plan – Claim back physio, therapies & more
- Dental Insurance
- Affinity Groups – Join employee-led communities
- Bounty Bonus – Refer a friend & get rewarded
- Inclusive and diverse workplace
- Equal opportunity employer
- Recruitment adjustments or accommodations available
