Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
FactSet

Lead Site Reliability Engineer, Kubernetes

FactSet

. Monitor, maintain, and improve the reliability and availability of production systems .

Posted 9/25/2026full-timeLondon • United KingdomSeniorWebsite

Tech Stack

Tools & technologies
AnsibleAWSAzureChefCloudGoogle Cloud PlatformGrafanaKubernetesPrometheusPuppetPythonTerraformGo

About the role

Key responsibilities & impact
  • Monitor, maintain, and improve the reliability and availability of production systems
  • Respond to and resolve incidents and conduct post-mortems to prevent recurrence
  • Define and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs)
  • Collaborate with development teams to build reliability into services
  • Design and implement automation to reduce toil and improve operational efficiency
  • Participate in an on-call rotation supporting critical systems
  • Contribute to capacity planning and performance optimization
  • Document systems, processes, and runbooks
  • Build and maintain robust infrastructure and drive engineering best practices

Requirements

What you’ll need
  • Hands-on experience deploying, managing, and troubleshooting workloads in Kubernetes
  • Strong understanding of Pods, Deployments, Services, ConfigMaps, and Ingress
  • Experience with Kubernetes cluster management and administration
  • Familiarity with Helm for application packaging and deployment
  • Understanding of Kubernetes networking, storage, and security best practices
  • Bachelor's degree in computer science or relevant degree
  • Willingness to work a hybrid model
  • Must be fluent in English, both verbal and written
  • Strong problem-solving and analytical skills
  • Excellent communication skills
  • Proactive mindset focused on automation and continuous improvement
  • Ability to work effectively under pressure during incident response
  • Commitment to a blameless culture and continuous learning
  • Experience with cloud platforms such as AWS, GCP, or Azure
  • Experience with CI/CD tooling such as GitHub Actions, ArgoCD, or Harness
  • Experience with monitoring and observability tools such as Prometheus, Grafana, Coralogix, or OpenTelemetry
  • Experience with Infrastructure as Code tools such as Terraform or Pulumi
  • Experience with configuration management tools such as Ansible, Puppet, or Chef
  • Experience with programming/scripting languages such as Python, Go, or Bash
  • Experience contributing to open-source projects
  • Familiarity with SRE principles as defined by the Google SRE handbook
  • Previous experience in a DevOps or Platform Engineering role

Benefits

Comp & perks
  • Hybrid work model
  • On-call rotation support for critical systems
  • Commitment to continuous learning
  • Blameless culture
  • Opportunity to contribute to open-source projects
  • Collaboration with global teams
  • Equal employment opportunity