FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Tech Stack
Tools & technologiesAnsibleAWSAzureChefCloudGoogle Cloud PlatformGrafanaKubernetesPrometheusPuppetPythonTerraformGo
About the role
Key responsibilities & impact- Monitor, maintain, and improve the reliability and availability of production systems
- Respond to and resolve incidents and conduct post-mortems to prevent recurrence
- Define and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs)
- Collaborate with development teams to build reliability into services
- Design and implement automation to reduce toil and improve operational efficiency
- Participate in an on-call rotation supporting critical systems
- Contribute to capacity planning and performance optimization
- Document systems, processes, and runbooks
- Build and maintain robust infrastructure and drive engineering best practices
Requirements
What you’ll need- Hands-on experience deploying, managing, and troubleshooting workloads in Kubernetes
- Strong understanding of Pods, Deployments, Services, ConfigMaps, and Ingress
- Experience with Kubernetes cluster management and administration
- Familiarity with Helm for application packaging and deployment
- Understanding of Kubernetes networking, storage, and security best practices
- Bachelor's degree in computer science or relevant degree
- Willingness to work a hybrid model
- Must be fluent in English, both verbal and written
- Strong problem-solving and analytical skills
- Excellent communication skills
- Proactive mindset focused on automation and continuous improvement
- Ability to work effectively under pressure during incident response
- Commitment to a blameless culture and continuous learning
- Experience with cloud platforms such as AWS, GCP, or Azure
- Experience with CI/CD tooling such as GitHub Actions, ArgoCD, or Harness
- Experience with monitoring and observability tools such as Prometheus, Grafana, Coralogix, or OpenTelemetry
- Experience with Infrastructure as Code tools such as Terraform or Pulumi
- Experience with configuration management tools such as Ansible, Puppet, or Chef
- Experience with programming/scripting languages such as Python, Go, or Bash
- Experience contributing to open-source projects
- Familiarity with SRE principles as defined by the Google SRE handbook
- Previous experience in a DevOps or Platform Engineering role
Benefits
Comp & perks- Hybrid work model
- On-call rotation support for critical systems
- Commitment to continuous learning
- Blameless culture
- Opportunity to contribute to open-source projects
- Collaboration with global teams
- Equal employment opportunity
