Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Circle

Senior Site Reliability Engineer – Infra Ops

Circle

. Design, build, and operate Kubernetes platforms for secure, highly available, and scalable critical production services across hybrid and public-cloud environments .

Posted 10/1/2026full-timeRemote • California • United StatesSenior💰 $152,500 - $205,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates extensive expertise in Kubernetes platform design and operation, with strong capabilities in Terraform for infrastructure as code. Proficient in backend development using Go, Python, or JavaScript/TypeScript, and skilled in implementing observability practices and incident management processes.

Highest-signal resume keywords
Kubernetes ExpertiseTerraform ProficiencyBackend Development in Go, Python, or JavaScript/TypeScriptObservability and Troubleshooting SkillsCI/CD and Deployment Automation

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
KubernetesTerraformGoPythonJavaScriptTypeScriptCI/CDGitOpsSLIsSLOs
Soft Skills
Clear CommunicationStrong OwnershipJudgment
Tools & Technologies
Cloud InfrastructureIncident Management ToolsAI-assisted Tooling
Industry Keywords
Site Reliability EngineeringDevOpsInfrastructure EngineeringProduction SystemsSecurity and Compliance

Tech Stack

Tools & technologies
CloudDistributed SystemsDNSJavaScriptKubernetesPythonTerraformTypeScriptGo

About the role

Key responsibilities & impact
  • Design, build, and operate Kubernetes platforms for secure, highly available, and scalable critical production services across hybrid and public-cloud environments
  • Build infrastructure as code with Terraform using reusable modules, safe delivery workflows, and governed infrastructure changes
  • Develop backend services, internal tools, and operational automation in Go, Python, or JavaScript/TypeScript
  • Partner with engineering and product teams to translate workload requirements into pragmatic solutions for reliability, performance, capacity, security, and cost
  • Improve the production lifecycle through CI/CD, deployment automation, progressive delivery, and clear operational ownership
  • Define and evolve observability practices across metrics, logs, traces, alerting, and dashboards
  • Participate in on-call, lead incident response, perform root-cause analysis, and drive blameless postmortems and corrective actions
  • Establish reliability targets using SLIs, SLOs, error budgets, capacity planning, disaster-recovery testing, and resilience improvements
  • Embed security and compliance into platform operations
  • Apply AI-assisted and data-driven operational techniques to improve signal detection, reduce alert noise, accelerate root-cause analysis, and identify automation opportunities
  • Contribute through code reviews, documentation, knowledge sharing, and mentorship
  • Mentor and support team growth

Requirements

What you’ll need
  • 5+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or a closely related software engineering role supporting production systems
  • Deep, hands-on Kubernetes expertise, including designing, operating, securing, and troubleshooting production clusters and containerized workloads at scale
  • Strong Terraform experience, including authoring reusable modules, managing state and environments, and delivering infrastructure changes through reviewable, automated workflows
  • Production software-development experience in Go, Python, or JavaScript/TypeScript
  • Demonstrated success improving the reliability, performance, scalability, or cost efficiency of distributed systems in production
  • Experience with cloud infrastructure and core networking concepts, including IAM, DNS, load balancing, routing, service networking, and secure connectivity
  • Strong observability and troubleshooting skills using metrics, logs, traces, alerting, and incident data
  • Experience defining and operating against SLIs, SLOs, error budgets, incident-management processes, postmortems, and disaster-recovery practices
  • Familiarity with CI/CD, GitOps or deployment automation, and canary or blue-green deployments
  • Security-minded approach to infrastructure and experience partnering with Security and engineering teams in regulated or high-availability environments
  • Clear written and verbal communication, strong ownership, and judgment to balance speed, risk, and operational excellence
  • Experience applying AI-assisted tooling to engineering or operations workflows is a plus

Benefits

Comp & perks
  • Flexible work environment
  • Remote-first work arrangement