FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Site Reliability Engineer – Infra Ops
Circle. Design, build, and operate Kubernetes platforms for secure, highly available, and scalable critical production services across hybrid and public-cloud environments .
Posted 10/1/2026full-timeRemote • California • United StatesSenior💰 $152,500 - $205,000 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates extensive expertise in Kubernetes platform design and operation, with strong capabilities in Terraform for infrastructure as code. Proficient in backend development using Go, Python, or JavaScript/TypeScript, and skilled in implementing observability practices and incident management processes.
Highest-signal resume keywords
Kubernetes ExpertiseTerraform ProficiencyBackend Development in Go, Python, or JavaScript/TypeScriptObservability and Troubleshooting SkillsCI/CD and Deployment Automation
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
KubernetesTerraformGoPythonJavaScriptTypeScriptCI/CDGitOpsSLIsSLOs
Soft Skills
Clear CommunicationStrong OwnershipJudgment
Tools & Technologies
Cloud InfrastructureIncident Management ToolsAI-assisted Tooling
Industry Keywords
Site Reliability EngineeringDevOpsInfrastructure EngineeringProduction SystemsSecurity and Compliance
Tech Stack
Tools & technologiesCloudDistributed SystemsDNSJavaScriptKubernetesPythonTerraformTypeScriptGo
About the role
Key responsibilities & impact- Design, build, and operate Kubernetes platforms for secure, highly available, and scalable critical production services across hybrid and public-cloud environments
- Build infrastructure as code with Terraform using reusable modules, safe delivery workflows, and governed infrastructure changes
- Develop backend services, internal tools, and operational automation in Go, Python, or JavaScript/TypeScript
- Partner with engineering and product teams to translate workload requirements into pragmatic solutions for reliability, performance, capacity, security, and cost
- Improve the production lifecycle through CI/CD, deployment automation, progressive delivery, and clear operational ownership
- Define and evolve observability practices across metrics, logs, traces, alerting, and dashboards
- Participate in on-call, lead incident response, perform root-cause analysis, and drive blameless postmortems and corrective actions
- Establish reliability targets using SLIs, SLOs, error budgets, capacity planning, disaster-recovery testing, and resilience improvements
- Embed security and compliance into platform operations
- Apply AI-assisted and data-driven operational techniques to improve signal detection, reduce alert noise, accelerate root-cause analysis, and identify automation opportunities
- Contribute through code reviews, documentation, knowledge sharing, and mentorship
- Mentor and support team growth
Requirements
What you’ll need- 5+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or a closely related software engineering role supporting production systems
- Deep, hands-on Kubernetes expertise, including designing, operating, securing, and troubleshooting production clusters and containerized workloads at scale
- Strong Terraform experience, including authoring reusable modules, managing state and environments, and delivering infrastructure changes through reviewable, automated workflows
- Production software-development experience in Go, Python, or JavaScript/TypeScript
- Demonstrated success improving the reliability, performance, scalability, or cost efficiency of distributed systems in production
- Experience with cloud infrastructure and core networking concepts, including IAM, DNS, load balancing, routing, service networking, and secure connectivity
- Strong observability and troubleshooting skills using metrics, logs, traces, alerting, and incident data
- Experience defining and operating against SLIs, SLOs, error budgets, incident-management processes, postmortems, and disaster-recovery practices
- Familiarity with CI/CD, GitOps or deployment automation, and canary or blue-green deployments
- Security-minded approach to infrastructure and experience partnering with Security and engineering teams in regulated or high-availability environments
- Clear written and verbal communication, strong ownership, and judgment to balance speed, risk, and operational excellence
- Experience applying AI-assisted tooling to engineering or operations workflows is a plus
Benefits
Comp & perks- Flexible work environment
- Remote-first work arrangement