Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Fingerprint

Senior Site Reliability Engineer

Fingerprint

. Own the reliability of core production systems end to end .

Posted 9/18/2026full-timeRemote • United StatesSenior💰 $152,000 - $205,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in Site Reliability Engineering (SRE) with a focus on managing production systems, defining SLIs and SLOs, and implementing effective incident response strategies. Proficient in cloud infrastructure management, particularly with AWS, and skilled in programming and observability tooling to enhance system reliability and performance.

Highest-signal resume keywords
Site Reliability Engineering (SRE)Infrastructure as Code (Terraform)Incident Response LeadershipCloud Infrastructure Management (AWS)Programming Skills (Go, Python)

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
SLIs and SLOs DefinitionError Budget PracticesLoad TestingCapacity PlanningDistributed SystemsContainerization (EKS/Kubernetes)Observability Tooling (Datadog, Prometheus, Grafana)Redis/ElastiCache ManagementSoftware Engineering Best PracticesAI Tools for Incident Investigation
Soft Skills
Strong Written CommunicationStrong Verbal CommunicationPersonal OwnershipPragmatism
Industry Keywords
Cloud-Based EnvironmentsProduction EngineeringHigh-Throughput SystemsLow-Latency EnvironmentsOn-Call Rotation

Tech Stack

Tools & technologies
AWSCloudDistributed SystemsGrafanaKubernetesPrometheusPythonRedisTerraformGo

About the role

Key responsibilities & impact
  • Own the reliability of core production systems end to end
  • Define and maintain SLIs and SLOs, dashboards, alerts, and error-budget practices
  • Improve alert quality and anomaly/correctness detection
  • Lead incident response, restore service, and write actionable postmortems
  • Build secure, resilient, and cost-efficient infrastructure with explicit failure-mode handling
  • Perform load testing, profiling, saturation analysis, and capacity planning
  • Improve change safety through progressive delivery, automated rollback, pre-production signals, and safe deployment practices
  • Manage infrastructure through code and configuration, primarily using Terraform
  • Design, write, and ship software and developer-facing tooling
  • Run game days and chaos exercises
  • Partner with product engineering teams on production readiness, capacity, failure modes, rollback plans, runbooks, and on-call handoff
  • Participate in and improve the on-call rotation
  • Apply a security lens to engineering work and peer reviews
  • Serve as the go-to person for difficult production problems and mentor engineers through code review, pairing, and design feedback

Requirements

What you’ll need
  • 6–10 years of experience in SRE, production engineering, infrastructure, or backend engineering within primarily cloud-based environments (AWS preferred)
  • Track record of owning a system end to end
  • Hands-on experience defining and operating against SLIs, SLOs, and error budgets
  • Experience leading or serving as a primary responder on high-severity, customer-facing incidents
  • Depth in distributed-systems failure modes in high-throughput, low-latency environments
  • Depth in cloud infrastructure fundamentals, including networking, load balancing, containerization (EKS/Kubernetes), and distributed systems
  • Strong hands-on experience managing infrastructure through code and configuration (Terraform or equivalent)
  • Solid programming skills in Go, Python, or a comparable language
  • Fluency with observability tooling such as Datadog, Prometheus, Grafana, or OpenTelemetry
  • Hands-on experience operating Redis/ElastiCache in production, including cluster/shard management, failover behavior, memory eviction policies, and scaling strategies
  • Fluency with software engineering best practices, including source control, code review, comprehensive test coverage, and safe deployment
  • High level of personal ownership and autonomy, with experience working without clearly defined requirements
  • Pragmatism in balancing reliability and delivery
  • Strong written and verbal communication in English
  • AI-native use of AI tools for incident investigation, telemetry analysis, runbooks, and tooling
  • Must be authorized to work from the home location
  • Visa sponsorship is not provided

Benefits

Comp & perks
  • 100% remote work
  • Ability to join the workforce from almost any country, subject to country restrictions
  • Visa sponsorship is not provided
  • Inclusive work environment
  • CCPA and GDPR notices for applicable residents