Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Horizon3.ai

Staff Site Reliability Engineer

Horizon3.ai

. Own and evolve the engineering-wide SRE strategy, operating model, and reliability standards .

Posted 9/25/2026full-timeRemote • United StatesLead💰 $199,750 - $270,000 per yearWebsite

Tech Stack

Tools & technologies
AWSDistributed SystemsGrafanaKubernetesPythonTerraform

About the role

Key responsibilities & impact
  • Own and evolve the engineering-wide SRE strategy, operating model, and reliability standards
  • Align reliability standards with customer impact, business priorities, and risk
  • Lead cross-functional alignment across Infrastructure, product, service, security, and business stakeholders
  • Improve reliability, observability, incident response, and operational readiness across multiple teams
  • Establish organization-wide service ownership, meaningful SLIs and SLOs, and error budgets for critical customer paths and services
  • Define and drive adoption of observability standards across pipelines and platform components
  • Set standards for dashboards, actionable alerting, runbooks, and escalation paths
  • Drive complex cross-functional reliability initiatives end to end
  • Raise engineering standards for incident management, incident command, on-call health, post-incident learning, and recovery readiness
  • Shape the technical direction, operating model, and growth path of the SRE function
  • Participate in a 24/7 on-call rotation
  • Help design a sustainable, appropriately staffed, and continuously improved on-call model

Requirements

What you’ll need
  • Experience designing, operating, and troubleshooting large-scale distributed systems in production environments
  • Deep knowledge of reliability engineering, observability, incident management, and production operations
  • Demonstrated ability to turn reliability knowledge into standards and practices adopted by others
  • Experience establishing SLIs, SLOs, actionable alerts, observability, and service ownership
  • Backend experience building systems and automation that reduce operational toil, strengthen safeguards, and improve operational efficiency
  • Experience leading high-severity incidents and improving incident response programs
  • Excellent written and verbal communication skills, including technical designs, runbooks, postmortems, and operational documentation
  • Experience with Python and Terraform, or equivalent automation and infrastructure-as-code tools
  • Experience with observability tools such as Datadog, New Relic, Grafana, or equivalent platforms
  • Experience operating production services in AWS and Kubernetes
  • Experience with CI/CD pipelines such as GitLab CI, ArgoCD, or GitOps workflows
  • Participation in a 24/7 on-call rotation
  • Legally authorized to work in the United States
  • Must answer whether employment visa sponsorship is required

Benefits

Comp & perks
  • Equity package in the form of stock options for all full-time roles
  • Health insurance for you and your family
  • Vision insurance for you and your family
  • Dental insurance for you and your family
  • Flexible vacation policy
  • Generous parental leave
  • Career development opportunities
  • Inclusive culture
  • Collaborative environment encouraging creativity and out-of-the-box thinking
  • Fully remote work model
  • Team off-sites and in-person project kick-offs