FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Staff Site Reliability Engineer
Horizon3.ai. Own and evolve the engineering-wide SRE strategy, operating model, and reliability standards .
Tech Stack
Tools & technologiesAWSDistributed SystemsGrafanaKubernetesPythonTerraform
About the role
Key responsibilities & impact- Own and evolve the engineering-wide SRE strategy, operating model, and reliability standards
- Align reliability standards with customer impact, business priorities, and risk
- Lead cross-functional alignment across Infrastructure, product, service, security, and business stakeholders
- Improve reliability, observability, incident response, and operational readiness across multiple teams
- Establish organization-wide service ownership, meaningful SLIs and SLOs, and error budgets for critical customer paths and services
- Define and drive adoption of observability standards across pipelines and platform components
- Set standards for dashboards, actionable alerting, runbooks, and escalation paths
- Drive complex cross-functional reliability initiatives end to end
- Raise engineering standards for incident management, incident command, on-call health, post-incident learning, and recovery readiness
- Shape the technical direction, operating model, and growth path of the SRE function
- Participate in a 24/7 on-call rotation
- Help design a sustainable, appropriately staffed, and continuously improved on-call model
Requirements
What you’ll need- Experience designing, operating, and troubleshooting large-scale distributed systems in production environments
- Deep knowledge of reliability engineering, observability, incident management, and production operations
- Demonstrated ability to turn reliability knowledge into standards and practices adopted by others
- Experience establishing SLIs, SLOs, actionable alerts, observability, and service ownership
- Backend experience building systems and automation that reduce operational toil, strengthen safeguards, and improve operational efficiency
- Experience leading high-severity incidents and improving incident response programs
- Excellent written and verbal communication skills, including technical designs, runbooks, postmortems, and operational documentation
- Experience with Python and Terraform, or equivalent automation and infrastructure-as-code tools
- Experience with observability tools such as Datadog, New Relic, Grafana, or equivalent platforms
- Experience operating production services in AWS and Kubernetes
- Experience with CI/CD pipelines such as GitLab CI, ArgoCD, or GitOps workflows
- Participation in a 24/7 on-call rotation
- Legally authorized to work in the United States
- Must answer whether employment visa sponsorship is required
Benefits
Comp & perks- Equity package in the form of stock options for all full-time roles
- Health insurance for you and your family
- Vision insurance for you and your family
- Dental insurance for you and your family
- Flexible vacation policy
- Generous parental leave
- Career development opportunities
- Inclusive culture
- Collaborative environment encouraging creativity and out-of-the-box thinking
- Fully remote work model
- Team off-sites and in-person project kick-offs