FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Tech Stack
Tools & technologiesCloudDistributed SystemsKubernetes
About the role
Key responsibilities & impact- Design, build, and maintain automation and tooling to reduce operational toil
- Mature reliability practices with Product & Engineering leaders
- Participate in the incident-management lifecycle, including detection, engagement, escalation, mitigation, stakeholder communication, post-incident review, and corrective-action follow-through
- Use alert, incident, support, and SLO trends to shift from reactive response toward proactive risk reduction
- Build feedback loops connecting incident learning to engineering standards, service maturity, product priorities, and vendor actions
- Collaborate with application engineering teams to improve developer experience and reduce toil
- Partner with service owners on production readiness standards
- Provide SRE guidance on capacity planning, resilience testing, game days, disaster-recovery readiness, and modernization of fragile or legacy workloads
- Document critical customer workflows, define health expectations, identify dependencies, and align reliability investment with business priorities
- Participate in cross-team reliability engagements and influence outcomes without direct authority
- Build strategic vendor relationships in observability, incident response, and cloud infrastructure
- Adopt AI-assisted and agentic workflows for alert triage, incident mitigation, postmortems, trend analysis, capacity planning, SLO analysis, and self-service knowledge
- Improve service metadata, observability data, incident records, runbooks, architecture documentation, and corrective-action quality
Requirements
What you’ll need- Experience in a production-facing SRE role supporting a complex SaaS environment
- Strong technical judgment across distributed systems, multi-cloud environments, Kubernetes, networking, infrastructure technologies, GitOps, and CI/CD
- Experience establishing and maturing SRE principles and practices, including SLOs, error budgets, observability, capacity planning, incident response, and toil elimination
- Proficiency using observability data to resolve high-severity incidents and investigate root causes during postmortems
- Experience leading high-pressure incidents and communicating clearly with technical teams, executives, customer-facing stakeholders, and third-party vendors
- Ability to participate in a 12-hour follow-the-sun on-call rotation
- Ability to responsibly adopt AI-assisted and agentic workflows while keeping qualified humans in the decision loop for production-impacting actions
Benefits
Comp & perks- Participation in Seismic's incentive plans in addition to base salary
- Remote work arrangement
- 12-hour follow-the-sun on-call rotation
