FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Director, Site Reliability Engineering
Early Warning. Lead the SRE function for an assigned product, platform, pillar, or business domain .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in Site Reliability Engineering (SRE) practices, including SLOs, SLIs, and incident management, while leading high-performing engineering teams and driving reliability improvements in business-critical systems. Capable of translating business priorities into reliability investments and fostering a culture of accountability and continuous improvement.
Highest-signal resume keywords
Site Reliability Engineering (SRE)SLOs/SLIs ManagementIncident ManagementCloud PlatformsDistributed Systems
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Software DevelopmentAutomation EngineeringPerformance AnalysisCapacity PlanningReliability MetricsObservabilityError BudgetsProduction ReadinessTechnical LeadershipEngineering Solutions
Soft Skills
Team LeadershipStakeholder CommunicationInclusive Team BuildingAccountabilityTraining and Development
Tools & Technologies
Infrastructure AutomationMonitoring ToolsTelemetry SystemsDashboardsAlerting Systems
Industry Keywords
Financial ServicesPaymentsHigh-Availability EnvironmentRegulatory ComplianceBusiness-Critical Systems
Tech Stack
Tools & technologiesCloudDistributed Systems
About the role
Key responsibilities & impact- Lead the SRE function for an assigned product, platform, pillar, or business domain
- Own reliability, scalability, performance, and operability of business-critical production services
- Develop engineering talent and establish domain reliability strategy and priorities
- Partner with Product, Software Engineering, Platform, Infrastructure, Security, and Risk teams
- Establish and measure SLIs and SLOs aligned with customer, product, and business outcomes
- Identify and prioritize reliability risk using SLO performance, error budgets, failure data, and production evidence
- Champion reusable software, tooling, and automation to improve reliability, scalability, and operability
- Develop code, contribute to automation and engineering solutions, and work directly with engineers when appropriate
- Establish expectations for metrics, logs, traces, dashboards, and telemetry
- Drive actionable alerting and observability for diagnosis, performance analysis, capacity planning, and continuous improvement
- Provide leadership for incident response, service restoration, and stakeholder communication
- Improve incident response through automation, runbooks, training, exercises, and recurring-pattern analysis
- Promote blameless post-incident reviews and systemic learning
- Ensure critical services meet production-readiness, resilience, recovery, capacity, and scalability expectations
- Reduce repetitive, manual, and low-value operational work through engineering, automation, and simplification
- Partner with Security, Risk, Compliance, and engineering teams to meet control, resilience, and regulatory obligations
- Build, develop, and retain high-performing SRE teams with clear accountability, career development, and succession
- Translate business and product priorities into reliability investments and a multi-quarter roadmap
- Set priorities and make evidence-based tradeoffs across reliability, delivery, operational risk, capacity, and business needs
- Develop managers and technical leaders
- Make reliability risk, SLO performance, operational load, and improvement progress visible through metrics and operating reviews
- Provide accountable leadership during significant incidents and ensure corrective actions are completed
Requirements
What you’ll need- Typically 12+ years of relevant software engineering, site reliability engineering, production engineering, platform engineering, or closely related experience, including significant technical leadership
- 5+ years of people leadership experience, with demonstrated success leading engineering teams and developing managers and/or senior technical leaders
- Experience operating highly available, business-critical distributed systems and leading reliability improvement across a major product, platform, pillar, or business domain
- Strong understanding of SRE practices, including SLOs/SLIs, error budgets, observability, incident management, automation, capacity, resilience, and production readiness
- Demonstrated hands-on technical capability and sufficient depth to develop code, contribute to automation and engineering solutions, and work directly with engineers when appropriate
- Ability to communicate technical risk, tradeoffs, and investment needs clearly to engineering, product, and senior business stakeholders
- Demonstrated ability to build inclusive, accountable, and high-performing engineering teams
- Candidates must independently possess eligibility to work in the United States for any employer at the date of hire
- Position is ineligible for employment visa sponsorship
- Preferred: experience in payments, financial services, or another highly regulated, high-availability environment
- Preferred: experience leading SRE or production engineering across multiple teams or a complex product or platform ecosystem
- Preferred: experience with cloud platforms, distributed systems, modern observability, infrastructure automation, and software delivery at scale
- Preferred: experience establishing reliability metrics, governance, and operating reviews across teams within a defined domain
Benefits
Comp & perks- Discretionary incentive plan
- Competitive medical (PPO/HDHP), dental, and vision plans
- Company contributions to Health Savings Account (HSA)
- Flexible spending accounts (FSA) for commuting, health and dependent care expenses
- 401(k) retirement plan with 100% Company Safe Harbor Match on first 6% deferral immediately upon eligibility
- Flexible Time Off for Exempt (salaried) employees
- Generous PTO for Non-Exempt (hourly) employees
- 11 paid company holidays
- Paid volunteer day
- 12 weeks of Paid Parental Leave
- Maven Family Planning support, including egg freezing, fertility, adoption, surrogacy, pregnancy, postpartum, early pediatrics, and returning to work