FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Engineering Manager – Site Reliability Engineering
Replit. Lead SRE across observability, incident management, load testing, performance engineering, cloud cost and capacity, and rollout infrastructure .
Posted 9/26/2026full-timeFoster City • California • United StatesSeniorLead💰 $250,000 - $325,000 per yearWebsite
Tech Stack
Tools & technologiesCloudDistributed SystemsGoogle Cloud PlatformKubernetes
About the role
Key responsibilities & impact- Lead SRE across observability, incident management, load testing, performance engineering, cloud cost and capacity, and rollout infrastructure
- Lead and grow an existing team building and operating production platforms across application and infrastructure boundaries
- Build and operate metrics, logs, traces, and alerting capabilities
- Help teams establish meaningful SLOs and use production telemetry to diagnose problems and verify improvements
- Own incident tooling and practices, coordinate cross-team response, and turn incident reviews into engineering improvements
- Build and maintain load and failure-testing capabilities
- Validate critical paths under expected demand, quantify headroom, and test recovery and production readiness with service owners
- Lead deep engagements on SLOs and end-to-end performance
- Use profiling, telemetry, and load tests to identify bottlenecks and deliver improvements with service owners
- Review designs and production changes, debug difficult failure modes, and use AI coding tools to prototype and automate
- Coach engineers, develop technical leaders, manage performance, and hire against agreed needs
- Establish deliberate distributed collaboration, mentoring, and backup coverage
- Track rollout safety, recovery time, repeat incidents, critical-path latency/throughput, test coverage, and cost/capacity improvements
- Agree success measures and continuing ownership with partner teams
Requirements
What you’ll need- Demonstrated engineering management: led and developed engineers, made prioritization and performance decisions, hired thoughtfully, and delivered through a team
- Experience building and operating distributed systems or reliability platforms
- Ability to reason across deployment behavior, Kubernetes, telemetry, service dependencies, and recovery mechanisms
- Experience leading consequential migrations or incidents
- Experience using measurement to diagnose reliability or performance problems and validate fixes under realistic conditions
- Ability to build capabilities adopted by other teams
- Ability to lead hands-on engagements without absorbing every service's operations
- Ability to make tradeoffs among reliability, performance, engineering effort, and cost
- Experience with GitOps or progressive-delivery platforms such as Harness, ArgoCD, or Kargo (nice to have)
- Experience with observability, profiling, load-testing, and failure-testing systems, including OpenTelemetry or comparable tooling (nice to have)
- Experience with cloud cost attribution, capacity planning, and provider coordination, particularly on GCP (nice to have)
- Experience growing distributed teams and using AI tools while preserving production safeguards (nice to have)
- Ability to work from the Foster City, CA headquarters 3 days per week, or willingness to relocate
- Legal authorization to work in the United States
- Must be at least 18 years of age
Benefits
Comp & perks- Competitive Salary & Equity
- 401(k) Program with a 4% match (US Only)
- Health, Dental, Vision and Life Insurance
- Short Term and Long Term Disability
- Paid Parental, Medical, Caregiver Leave
- Flexible Time Off (FTO) + Holidays
- Commuter Benefits (In-Office & US Only)
- Monthly Wellness Stipend
- Autonomous Work Environment
- In Office Set-Up Reimbursement (In-Office Only)
- Quarterly Team Gatherings
- In Office Amenities (In-Office Only)