FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Principal Site Reliability Engineer
Early Warning. Apply software engineering and systems engineering practices to improve reliability, resilience, scalability, and operational health of production services .
Posted 9/17/2026full-timeSan Francisco • Arizona • United StatesLead💰 $207,000 - $276,000 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates extensive experience in Software Engineering, Site Reliability Engineering, and Cloud Technologies, with a strong focus on improving system reliability, observability, and operational readiness. Proficient in applying automation, CI/CD practices, and incident management to enhance engineering effectiveness and service health.
Highest-signal resume keywords
15+ Years Experience in Software EngineeringExpertise in AWS or Comparable Cloud TechnologiesProficient in CI/CD and Infrastructure as CodeExperience with SLIs, SLOs, and Incident ManagementStrong Analytical and Problem-Solving Skills
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Software DevelopmentScriptingDistributed SystemsAutomationObservabilityIncident ManagementPerformance AnalysisCapacity ManagementDisaster RecoveryCloud Architecture
Soft Skills
Analytical SkillsProblem-SolvingCommunicationCollaboration
Tools & Technologies
AWSMicrosoft AzureGoogle Cloud PlatformOracle Cloud InfrastructureCI/CD ToolsMonitoring ToolsLogging ToolsDashboardsContainersOrchestration
Industry Keywords
Site Reliability EngineeringDevOpsInfrastructure EngineeringTechnical LeadershipOperational Readiness
Tech Stack
Tools & technologiesAWSAzureCloudDistributed SystemsGoogle Cloud PlatformLinuxOracleUnix
About the role
Key responsibilities & impact- Apply software engineering and systems engineering practices to improve reliability, resilience, scalability, and operational health of production services
- Partner with Software Engineering and technology teams to engineer reliability, observability, recoverability, performance, and operational readiness throughout the system lifecycle
- Establish technical direction and apply evidence-driven engineering, technical rigor, automation, sound judgment, and broad systems expertise across organizational boundaries
- Improve how services are built, tested, deployed, observed, operated, and recovered using software engineering, automation, and DevOps practices
- Use data, evidence, experimentation, and rigorous engineering analysis to identify reliability risks, test assumptions, and guide technical decisions
- Define, implement, or improve SLIs, SLOs, error budgets, and service-health measures
- Improve observability through metrics, logging, tracing, monitoring, alerting, dashboards, and service-health instrumentation
- Drive continuous improvement across CI/CD, observability, deployment practices, Infrastructure as Code, automation, testing, incident response, capacity management, resilience, and operational readiness
- Identify recurring or systemic production issues and translate operational experience into improvements in code, architecture, automation, tooling, and engineering practices
- Participate in or lead incident response, troubleshooting, service restoration, and blameless post-incident learning
- Provide enterprise-level technical leadership for critical production incidents and improve incident response, escalation, service restoration, and sustainable on-call operations
- Reduce operational toil and manual intervention through software, automation, reusable patterns, and improved engineering practices
- Raise the effectiveness and technical capability of engineers and teams across the organization
- Establish enterprise technical direction, develop senior technical leaders, and operate autonomously across consequential reliability challenges
Requirements
What you’ll need- Typically 15+ years of relevant professional experience in Software Engineering, Site Reliability Engineering, Systems Engineering, Cloud/Platform Engineering, DevOps, Infrastructure Engineering, Architecture, or a comparable technical discipline
- Experience with software development or scripting using one or more modern programming languages
- Experience with software engineering principles, distributed systems, production troubleshooting, automation, and observability
- Experience with public cloud technologies and architectures, preferably AWS, along with infrastructure, networking, Linux/Unix, and modern application architectures
- Demonstrated analytical, problem-solving, communication, and collaboration skills
- Hands-on experience with AWS is preferred, or comparable experience with Microsoft Azure, Google Cloud Platform (GCP), or Oracle Cloud Infrastructure (OCI)
- Experience developing, deploying, operating, or improving highly available production software or distributed systems
- Experience with CI/CD, Infrastructure as Code, containers or orchestration, observability, monitoring, alerting, and software-delivery automation
- Experience with SLIs, SLOs, error budgets, incident management, performance analysis, capacity management, resilience testing, disaster recovery, or operational readiness
- Experience creating reusable automation, tooling, platforms, patterns, or practices that improve engineering effectiveness
- Bachelor's degree in Computer Science, Software Engineering, Computer Engineering, Information Systems, or a related technical field, or equivalent practical experience
- Independently possess eligibility to work in the United States at the date of hire
- Position is ineligible for employment Visa sponsorship
- Ability to perform essential functions and physical requirements with or without reasonable accommodation
- Ability to lift 10 pounds occasionally and/or negligible amount of force frequently
- Requires visual acuity and dexterity to use computers and office equipment
Benefits
Comp & perks- Competitive medical (PPO/HDHP), dental, and vision plans
- Company contributions to Health Savings Account (HSA)
- Flexible spending accounts (FSA) for commuting, health, and dependent care expenses
- 401(k) Retirement Plan with a 100% Company Safe Harbor Match on the first 6% deferral immediately upon eligibility
- Flexible Time Off for Exempt (salaried) employees
- Generous PTO for Non-Exempt (hourly) employees
- 11 paid company holidays
- Paid volunteer day
- 12 weeks of Paid Parental Leave
- Maven Family Planning support, including egg freezing, fertility, adoption, surrogacy, pregnancy, postpartum, early pediatrics, and returning to work
- Discretionary incentive plan
- Reasonable accommodation for essential functions and physical requirements