Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Early Warning

Site Reliability Engineer II

Early Warning

. Improve the reliability, resilience, scalability, and operational health of production services .

Posted 10/3/2026full-timeUnited StatesJuniorMid-Level💰 $83,000 - $132,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in improving the reliability, scalability, and operational health of production services through software engineering, automation, and DevOps practices. Proficient in defining SLIs, SLOs, and implementing observability measures to enhance service performance and incident management.

Highest-signal resume keywords
Site Reliability EngineeringAWS Cloud TechnologiesCI/CD AutomationObservability and MonitoringIncident Management

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Software DevelopmentScriptingDistributed SystemsAutomationInfrastructure as CodePerformance AnalysisCapacity ManagementDisaster RecoveryResilience TestingCloud Architecture
Soft Skills
Analytical SkillsProblem-SolvingCommunicationCollaboration
Tools & Technologies
AWSMicrosoft AzureGoogle Cloud PlatformOracle Cloud InfrastructureLinux/UnixContainersOrchestrationMetricsLoggingDashboards
Industry Keywords
Service LifecycleOperational ReadinessIncident ResponseBlameless Post-Incident LearningReusable Automation

Tech Stack

Tools & technologies
AWSAzureCloudDistributed SystemsGoogle Cloud PlatformLinuxOracleUnix

About the role

Key responsibilities & impact
  • Improve the reliability, resilience, scalability, and operational health of production services
  • Partner with Software Engineering and technology teams to engineer reliability, observability, recoverability, performance, and operational readiness throughout the service lifecycle
  • Use software engineering, automation, and DevOps practices to improve how services are built, tested, deployed, observed, operated, and recovered
  • Use data, experimentation, and engineering analysis to identify reliability risks and guide technical decisions
  • Define and improve SLIs, SLOs, error budgets, and service-health measures
  • Improve observability through metrics, logging, tracing, monitoring, alerting, dashboards, and instrumentation
  • Drive continuous improvement across CI/CD, Infrastructure as Code, automation, testing, incident response, capacity management, resilience, and operational readiness
  • Identify systemic production issues and improve code, architecture, automation, tooling, and engineering practices
  • Participate in or lead incident response, troubleshooting, service restoration, and blameless post-incident learning
  • Independently participate in an on-call rotation, diagnose and resolve incidents and service degradation, and escalate complex issues
  • Reduce operational toil through software, automation, reusable patterns, and improved engineering practices
  • Create reusable solutions, share expertise, transfer knowledge, and improve team effectiveness

Requirements

What you’ll need
  • Typically 2–5 years of relevant professional experience in Software Engineering, Site Reliability Engineering, Systems Engineering, Cloud/Platform Engineering, DevOps, Infrastructure Engineering, Architecture, or a comparable technical discipline
  • Experience with software development or scripting using one or more modern programming languages
  • Experience with software engineering principles, distributed systems, production troubleshooting, automation, and observability
  • Experience with public cloud technologies and architectures, preferably AWS, along with infrastructure, networking, Linux/Unix, and modern application architectures
  • Demonstrated analytical, problem-solving, communication, and collaboration skills
  • Experience with AWS or comparable experience with Microsoft Azure, Google Cloud Platform (GCP), or Oracle Cloud Infrastructure (OCI) preferred
  • Experience developing, deploying, operating, or improving highly available production software or distributed systems
  • Experience with CI/CD, Infrastructure as Code, containers or orchestration, observability, monitoring, alerting, and software-delivery automation
  • Experience with SLIs, SLOs, error budgets, incident management, performance analysis, capacity management, resilience testing, disaster recovery, or operational readiness
  • Experience creating reusable automation, tooling, platforms, patterns, or practices
  • Bachelor's degree in Computer Science, Software Engineering, Computer Engineering, Information Systems, or a related technical field, or equivalent practical experience
  • Must independently possess eligibility to work in the United States for any employer at the date of hire
  • Position is ineligible for employment Visa sponsorship

Benefits

Comp & perks
  • Discretionary incentive plan
  • Medical, dental, and vision plans
  • Company contributions to HSA or pre-tax FSA savings accounts
  • 401(k) retirement plan with 100% Company Safe Harbor Match on first 6% deferral immediately upon eligibility
  • Flexible Time Off for exempt employees
  • Generous PTO for non-exempt employees
  • 11 paid company holidays
  • Paid volunteer day
  • 12 weeks of paid parental leave
  • Maven Family Planning support, including egg freezing, fertility, adoption, surrogacy, pregnancy, postpartum, early pediatrics, and returning to work