FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Principal Site Reliability Engineer
Tandem Diabetes Care. Lead day-to-day production support, including intake, triage, prioritization, escalation, queue health, and change execution .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in Site Reliability Engineering and incident management, with a strong focus on automation, cloud services, and compliance. Proven ability to lead production support teams, optimize operational processes, and ensure system reliability through effective communication and collaboration.
Highest-signal resume keywords
Site Reliability Engineering PrinciplesIncident Management LeadershipTerraform Infrastructure as CodeCloud Platform Expertise (AWS/Azure/GCP)CI/CD Pipeline Development
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Incident CommandSLIs/SLOs DefinitionDisaster Recovery DesignAutomation and ToolingProduction SupportCloud Security PracticesScripting (Python/Go/Bash)Observability Tooling (Prometheus/Grafana)Kubernetes ManagementChange Execution
Soft Skills
Team LeadershipEffective CommunicationMentoring EngineersCollaboration Across TeamsProblem-Solving
Tools & Technologies
PagerDutyNew RelicGitHub ActionsOctopus DeployAzure DevOpsDockerCloudWatchELK/OpenSearchTerraformKubernetes
Certifications & Qualifications
AWS Professional CertificationAzure Professional CertificationGCP Professional Certification
Industry Keywords
FDA Regulated IndustriesISO ComplianceAgile MethodologiesPrivacy/HIPAA ComplianceCloud Cost Optimization
Tech Stack
Tools & technologiesAWSAzureCloudDockerGoogle Cloud PlatformGrafanaKubernetesPrometheusPythonTerraformGo
About the role
Key responsibilities & impact- Lead day-to-day production support, including intake, triage, prioritization, escalation, queue health, and change execution
- Establish consistent support practices across a distributed team, including shift handoffs, ticket quality standards, and ownership of open issues
- Lead incident management end-to-end, including incident command, stakeholder communication, and blameless postmortems with corrective actions tracked to closure
- Coordinate production security incident response with Security teams
- Own on-call strategy, rotation design, escalation paths, alert tuning, and tooling such as PagerDuty and New Relic
- Serve as a senior escalation tier for high-severity incidents
- Build runbooks for common failure modes and first-line resolution
- Define and own SLIs and SLOs for critical services
- Reduce MTTD and MTTR through instrumentation, alerting, diagnostics, and automation
- Convert recurring support burden into permanent fixes, automation, or documentation
- Maintain technology currency and lifecycle management for cloud services, Kubernetes clusters, operating systems, runtimes, and infrastructure components
- Own business continuity and disaster recovery readiness, including backup and recovery strategies, recovery testing, failover capabilities, recovery runbooks, and RTO/RPO objectives
- Lead infrastructure automation with Terraform
- Eliminate toil through automation
- Add reliability guardrails to CI/CD pipelines, including automated rollback, change-risk checks, and progressive delivery
- Maintain production systems in accordance with regulatory and compliance requirements
- Maintain business continuity and disaster recovery documentation and audit evidence
- Partner with Security, Quality, and Compliance teams on audits, compliance, and remediation
- Grow SRE and DevOps engineers through pairing, design and code review, and incident debriefs
- Foster open communication among junior and contract engineers
- Communicate documentation-first across distributed, multi-time-zone teams
- Partner with software engineering, QA, and architecture to embed reliability into the development lifecycle
- Inform capacity planning and scaling strategy with the Test team
- Introduce proactive resilience testing such as game days
- Align business continuity and disaster recovery capabilities with application requirements
- Support cloud cost optimization through rightsizing, reserved capacity, and observability spend governance
- Ensure work complies with company policies and applicable Privacy/HIPAA, regulatory, legal, and safety requirements
Requirements
What you’ll need- Demonstrated experience leading production support and incident management for production systems, including incident command during high-severity events
- Strong grounding in SRE principles, including SLIs/SLOs, blameless postmortems, toil reduction, and reliability engineering
- Demonstrated experience owning on-call strategy, including rotation design, alert tuning, and escalation
- Expertise with Terraform or comparable IaC at scale, including module design, state management, and policy-as-code guardrails
- Hands-on experience building CI/CD pipelines with reliability guardrails using GitHub Actions, Octopus Deploy, or Azure DevOps
- Deep experience with at least one major cloud platform: AWS, Azure, or GCP
- Experience with Docker and Kubernetes
- Working knowledge of observability tooling such as Prometheus, Grafana, Datadog, CloudWatch, and ELK/OpenSearch
- Experience designing and testing disaster recovery, including backup/restore, failover, and RTO/RPO validation
- Working knowledge of cloud security and compliance practices, including IAM, network segmentation, encryption, vulnerability and patch management, and cloud cost optimization
- Proficiency in at least one scripting or programming language such as Python, Go, or Bash
- Experience in FDA and ISO regulated industries and agile methodologies preferred
- B.S. in Computer Science or equivalent combination of education and applicable job experience, including technical school training and certifications in networks, servers, and cloud infrastructure; demonstrated production experience weighs more heavily than degree
- Relevant cloud certifications such as AWS/Azure/GCP Professional or Architect level preferred
- 10+ years in Site Reliability Engineering, DevOps, or infrastructure engineering
- 2+ years mentoring or technically leading other engineers, including remote, offshore, or contracted partner engineers
- Must be within the United States
- Successful completion of pre-employment drug test and background check
- Compliance with applicable company, Privacy/HIPAA, regulatory, legal, and safety requirements
Benefits
Comp & perks- Medical, dental, and vision benefits available the first day
- Health savings accounts
- Flexible savings accounts
- 11 paid holidays per year
- Minimum of 20 days of paid time off, with accrual starting on day 1
- 401(k) plan with company match
- Employee Stock Purchase plan
- Equipment provided
- Virtual training
- Bonus and competitive compensation package
- Joy-focused workplace supporting well-being, achievement, growth, fun, and camaraderie