Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Tandem Diabetes Care

Principal Site Reliability Engineer

Tandem Diabetes Care

. Lead day-to-day production support, including intake, triage, prioritization, escalation, queue health, and change execution .

Posted 9/18/2026full-timeRemote • United StatesLead💰 $165,000 - $185,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in Site Reliability Engineering and incident management, with a strong focus on automation, cloud services, and compliance. Proven ability to lead production support teams, optimize operational processes, and ensure system reliability through effective communication and collaboration.

Highest-signal resume keywords
Site Reliability Engineering PrinciplesIncident Management LeadershipTerraform Infrastructure as CodeCloud Platform Expertise (AWS/Azure/GCP)CI/CD Pipeline Development

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Incident CommandSLIs/SLOs DefinitionDisaster Recovery DesignAutomation and ToolingProduction SupportCloud Security PracticesScripting (Python/Go/Bash)Observability Tooling (Prometheus/Grafana)Kubernetes ManagementChange Execution
Soft Skills
Team LeadershipEffective CommunicationMentoring EngineersCollaboration Across TeamsProblem-Solving
Tools & Technologies
PagerDutyNew RelicGitHub ActionsOctopus DeployAzure DevOpsDockerCloudWatchELK/OpenSearchTerraformKubernetes
Certifications & Qualifications
AWS Professional CertificationAzure Professional CertificationGCP Professional Certification
Industry Keywords
FDA Regulated IndustriesISO ComplianceAgile MethodologiesPrivacy/HIPAA ComplianceCloud Cost Optimization

Tech Stack

Tools & technologies
AWSAzureCloudDockerGoogle Cloud PlatformGrafanaKubernetesPrometheusPythonTerraformGo

About the role

Key responsibilities & impact
  • Lead day-to-day production support, including intake, triage, prioritization, escalation, queue health, and change execution
  • Establish consistent support practices across a distributed team, including shift handoffs, ticket quality standards, and ownership of open issues
  • Lead incident management end-to-end, including incident command, stakeholder communication, and blameless postmortems with corrective actions tracked to closure
  • Coordinate production security incident response with Security teams
  • Own on-call strategy, rotation design, escalation paths, alert tuning, and tooling such as PagerDuty and New Relic
  • Serve as a senior escalation tier for high-severity incidents
  • Build runbooks for common failure modes and first-line resolution
  • Define and own SLIs and SLOs for critical services
  • Reduce MTTD and MTTR through instrumentation, alerting, diagnostics, and automation
  • Convert recurring support burden into permanent fixes, automation, or documentation
  • Maintain technology currency and lifecycle management for cloud services, Kubernetes clusters, operating systems, runtimes, and infrastructure components
  • Own business continuity and disaster recovery readiness, including backup and recovery strategies, recovery testing, failover capabilities, recovery runbooks, and RTO/RPO objectives
  • Lead infrastructure automation with Terraform
  • Eliminate toil through automation
  • Add reliability guardrails to CI/CD pipelines, including automated rollback, change-risk checks, and progressive delivery
  • Maintain production systems in accordance with regulatory and compliance requirements
  • Maintain business continuity and disaster recovery documentation and audit evidence
  • Partner with Security, Quality, and Compliance teams on audits, compliance, and remediation
  • Grow SRE and DevOps engineers through pairing, design and code review, and incident debriefs
  • Foster open communication among junior and contract engineers
  • Communicate documentation-first across distributed, multi-time-zone teams
  • Partner with software engineering, QA, and architecture to embed reliability into the development lifecycle
  • Inform capacity planning and scaling strategy with the Test team
  • Introduce proactive resilience testing such as game days
  • Align business continuity and disaster recovery capabilities with application requirements
  • Support cloud cost optimization through rightsizing, reserved capacity, and observability spend governance
  • Ensure work complies with company policies and applicable Privacy/HIPAA, regulatory, legal, and safety requirements

Requirements

What you’ll need
  • Demonstrated experience leading production support and incident management for production systems, including incident command during high-severity events
  • Strong grounding in SRE principles, including SLIs/SLOs, blameless postmortems, toil reduction, and reliability engineering
  • Demonstrated experience owning on-call strategy, including rotation design, alert tuning, and escalation
  • Expertise with Terraform or comparable IaC at scale, including module design, state management, and policy-as-code guardrails
  • Hands-on experience building CI/CD pipelines with reliability guardrails using GitHub Actions, Octopus Deploy, or Azure DevOps
  • Deep experience with at least one major cloud platform: AWS, Azure, or GCP
  • Experience with Docker and Kubernetes
  • Working knowledge of observability tooling such as Prometheus, Grafana, Datadog, CloudWatch, and ELK/OpenSearch
  • Experience designing and testing disaster recovery, including backup/restore, failover, and RTO/RPO validation
  • Working knowledge of cloud security and compliance practices, including IAM, network segmentation, encryption, vulnerability and patch management, and cloud cost optimization
  • Proficiency in at least one scripting or programming language such as Python, Go, or Bash
  • Experience in FDA and ISO regulated industries and agile methodologies preferred
  • B.S. in Computer Science or equivalent combination of education and applicable job experience, including technical school training and certifications in networks, servers, and cloud infrastructure; demonstrated production experience weighs more heavily than degree
  • Relevant cloud certifications such as AWS/Azure/GCP Professional or Architect level preferred
  • 10+ years in Site Reliability Engineering, DevOps, or infrastructure engineering
  • 2+ years mentoring or technically leading other engineers, including remote, offshore, or contracted partner engineers
  • Must be within the United States
  • Successful completion of pre-employment drug test and background check
  • Compliance with applicable company, Privacy/HIPAA, regulatory, legal, and safety requirements

Benefits

Comp & perks
  • Medical, dental, and vision benefits available the first day
  • Health savings accounts
  • Flexible savings accounts
  • 11 paid holidays per year
  • Minimum of 20 days of paid time off, with accrual starting on day 1
  • 401(k) plan with company match
  • Employee Stock Purchase plan
  • Equipment provided
  • Virtual training
  • Bonus and competitive compensation package
  • Joy-focused workplace supporting well-being, achievement, growth, fun, and camaraderie