Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Veeam Software

Staff Site Reliability Engineer

Veeam Software

. Build reliability features as productized code, including libraries, services, controllers, deployment safety guards, rate-limiters, circuit breakers, load-shedding adapters, and back-pressure controls .

Posted 9/30/2026full-timeCalifornia • United StatesLead💰 $172,400 - $441,500 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in building reliability features and observability platforms for cloud-based products, with strong proficiency in backend programming languages and hands-on experience with Kubernetes and infrastructure as code tools. Capable of leading cross-team initiatives and driving measurable reliability outcomes through strategic architectural practices.

Highest-signal resume keywords
Cloud-Based Software EngineeringDistributed Systems DesignKubernetes ManagementInfrastructure as Code (Terraform/Pulumi)CI/CD Tools (GitHub Actions, GitLab, ArgoCD)

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
GoTypeScriptNode.jsC#JavaAPI DesignProduction-Grade ServicesObservability (Metrics, Tracing, Logging)Error BudgetsFault-Injection APIs
Soft Skills
MentoringCross-Team CollaborationLeadership
Tools & Technologies
KubernetesTerraformPulumiGitHub ActionsGitLabArgoCD
Industry Keywords
Reliability EngineeringObservability PlatformsProgressive DeliveryIncident Learning AutomationBlameless Learning

Tech Stack

Tools & technologies
AzureCloudDistributed SystemsJavaJavaScriptKubernetesNode.jsTerraformTypeScriptGo

About the role

Key responsibilities & impact
  • Build reliability features as productized code, including libraries, services, controllers, deployment safety guards, rate-limiters, circuit breakers, load-shedding adapters, and back-pressure controls
  • Define and implement observability platforms with telemetry pipelines, SLIs/SLOs, error-budget policies, SDKs, CLIs, plugins, and release gates
  • Build change-safety tooling using progressive delivery, automated rollback, release validation, reusable services/operators, and CI/CD integrations
  • Develop resilience automation including fault-injection APIs, chaos experiments, traffic shadowing, and load/performance harnesses
  • Create golden-path platform components using Terraform/Pulumi modules, Kubernetes operators, Helm charts, and reference microservice templates
  • Develop incident-learning automation for context capture, timelines, action tracking, and systemic code fixes
  • Write high-quality code in Go, TypeScript/Node.js, C#, or Java; design APIs, write tests, and ship iteratively
  • Lead designs for distributed, multi-region services initially on Azure
  • Partner with Staff/Principal peers across product and platform to align reliability standards and drive adoption
  • Instrument systems, automate detection and response, and maintain actionable alerting
  • Lead complex incidents, drive blameless learning, and implement systemic fixes
  • Mentor senior engineers through design reviews, ADRs, and pair programming
  • Drive strategic initiatives and define architectural best practices for Veeam’s global SRE function

Requirements

What you’ll need
  • 8+ years in software engineering for cloud-based products
  • Significant experience designing and operating distributed systems at scale
  • Strong proficiency in at least one backend language: C#, Java, Go, or TypeScript/Node.js
  • Experience writing production-grade services and libraries
  • Hands-on experience with Kubernetes
  • Hands-on experience with infrastructure as code using Terraform or Pulumi
  • Experience with CI/CD tools such as GitHub Actions, GitLab, or ArgoCD
  • Practical observability expertise in metrics, tracing, and logging
  • Experience turning SLOs and error budgets into engineering workflows
  • Ability to lead cross-team initiatives, influence architecture, and deliver measurable reliability outcomes
  • Comfortable with a follow-the-sun on-call model with 8-hour daytime rotations and coverage

Benefits

Comp & perks
  • Unlimited paid time off
  • 12 paid holidays, including 4 global VeeaMe Days for self-care
  • 24 paid volunteer hours annually through Veeam Cares
  • Paid parental leave: 8 weeks for all parents, 16 weeks for birthing parents
  • Medical, dental, and vision coverage starting on your first day
  • Mental health support, therapy sessions, and digital wellness tools via the Employee Assistance Program
  • 401(k) retirement plan with company matching contributions
  • Fertility, adoption, and surrogacy support through Maven
  • AirVet 24/7 virtual veterinary care at no cost
  • Legal services, identity protection, and supplemental health insurance options
  • Tax-advantaged spending accounts for healthcare, dependent care, and commuting
  • Learning and development through LinkedIn Learning, O’Reilly, mentoring, workshops, and learning events
  • Flexible work arrangements
  • Competitive pay and benefits
  • Compensatory benefits per local policy for on-call rotations