See all jobs on Scoutfield
Search thousands of fresh jobs every day.
- Fresh listings
- Fast filters
- No subscription required

Staff Site Reliability Engineer
Veeam Software. Build reliability features as productized code, including libraries, services, controllers, deployment safety guards, rate-limiters, circuit breakers, load-shedding adapters, and back-pressure controls .
Core Competencies
Role fitUse this summary to align your resume positioning with the role.
Demonstrates expertise in building reliability features and observability platforms for cloud-based products, with strong proficiency in backend programming languages and hands-on experience with Kubernetes and infrastructure as code tools. Capable of leading cross-team initiatives and driving measurable reliability outcomes through strategic architectural practices.
ATS Keywords
Tailor your resumeTip: use these terms in your resume and cover letter to boost ATS matches.
Tech Stack
Tools & technologiesAbout the role
Key responsibilities & impact- Build reliability features as productized code, including libraries, services, controllers, deployment safety guards, rate-limiters, circuit breakers, load-shedding adapters, and back-pressure controls
- Define and implement observability platforms with telemetry pipelines, SLIs/SLOs, error-budget policies, SDKs, CLIs, plugins, and release gates
- Build change-safety tooling using progressive delivery, automated rollback, release validation, reusable services/operators, and CI/CD integrations
- Develop resilience automation including fault-injection APIs, chaos experiments, traffic shadowing, and load/performance harnesses
- Create golden-path platform components using Terraform/Pulumi modules, Kubernetes operators, Helm charts, and reference microservice templates
- Develop incident-learning automation for context capture, timelines, action tracking, and systemic code fixes
- Write high-quality code in Go, TypeScript/Node.js, C#, or Java; design APIs, write tests, and ship iteratively
- Lead designs for distributed, multi-region services initially on Azure
- Partner with Staff/Principal peers across product and platform to align reliability standards and drive adoption
- Instrument systems, automate detection and response, and maintain actionable alerting
- Lead complex incidents, drive blameless learning, and implement systemic fixes
- Mentor senior engineers through design reviews, ADRs, and pair programming
- Drive strategic initiatives and define architectural best practices for Veeam’s global SRE function
Requirements
What you’ll need- 8+ years in software engineering for cloud-based products
- Significant experience designing and operating distributed systems at scale
- Strong proficiency in at least one backend language: C#, Java, Go, or TypeScript/Node.js
- Experience writing production-grade services and libraries
- Hands-on experience with Kubernetes
- Hands-on experience with infrastructure as code using Terraform or Pulumi
- Experience with CI/CD tools such as GitHub Actions, GitLab, or ArgoCD
- Practical observability expertise in metrics, tracing, and logging
- Experience turning SLOs and error budgets into engineering workflows
- Ability to lead cross-team initiatives, influence architecture, and deliver measurable reliability outcomes
- Comfortable with a follow-the-sun on-call model with 8-hour daytime rotations and coverage
Benefits
Comp & perks- Unlimited paid time off
- 12 paid holidays, including 4 global VeeaMe Days for self-care
- 24 paid volunteer hours annually through Veeam Cares
- Paid parental leave: 8 weeks for all parents, 16 weeks for birthing parents
- Medical, dental, and vision coverage starting on your first day
- Mental health support, therapy sessions, and digital wellness tools via the Employee Assistance Program
- 401(k) retirement plan with company matching contributions
- Fertility, adoption, and surrogacy support through Maven
- AirVet 24/7 virtual veterinary care at no cost
- Legal services, identity protection, and supplemental health insurance options
- Tax-advantaged spending accounts for healthcare, dependent care, and commuting
- Learning and development through LinkedIn Learning, O’Reilly, mentoring, workshops, and learning events
- Flexible work arrangements
- Competitive pay and benefits
- Compensatory benefits per local policy for on-call rotations