Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
CVS Health

Staff Observability Platform Engineer – SRE

CVS Health

. Define, implement, and maintain key performance metrics, SLOs, and SLIs .

Posted 9/17/2026full-timeScottsdale • Arizona • United StatesLead💰 $118,450 - $236,900 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in defining and implementing key performance metrics, SLOs, and SLIs while managing cloud infrastructure and observability practices. Proficient in leading DevOps teams and integrating automated quality gates into CI/CD pipelines to enhance reliability and performance.

Highest-signal resume keywords
10+ Years Experience In Software Engineering7+ Years Experience With Observability Practices7+ Years Building Production-Grade Backend Services In Java/Python7+ Years With Cloud-Native And Containerized PlatformsExperience With Infrastructure As Code Tools

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Key Performance MetricsSLOsSLIsJavaPythonOpenTelemetryDockerKubernetesTerraformChaos Engineering
Soft Skills
Excellent Analytical SkillsAbility To Communicate Complex Technical Concepts
Tools & Technologies
PrometheusGrafanaLokiTempoAWSGCPAzureKafkaSplunkAppDynamics
Industry Keywords
Cloud InfrastructureObservability PracticesIncident ManagementQuality GatesData Pipelines

Tech Stack

Tools & technologies
AWSAzureCassandraCloudDockerGoogle Cloud PlatformGrafanaJavaKafkaKubernetesMySQLNoSQLPostgresPrometheusPulsarPythonSplunkTerraform

About the role

Key responsibilities & impact
  • Define, implement, and maintain key performance metrics, SLOs, and SLIs
  • Align reliability metrics with business objectives and operational goals
  • Manage error budgets and collaborate with development teams to balance reliability and feature delivery
  • Analyze incidents and outages to inform error-budget adjustments
  • Design and implement comprehensive monitoring solutions, dashboards, and alerts using Prometheus, Grafana, Loki, Tempo, and other observability platforms
  • Architect, design, and implement scalable cloud infrastructure for multiple business applications
  • Develop and implement automated quality gates for release reliability and performance standards
  • Lead the release DevOps team in integrating quality gates into CI/CD pipelines
  • Assist with incident response using metrics and monitoring insights
  • Conduct post-mortem analyses, identify root causes, and recommend preventive measures
  • Use AI to identify changes, anomalies, and next-best actions from metrics, logs, and traces
  • Apply GenAI to accelerate incident triage and root-cause analysis
  • Monitor AI workloads for quality, safety, cost, latency, and reliability with end-to-end tracing and secure logging/redaction
  • Embed AI signal checks into CI/CD to identify SLO risk, latency/error drift, and regression patterns before production release
  • Work closely with cross-functional teams to improve reliability, performance, and scalable growth of cloud-based systems

Requirements

What you’ll need
  • 10+ years of experience in Software Engineering, Platform Engineering, or SRE
  • 7+ years of experience with observability practices, including SLIs/SLOs/SLAs, alerting, and incident management
  • 7+ years building production-grade backend services in Java/python
  • 7+ years implementing and operating OpenTelemetry, including OTLP, semantic conventions, and instrumentation patterns
  • 7+ years with cloud-native and containerized platforms, including Docker, Kubernetes, and Argo CD
  • 7+ years working with public cloud platforms, including AWS, GCP, or Azure
  • 5+ years designing and scaling distributed, high-volume data pipelines
  • 5+ years working with Grafana OSS or comparable observability backends, such as Grafana, Loki, Tempo, and Prometheus
  • 5+ years with relational databases, including PostgreSQL and MySQL
  • Experience with service meshes and networking technologies such as Envoy and Istio
  • Experience integrating or operating commercial observability platforms such as Splunk and AppDynamics
  • Experience with streaming and data platforms such as Kafka and Pulsar
  • Familiarity with time-series, NoSQL, or analytical databases such as ClickHouse, Bigtable, and Cassandra
  • Experience with Infrastructure as Code tools such as Terraform or CloudFormation
  • Experience with cost optimization and capacity planning for large-scale cloud infrastructure
  • Experience with chaos engineering, resiliency testing, or fault injection
  • Background in security-aware platform design, including secure service-to-service communication
  • Experience mentoring senior engineers and influencing platform standards across organizations
  • Strong operational experience supporting 24x7 production systems, including on-call responsibilities
  • Knowledge of security best practices in cloud environments
  • Bachelor’s degree or equivalent experience (HS diploma + 4 years relevant experience)
  • Excellent analytical skills and ability to communicate complex technical concepts to non-technical stakeholders

Benefits

Comp & perks
  • CVS Health bonus, commission or short-term incentive program
  • Equity award program
  • Medical coverage
  • Dental coverage
  • Vision coverage
  • Paid time off
  • Retirement savings options
  • Wellness programs
  • Other resources supporting physical, emotional, and financial well-being