FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates extensive experience in Site Reliability Engineering (SRE), DevOps, and Observability, with a strong focus on designing and operating production observability platforms using AWS services and OpenTelemetry. Capable of leading technical teams, conducting assessments, and delivering strategic recommendations to enhance observability and performance.
Highest-signal resume keywords
Production Observability PlatformsOpenTelemetry Architecture DesignPrometheus and Thanos ExpertiseAWS Services ProficiencyTerraform Advanced Proficiency
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Metrics ManagementLogging and Tracing PlatformsPromQLSLO-Based AlertingData GovernancePython ScriptingGo ScriptingBash ScriptingLog Pipeline DesignCost Model Development
Soft Skills
Technical LeadershipStakeholder CommunicationAnalytical Problem SolvingCollaboration
Tools & Technologies
DatadogCloudWatchKubernetesGrafanaVectorFluent BitLogstashAmazon Managed Service for PrometheusAmazon Managed GrafanaPagerDuty
Certifications & Qualifications
Prometheus Certified AssociateOpenTelemetry Certified AssociateAWS Certified DevOps Engineer – ProfessionalAWS Certified Solutions Architect – ProfessionalCKA
Industry Keywords
Observability Maturity AssessmentTelemetry ClassificationPerformance-Parity TestsVendor MigrationCost Attribution
Tech Stack
Tools & technologiesAWSGrafanaKubernetesLogstashPrometheusPythonRayTerraformGo
About the role
Key responsibilities & impact- Lead an embedded TechPod as Pod Leader and serve as the primary contact for customer engineering leadership
- Lead a two-month Observability Maturity Assessment
- Assess the observability estate across vendors, agents, collectors, query surfaces, data volumes, and operating model
- Validate metric cardinality, active series, scrape target health, log volumes, trace sampling, retention, and telemetry fidelity
- Build total cost of ownership comparisons for current and target states
- Design an AWS-native observability target architecture using OpenTelemetry and AWS services
- Partner with Security, Legal, and Engineering on telemetry classification and policy-driven retention
- Map current capabilities to AWS-native equivalents and identify gaps
- Define and run performance-parity tests for query performance, alert latency, and data fidelity
- Plan and lead migration execution, including pipeline cutover, porting log transforms to OpenTelemetry, rebuilding dashboards, and deduplicating/rebuilding alerts
- Define OpenTelemetry conventions, collector topology, and sampling strategy
- Rebuild service ownership tagging for telemetry cost attribution
- Analyze platform spend, licensing, and commit structures and inform vendor renewal strategy
- Lead working cadence with customer observability leadership
- Produce assessment reports, target architecture, cost model, migration plan, executive summary, and present findings to engineering leadership
Requirements
What you’ll need- 8+ years in SRE, DevOps, Observability, or Platform Engineering
- 4+ years owning production observability platforms
- Prior experience in a technical lead, staff, or principal-level role
- Deep experience running metrics, logging, and tracing platforms at large scale
- Advanced production experience with Prometheus and Thanos, Cortex, or Mimir
- Experience with PromQL, recording rules, remote write, and cardinality management
- Hands-on experience with Datadog or a comparable commercial platform
- Production experience with CloudWatch, CloudWatch Logs, Container Insights, X-Ray or Application Signals, Amazon Managed Service for Prometheus, and Amazon Managed Grafana
- Production experience with OpenTelemetry Collector or ADOT, including pipeline design, processors, tail and head sampling, and multi-backend export
- Experience designing and operating log pipelines with Vector, Fluent Bit, Logstash, or Firehose
- Strong production experience with EKS or Kubernetes
- Advanced proficiency with Terraform
- Experience designing SLO-based alerting and integrating with incident management tooling such as PagerDuty
- Ability to build defensible TCO models from usage data and pricing
- Strong scripting ability using Python, Go, or Bash
- Demonstrated ability to assess unfamiliar environments and produce recommendations quickly
- Ability to explain technical, cost, and compliance tradeoffs to technical and executive stakeholders
- Experience with vendor migrations, dashboards and alerts as code, Grafana ecosystem tools, continuous profiling, data governance, analytics platforms, consumer-scale platforms, or consulting is a plus
- Certifications such as Prometheus Certified Associate, OpenTelemetry Certified Associate, AWS Certified DevOps Engineer – Professional, AWS Certified Solutions Architect – Professional, or CKA are extra awesome
- EverOps hires remotely in the United States
- Must be legally authorized to work for any employer in the U.S.
Benefits
Comp & perks- 100% Remote Workplace
- Unlimited Paid Time Off
- Equity: Become a true owner of the company
- 401K with company contribution
- Sponsored healthcare
- Access to training and certification programs
