FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in defining and managing service level indicators and error budgets, with a strong focus on observability using tools like Prometheus, Loki, and Grafana. Proven ability to lead incident response practices and performance engineering in production environments, particularly within B2B SaaS.
Highest-signal resume keywords
Service Level Objectives ManagementProduction Observability with PrometheusIncident Response LeadershipDatabase Performance OptimizationKubernetes Production Experience
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Service Level ObjectivesError BudgetsObservability ToolsDatabase Performance SkillsProduction KubernetesCoding in GoPythonC#TypeScriptLoad Testing with k6
Soft Skills
Excellent Written CommunicationJudgment Under Pressure
Tools & Technologies
PrometheusLokiTempoGrafanaAzureAKSMongoDBSentryPagerDutyTemporal
Industry Keywords
B2B SaaSIncident ManagementDisaster RecoveryMulti-Region ArchitectureDORA MetricsSOC 2NIST 800-171Construction Technology
Tech Stack
Tools & technologiesAWSAzureFluxGrafanaJMeterKubernetesMongoDBPrometheusPythonTypeScriptGo.NET
About the role
Key responsibilities & impact- Define and own service level indicators, objectives, and error budgets for customer-facing services
- Build measurement pipelines and publish availability and latency against targets
- Build and own production observability, including instrumentation standards, dashboards, and actionable alerting
- Operate observability across AKS on Azure, Prometheus, Loki, Tempo, Grafana, Istio, and Flux
- Establish and run incident response practices, including on-call rotation, paging paths, severity definitions, incident ownership, escalation, and blameless post-mortems
- Define rollback and recovery standards and verify them through regular exercises
- Lead performance and capacity engineering, including query and index tuning, connection-pool sizing, capacity modeling, and graceful degradation design
- Work with the team operating k6 load, stress, spike, and soak testing
- Set endpoint latency thresholds tied to SLOs and make test results a delivery gate
- Drive production readiness reviews covering instrumentation, alerting, failure modes, resource limits, and rollback
- Track post-mortem remediation items through verified production changes
- Contribute to business continuity and disaster recovery planning, including backup/restore validation, failover design, and recovery objectives
- Coach engineering teams on instrumenting and operating their own services
- Write tooling, automation, instrumentation libraries, and production-code fixes
- Work across engineering pods, platform and security functions, and customer-facing teams
- Report to the Senior Director, Platform Engineering
Requirements
What you’ll need- 6+ years of professional engineering experience
- At least 3+ years in a dedicated SRE or production engineering role at a B2B SaaS company
- Demonstrated ownership of SLOs and error budgets in production
- Deep, practical observability skills with Prometheus, Loki, Tempo, and Grafana or comparable tools
- Hands-on incident command experience at meaningful severity
- Experience building on-call and escalation practice from the ground up
- Strong database performance skills: query profiling, index design, connection pooling, and diagnosing saturation under load
- MongoDB experience strongly preferred; comparable document or relational depth acceptable
- Production Kubernetes experience; AKS preferred
- Ability to debug pod scheduling, resource limits, networking, and service mesh behavior; Istio preferred
- Solid coding ability in at least one of Go, Python, C#, or TypeScript
- Willingness to work in a C#/.NET codebase
- Experience with load and performance testing tooling such as k6, JMeter, or Gatling
- Production experience on Azure or AWS, with understanding of managed-service failure modes
- Fluency with AI-assisted engineering tooling and a track record of designing AI-leveraged workflows
- Excellent written communication for post-mortems, runbooks, and reliability reports
- Judgment and steadiness under pressure
- Experience establishing or maturing an SRE practice preferred
- Experience with Sentry or comparable application error-monitoring platforms preferred
- Experience operating event-driven and real-time systems preferred
- Experience operating MongoDB Atlas at production scale preferred
- Experience with durable workflow orchestration such as Temporal preferred
- Background in multi-region or multi-zone architecture and DR design preferred
- Experience with incident.io, PagerDuty, or comparable incident management platforms preferred
- Familiarity with DORA metrics and reliability work inside SOC 2 or NIST 800-171 scope preferred
- Experience with legacy monolith reliability preferred
- Domain interest in MEP, BIM, AEC, or construction technology preferred
- Prior experience in a Series B/growth-stage company preferred
Benefits
Comp & perks- Comprehensive and competitive health benefits plan
- Matching 401k contributions
- 20 days annual PTO
- Primarily remote work
- Occasional annual team onsites
