FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and implementing observability instrumentation, building telemetry pipelines, and operationalizing SLIs and SLOs. Proficient in backend software engineering with a strong focus on distributed systems and performance optimization.
Highest-signal resume keywords
Backend Software EngineeringProficiency in Go, C++, Rust, Java, PythonDistributed Systems UnderstandingExperience with OpenTelemetry, Prometheus, GrafanaDesigning Scalable Telemetry Pipelines
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Observability InstrumentationTelemetry Pipeline DevelopmentService-Level IndicatorsService-Level ObjectivesMetrics, Logs, and Distributed TracingProduction Monitoring and AlertingConcurrency and Performance TradeoffsRoot-Cause AnalysisHigh-Signal Alert DesignDebugging Large-Scale Production Incidents
Soft Skills
Collaboration with EngineersClear CommunicationProblem-Solving
Tools & Technologies
OpenTelemetryPrometheusGrafanaDatadogElasticJaegerTempo
Industry Keywords
High-Performance ComputingAI/ML SystemsHardware-Aware ObservabilitySREPlatform Engineering
Tech Stack
Tools & technologiesDistributed SystemsGrafanaJavaPrometheusPythonRustC++Go
About the role
Key responsibilities & impact- Design and implement observability instrumentation across services and platforms
- Build and maintain telemetry pipelines for metrics, logs, and traces at scale
- Develop internal observability platforms, libraries, and tooling
- Define and operationalize SLIs, SLOs, and alerting strategies
- Partner with engineers to make systems debuggable by design
- Reduce MTTR by enabling fast root-cause analysis during incidents
- Create clear, actionable dashboards and alerts that reflect real system health
- Balance telemetry signal against cost, noise, and performance impact
- Improve the developer experience around observability and debugging
- Write production software, shape internal platforms, and work with engineers across the stack
Requirements
What you’ll need- Strong experience in backend or systems software engineering
- Proficiency in one or more of: Go, C++, Rust, Java, Python
- Solid understanding of distributed systems
- Solid understanding of networking fundamentals
- Solid understanding of concurrency and performance tradeoffs
- Hands-on experience with metrics, logs, and distributed tracing
- Experience with production monitoring and alerting
- Familiarity with OpenTelemetry, Prometheus, Grafana, Datadog, Elastic, Jaeger, Tempo, or similar tools
- Experience designing high-signal alerts
- Experience designing scalable telemetry pipelines
- Experience designing service-level indicators and objectives
- Preferred: experience in high-performance computing, AI/ML systems, or inference platforms
- Preferred: hardware-aware observability, including accelerators, GPUs, or custom hardware
- Preferred: prior SRE or platform engineering background
- Preferred: experience debugging large-scale production incidents
- Preferred: building internal developer platforms or shared libraries
Benefits
Comp & perks- Job stability with startup vitality
- Opportunity to publish and open source cutting-edge AI research
- Work on one of the fastest AI supercomputers in the world
- Simple, non-corporate work culture that respects individual beliefs
- Continuous learning, growth and support
- Equal and diverse work environment
