Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
OpenRouter

Site Reliability Engineer, Provider Operations

OpenRouter

. Build and own monitoring for every provider and endpoint, covering latency, throughput, error rates, uptime, and output correctness .

Posted 10/9/2026full-timeRemote • United StatesMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in Site Reliability Engineering (SRE) and production engineering, focusing on observability practices, incident management, and automation of provider operations. Proficient in TypeScript and Python, with a strong understanding of distributed systems and monitoring tools.

Highest-signal resume keywords
Site Reliability Engineering (SRE)Observability PracticesTypeScriptIncident ManagementDistributed Systems

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
MonitoringSLOsError BudgetsLoad TestingAutomationMetricsTracingLogsTraffic ManagementSynthetic Monitoring
Soft Skills
Calm Under PressureClear CommunicationIncident Command
Tools & Technologies
PythonTypeScriptCloudflare WorkersPostgresClickHouseGCPVercel
Industry Keywords
High-Traffic SystemsProvider OperationsAPI GatewayGPU CloudLLM Inference

Tech Stack

Tools & technologies
CloudDistributed SystemsGoogle Cloud PlatformPostgresPythonTypeScript

About the role

Key responsibilities & impact
  • Build and own monitoring for every provider and endpoint, covering latency, throughput, error rates, uptime, and output correctness
  • Set SLOs per provider tier and create actionable alerts
  • Improve detection of degraded endpoints and coordinate automatic traffic failover with the routing team
  • Own on-call for provider incidents, including triage, mitigation, provider communication, postmortems, and follow-up closure
  • Create provider scorecards and SLO reporting
  • Serve as the technical escalation point for misbehaving provider endpoints
  • Build continuous canaries and evaluations to detect silent quality regressions
  • Automate provider-operations tasks such as disabling endpoints, capacity changes, deprecations, and rate-limit tuning
  • Build endpoint load-testing tools for launch readiness and day-zero traffic
  • Report to the Provider Operations Manager

Requirements

What you’ll need
  • 4+ years in SRE, production engineering, or infrastructure roles running high-traffic, customer-facing systems
  • Strong observability practice with metrics, tracing, logs, SLOs/error budgets, and actionable alerting
  • Capable software engineer who prefers writing tools over executing runbooks
  • Experience with TypeScript and/or Python
  • Experience with distributed systems failure modes including timeouts, retries, backpressure, and partial outages
  • Ability to serve as a calm, clear incident commander and communicate with external partners under pressure
  • Understanding of, or eagerness to learn, LLM inference concepts including streaming, tool calling, prompt caching, throughput/latency tradeoffs, and provider API differences
  • Experience at an inference provider, model lab, GPU cloud, or API gateway/CDN company is nice to have
  • Experience with TypeScript, Cloudflare Workers, Postgres, ClickHouse, GCP, or Vercel is nice to have
  • Background in routing, load balancing, or traffic management systems is nice to have
  • Experience with evals or synthetic monitoring for ML systems is nice to have