Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Salad

Staff Software Engineer, AI Inference Gateway

Salad

. Own the AI Gateway system end to end, including request routing, streaming, batch/async job handling, and horizontal scaling across multiple gateway servers .

Posted 9/19/2026full-timeRemote • United StatesLead💰 $180,000 - $220,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in managing AI Gateway systems, including request routing and performance optimization, while effectively collaborating with cross-functional teams. Proficient in operating LLM inference servers and implementing observability solutions to ensure system reliability and efficiency.

Highest-signal resume keywords
Rust ProgrammingLLM Inference ServersDistributed Systems FundamentalsOpenTelemetry ObservabilityGPU Hardware Benchmarking

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
RustLLM Inference ServersDistributed SystemsPerformance ScoringNode AutoscalingKV Cache BehaviorMetrics and TracingQuantization FormatsLoad BalancingFailure Handling
Soft Skills
Clear CommunicationCollaboration
Tools & Technologies
VLLMLlama.cppTensorRT-LLMOpenTelemetryCUDAROCm
Industry Keywords
AI GatewayFleet Efficiency AlgorithmsSubscription Billing IntegrationsCross-System DesignOn-Call Responsibilities

Tech Stack

Tools & technologies
Distributed SystemsNode.jsRust

About the role

Key responsibilities & impact
  • Own the AI Gateway system end to end, including request routing, streaming, batch/async job handling, and horizontal scaling across multiple gateway servers
  • Build and tune fleet-efficiency algorithms for node autoscaling, performance scoring, and eviction of underperformers
  • Operate and scale inference nodes running vLLM, llama.cpp, and similar servers
  • Design and maintain OpenTelemetry-based observability across the gateway and fleet
  • Maintain accurate pay-per-token and subscription billing integrations based on telemetry
  • Benchmark new models, quantizations, and GPU hardware for approximately 25% of working time
  • Work with product and marketing teams to decide what enters production
  • Collaborate on cross-system design and integrations with Salad engineering
  • Explain technical trade-offs to technical and non-technical teammates
  • Report directly to the CTO
  • Participate in shared after-hours support

Requirements

What you’ll need
  • Strong production Rust experience, ideally on high-throughput networked services
  • Hands-on experience running LLM inference servers (vLLM, llama.cpp, TGI, TensorRT-LLM, or similar)
  • Knowledge of KV cache behavior
  • Distributed systems fundamentals: load balancing, backpressure, failure handling, and unreliable nodes
  • Experience operating systems with metrics, tracing, and on-call responsibilities
  • Willingness to be on-call
  • Clear written and verbal communication
  • Nice to have: Pingora, Tokio, or proxy/gateway internals
  • Nice to have: experience with heterogeneous or consumer-grade GPU fleets, CUDA, or ROCm
  • Nice to have: quantization formats including GGUF, AWQ, GPTQ, and FP8
  • Strong in at least two of Rust, LLM inference servers, and distributed systems

Benefits

Comp & perks
  • Unlimited PTO
  • 75% of health insurance premiums covered for you and your dependents
  • Dental and vision coverage
  • 401(k) plan
  • Stock options
  • Company-provided computer
  • $500 WFH budget
  • Fully remote, with flexible hours