FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in managing AI Gateway systems, including request routing and performance optimization, while effectively collaborating with cross-functional teams. Proficient in operating LLM inference servers and implementing observability solutions to ensure system reliability and efficiency.
Highest-signal resume keywords
Rust ProgrammingLLM Inference ServersDistributed Systems FundamentalsOpenTelemetry ObservabilityGPU Hardware Benchmarking
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
RustLLM Inference ServersDistributed SystemsPerformance ScoringNode AutoscalingKV Cache BehaviorMetrics and TracingQuantization FormatsLoad BalancingFailure Handling
Soft Skills
Clear CommunicationCollaboration
Tools & Technologies
VLLMLlama.cppTensorRT-LLMOpenTelemetryCUDAROCm
Industry Keywords
AI GatewayFleet Efficiency AlgorithmsSubscription Billing IntegrationsCross-System DesignOn-Call Responsibilities
Tech Stack
Tools & technologiesDistributed SystemsNode.jsRust
About the role
Key responsibilities & impact- Own the AI Gateway system end to end, including request routing, streaming, batch/async job handling, and horizontal scaling across multiple gateway servers
- Build and tune fleet-efficiency algorithms for node autoscaling, performance scoring, and eviction of underperformers
- Operate and scale inference nodes running vLLM, llama.cpp, and similar servers
- Design and maintain OpenTelemetry-based observability across the gateway and fleet
- Maintain accurate pay-per-token and subscription billing integrations based on telemetry
- Benchmark new models, quantizations, and GPU hardware for approximately 25% of working time
- Work with product and marketing teams to decide what enters production
- Collaborate on cross-system design and integrations with Salad engineering
- Explain technical trade-offs to technical and non-technical teammates
- Report directly to the CTO
- Participate in shared after-hours support
Requirements
What you’ll need- Strong production Rust experience, ideally on high-throughput networked services
- Hands-on experience running LLM inference servers (vLLM, llama.cpp, TGI, TensorRT-LLM, or similar)
- Knowledge of KV cache behavior
- Distributed systems fundamentals: load balancing, backpressure, failure handling, and unreliable nodes
- Experience operating systems with metrics, tracing, and on-call responsibilities
- Willingness to be on-call
- Clear written and verbal communication
- Nice to have: Pingora, Tokio, or proxy/gateway internals
- Nice to have: experience with heterogeneous or consumer-grade GPU fleets, CUDA, or ROCm
- Nice to have: quantization formats including GGUF, AWQ, GPTQ, and FP8
- Strong in at least two of Rust, LLM inference servers, and distributed systems
Benefits
Comp & perks- Unlimited PTO
- 75% of health insurance premiums covered for you and your dependents
- Dental and vision coverage
- 401(k) plan
- Stock options
- Company-provided computer
- $500 WFH budget
- Fully remote, with flexible hours
