FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Staff Inference Engineer
Designworks Talent LLC. Build and operate production-grade model-serving and inference systems for high-throughput, low-latency AI workloads .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and operating production-grade machine learning inference systems, optimizing for latency, throughput, and cost efficiency. Proficient in designing reliable distributed systems and managing GPU-backed AI workloads to ensure scalability and performance.
Highest-signal resume keywords
Production Machine Learning Inference SystemsGPU Utilization OptimizationDistributed Systems DesignInference-Serving FrameworksAPI-Based AI Products
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Model-Serving SystemsLatency OptimizationThroughput ManagementMemory UtilizationCost-Efficiency Trade-OffsGPU SchedulingPerformance TuningQuantizationBatchingCloud Infrastructure
Soft Skills
Problem SolvingAdaptabilityCollaboration
Tools & Technologies
VLLMSGLangTensorRT-LLMTriton Inference ServerKubernetes
Industry Keywords
AI WorkloadsHigh-Volume Production ServicesHyperscalerAI LabML Infrastructure
Tech Stack
Tools & technologiesCloudDistributed SystemsKubernetes
About the role
Key responsibilities & impact- Build and operate production-grade model-serving and inference systems for high-throughput, low-latency AI workloads
- Optimize inference infrastructure for token throughput, latency, scalability, and cost efficiency across model architectures and workloads
- Design systems that maximize GPU utilization while maintaining predictable performance and reliability
- Improve inference-platform scalability and operational maturity as customer demand grows
- Partner with AI training, GPU performance, orchestration, and infrastructure teams to transition models from development to production serving
- Develop monitoring, alerting, and operational practices for reliable inference services
- Investigate and resolve performance, reliability, and capacity challenges across inference workloads
- Contribute to architecture decisions and engineering standards as the platform evolves
Requirements
What you’ll need- Experience building and operating production machine learning inference or model-serving systems at scale
- Strong understanding of latency, throughput, memory utilization, and cost-efficiency trade-offs when serving large AI models
- Experience designing reliable distributed systems or production infrastructure
- Understanding of GPU-backed AI workloads and the challenges of scaling inference systems
- Strong engineering fundamentals and ability to independently own complex technical problems
- Comfortable working in a fast-moving environment where systems and processes are built from the ground up
- Experience with modern inference-serving frameworks such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or similar technologies
- Experience optimizing LLM inference workloads or large-scale AI serving platforms
- Background operating API-based AI products or high-volume production services
- Experience with GPU scheduling, distributed systems, Kubernetes, or cloud infrastructure platforms
- Familiarity with quantization, batching, caching, or performance tuning
- Experience working at a hyperscaler, AI lab, GPU cloud provider, or large-scale ML infrastructure organization
- U.S. work authorization is required
- Visa sponsorship is not currently available
- Willingness to work in a hybrid situation based in downtown Bellevue, WA, with a minimum of three days per week in-office once the permanent office is established
Benefits
Comp & perks- Competitive base pay for Bellevue market
- Merit increases
- Annual bonus
- Long-term incentives
- Medical insurance
- Dental insurance
- Vision insurance
- 401(k) plan
- Company 401(k) match
- Paid holidays
- Startup ownership and technical impact
- Collaboration with an experienced AI infrastructure team