Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Designworks Talent LLC

Staff Inference Engineer

Designworks Talent LLC

. Build and operate production-grade model-serving and inference systems for high-throughput, low-latency AI workloads .

Posted 10/2/2026full-timeBellevue • Washington • United StatesLeadWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in building and operating production-grade machine learning inference systems, optimizing for latency, throughput, and cost efficiency. Proficient in designing reliable distributed systems and managing GPU-backed AI workloads to ensure scalability and performance.

Highest-signal resume keywords
Production Machine Learning Inference SystemsGPU Utilization OptimizationDistributed Systems DesignInference-Serving FrameworksAPI-Based AI Products

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Model-Serving SystemsLatency OptimizationThroughput ManagementMemory UtilizationCost-Efficiency Trade-OffsGPU SchedulingPerformance TuningQuantizationBatchingCloud Infrastructure
Soft Skills
Problem SolvingAdaptabilityCollaboration
Tools & Technologies
VLLMSGLangTensorRT-LLMTriton Inference ServerKubernetes
Industry Keywords
AI WorkloadsHigh-Volume Production ServicesHyperscalerAI LabML Infrastructure

Tech Stack

Tools & technologies
CloudDistributed SystemsKubernetes

About the role

Key responsibilities & impact
  • Build and operate production-grade model-serving and inference systems for high-throughput, low-latency AI workloads
  • Optimize inference infrastructure for token throughput, latency, scalability, and cost efficiency across model architectures and workloads
  • Design systems that maximize GPU utilization while maintaining predictable performance and reliability
  • Improve inference-platform scalability and operational maturity as customer demand grows
  • Partner with AI training, GPU performance, orchestration, and infrastructure teams to transition models from development to production serving
  • Develop monitoring, alerting, and operational practices for reliable inference services
  • Investigate and resolve performance, reliability, and capacity challenges across inference workloads
  • Contribute to architecture decisions and engineering standards as the platform evolves

Requirements

What you’ll need
  • Experience building and operating production machine learning inference or model-serving systems at scale
  • Strong understanding of latency, throughput, memory utilization, and cost-efficiency trade-offs when serving large AI models
  • Experience designing reliable distributed systems or production infrastructure
  • Understanding of GPU-backed AI workloads and the challenges of scaling inference systems
  • Strong engineering fundamentals and ability to independently own complex technical problems
  • Comfortable working in a fast-moving environment where systems and processes are built from the ground up
  • Experience with modern inference-serving frameworks such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or similar technologies
  • Experience optimizing LLM inference workloads or large-scale AI serving platforms
  • Background operating API-based AI products or high-volume production services
  • Experience with GPU scheduling, distributed systems, Kubernetes, or cloud infrastructure platforms
  • Familiarity with quantization, batching, caching, or performance tuning
  • Experience working at a hyperscaler, AI lab, GPU cloud provider, or large-scale ML infrastructure organization
  • U.S. work authorization is required
  • Visa sponsorship is not currently available
  • Willingness to work in a hybrid situation based in downtown Bellevue, WA, with a minimum of three days per week in-office once the permanent office is established

Benefits

Comp & perks
  • Competitive base pay for Bellevue market
  • Merit increases
  • Annual bonus
  • Long-term incentives
  • Medical insurance
  • Dental insurance
  • Vision insurance
  • 401(k) plan
  • Company 401(k) match
  • Paid holidays
  • Startup ownership and technical impact
  • Collaboration with an experienced AI infrastructure team