FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Research Engineer, Inference Foundation
Mistral AI. Develop and fix the core inference stack, including the engine and orchestrator, feature selection, configuration, and performance tuning at scale .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Expertise in building and optimizing machine learning and large language model services at scale, with a strong focus on inference performance, serving infrastructure, and hardware-aware optimizations. Proficient in Python, PyTorch, and Kubernetes, with hands-on experience in debugging and optimizing distributed systems.
Highest-signal resume keywords
ML/LLM Services DevelopmentInference Engine OptimizationKubernetes Infrastructure ManagementCUDA/NCCL DebuggingPerformance Tuning
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
PythonPyTorchCUDANCCLVLLMSGLangTensorRT-LLMRustC++Triton
Tools & Technologies
KubernetesNsight SystemsNsight Compute
Industry Keywords
Machine LearningLarge Language ModelsDistributed ArchitectureInference PerformanceServing Topology
Tech Stack
Tools & technologiesKubernetesPythonPyTorchRustC++
About the role
Key responsibilities & impact- Develop and fix the core inference stack, including the engine and orchestrator, feature selection, configuration, and performance tuning at scale
- Own validated, regression-free serving-stack releases through automated performance gates and progressive rollout
- Drive improvements and fixes upstream when the open-source engine is the right place
- Optimize serving efficiency across the fleet by reducing pod startup time, addressing cold-cache regressions during scale-up, and improving caching and offloading
- Optimize and maintain serving topology by overlapping communication and transfers with computation and ensuring optimal placement, connectivity, and routing
- Build serving infrastructure powering RL and post-training for frontier models
- Optimize inference performance across the full spectrum of workloads
Requirements
What you’ll need- Experience building and running ML/LLM services at scale, with clear latency and availability targets
- Hands-on experience with inference engines such as vLLM, SGLang, TensorRT-LLM, or others
- Solid grasp of inference internals: prefill vs. decode, KV-cache behavior, batching, scheduling, speculative decoding, and parallelism strategies
- Familiarity with distributed and disaggregated serving architectures
- Comfortable debugging across CUDA/NCCL, kernels, containers, networking, and storage
- Python for systems tooling and backend services
- PyTorch
- Kubernetes for running infrastructure at scale
- GPU and networking fundamentals: CUDA runtime, NCCL, InfiniBand/RDMA
- Demonstrated vLLM/sglang know-how, ideally through upstream contributions or demanding production environments
- Hardware-aware optimization for various model architectures
- Experience serving MoE models at scale, including expert parallelism, expert placement, and load balancing
- CUDA/Triton kernel development
- Nsight Systems/Compute profiling
- Rust and/or C++ in production systems
Benefits
Comp & perks- Healthcare coverage
- Parental leave
- Retirement plans
- Relocation support
- Wellness programs
- Meal allowances
- Transportation allowances
- Other location-specific perks