Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Mistral AI

Research Engineer, Inference Foundation

Mistral AI

. Develop and fix the core inference stack, including the engine and orchestrator, feature selection, configuration, and performance tuning at scale .

Posted 10/6/2026full-timeParis • FranceMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Expertise in building and optimizing machine learning and large language model services at scale, with a strong focus on inference performance, serving infrastructure, and hardware-aware optimizations. Proficient in Python, PyTorch, and Kubernetes, with hands-on experience in debugging and optimizing distributed systems.

Highest-signal resume keywords
ML/LLM Services DevelopmentInference Engine OptimizationKubernetes Infrastructure ManagementCUDA/NCCL DebuggingPerformance Tuning

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
PythonPyTorchCUDANCCLVLLMSGLangTensorRT-LLMRustC++Triton
Tools & Technologies
KubernetesNsight SystemsNsight Compute
Industry Keywords
Machine LearningLarge Language ModelsDistributed ArchitectureInference PerformanceServing Topology

Tech Stack

Tools & technologies
KubernetesPythonPyTorchRustC++

About the role

Key responsibilities & impact
  • Develop and fix the core inference stack, including the engine and orchestrator, feature selection, configuration, and performance tuning at scale
  • Own validated, regression-free serving-stack releases through automated performance gates and progressive rollout
  • Drive improvements and fixes upstream when the open-source engine is the right place
  • Optimize serving efficiency across the fleet by reducing pod startup time, addressing cold-cache regressions during scale-up, and improving caching and offloading
  • Optimize and maintain serving topology by overlapping communication and transfers with computation and ensuring optimal placement, connectivity, and routing
  • Build serving infrastructure powering RL and post-training for frontier models
  • Optimize inference performance across the full spectrum of workloads

Requirements

What you’ll need
  • Experience building and running ML/LLM services at scale, with clear latency and availability targets
  • Hands-on experience with inference engines such as vLLM, SGLang, TensorRT-LLM, or others
  • Solid grasp of inference internals: prefill vs. decode, KV-cache behavior, batching, scheduling, speculative decoding, and parallelism strategies
  • Familiarity with distributed and disaggregated serving architectures
  • Comfortable debugging across CUDA/NCCL, kernels, containers, networking, and storage
  • Python for systems tooling and backend services
  • PyTorch
  • Kubernetes for running infrastructure at scale
  • GPU and networking fundamentals: CUDA runtime, NCCL, InfiniBand/RDMA
  • Demonstrated vLLM/sglang know-how, ideally through upstream contributions or demanding production environments
  • Hardware-aware optimization for various model architectures
  • Experience serving MoE models at scale, including expert parallelism, expert placement, and load balancing
  • CUDA/Triton kernel development
  • Nsight Systems/Compute profiling
  • Rust and/or C++ in production systems

Benefits

Comp & perks
  • Healthcare coverage
  • Parental leave
  • Retirement plans
  • Relocation support
  • Wellness programs
  • Meal allowances
  • Transportation allowances
  • Other location-specific perks