Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Luma AI

Software Engineer, Inference

Luma AI

. Own how Luma's models get served by integrating new architectures into the inference engine .

Posted 10/8/2026full-timeLondon • United KingdomMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in deploying and optimizing machine learning models at scale, with a strong focus on Python, system architecture, and high-performance computing. Proficient in managing inference workloads and ensuring reliability across diverse hardware environments.

Highest-signal resume keywords
Python ProgrammingModel Deployment with PyTorchKubernetes OrchestrationCI/CD for Model CheckpointsHigh-Performance ML Systems

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
System ArchitectureModel DeploymentTraffic ControlFleet ManagementLinuxDockerSchedulingCUDAFFmpegRedis
Tools & Technologies
Hugging FaceVLLMSGLangTensorRT-LLMS3-Compatible StorageRDMARoCEInfiniBandNVLink
Industry Keywords
Inference EngineGPU UtilizationModel EfficiencyInternal ToolingInference Workflows

Tech Stack

Tools & technologies
DockerFFmpegKubernetesLinuxPythonPyTorchRedis

About the role

Key responsibilities & impact
  • Own how Luma's models get served by integrating new architectures into the inference engine
  • Scale deployments across thousands of machines
  • Keep expensive GPU fleets utilized while meeting internal SLOs
  • Collaborate across research, engineering, and infrastructure to optimize model efficiency and deployments
  • Build internal tooling to measure, profile, and track inference jobs and workflows
  • Automate, test, and maintain inference services for uptime and reliability
  • Manage and optimize inference workloads across clusters and hardware providers
  • Build scheduling systems that optimize GPU resources while meeting SLOs
  • Maintain CI/CD for model checkpoints and SDKs
  • Learn the inference stack, fleets, and reliability or utilization issues during the first 30 days
  • Integrate a model or ship tooling/scheduling improvements during days 30–60
  • Harden deployment pipelines and scheduling across clusters and providers during days 60–90

Requirements

What you’ll need
  • Strong Python and system-architecture skills
  • Experience deploying models with PyTorch, Hugging Face, vLLM, SGLang, TensorRT-LLM, or similar
  • Experience with queues, scheduling, traffic control, and fleet management at scale
  • Experience with Linux, Docker, and Kubernetes
  • Experience with orchestration, deployment, and scheduling
  • Familiarity with Redis and S3-compatible storage
  • Modern networking stacks including RDMA (RoCE, InfiniBand, NVLink) preferred
  • Experience with high-performance large-scale ML systems involving 100+ GPUs preferred
  • CUDA and FFmpeg or multimedia processing preferred

Benefits

Comp & perks
  • Equal opportunity employer
  • Optional voluntary diversity and inclusion survey; participation does not affect the job application
  • Global Equal Employment Opportunity protections