Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Luma AI

Software Engineer, Inference

Luma AI

. Own how Luma's models get served through inference-engine integration, deployment scaling, and GPU fleet utilization .

Posted 10/8/2026full-timeRemoteMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in deploying and optimizing machine learning models at scale, with a strong focus on Python, system architecture, and GPU resource management. Proficient in building and maintaining reliable inference services and CI/CD pipelines across diverse hardware environments.

Highest-signal resume keywords
Python ProgrammingModel Deployment with PyTorchKubernetes OrchestrationCI/CD for Model CheckpointsGPU Resource Management

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
System ArchitectureModel IntegrationInference Engine OptimizationTraffic ControlScheduling SystemsLinuxDockerRedisS3-Compatible StorageCUDA
Tools & Technologies
Hugging FaceVLLMSGLangTensorRT-LLMFFmpeg
Industry Keywords
Inference WorkflowsHigh-Performance ML SystemsFleet ManagementDeployment ScalingModern Networking Stacks

Tech Stack

Tools & technologies
DockerFFmpegKubernetesLinuxPythonPyTorchRedis

About the role

Key responsibilities & impact
  • Own how Luma's models get served through inference-engine integration, deployment scaling, and GPU fleet utilization
  • Integrate new model architectures into the inference engine
  • Collaborate across research, engineering, and infrastructure to optimize model efficiency and deployments
  • Build internal tooling to measure, profile, and track inference jobs and workflows
  • Automate, test, and maintain inference services for uptime and reliability
  • Manage and optimize inference workloads across clusters and hardware providers
  • Scale deployments across thousands of machines
  • Build scheduling systems that optimize GPU resources while meeting SLOs
  • Maintain CI/CD for model checkpoints and SDKs
  • Learn the inference stack and diagnose reliability or utilization issues during the first 30 days
  • Integrate a model or ship tooling/scheduling improvements during days 30–60
  • Harden deployment pipelines and scheduling across clusters and providers during days 60–90

Requirements

What you’ll need
  • Strong Python and system-architecture skills
  • Experience deploying models with PyTorch, Hugging Face, vLLM, SGLang, TensorRT-LLM, or similar
  • Experience with queues, scheduling, traffic control, and fleet management at scale
  • Experience with Linux, Docker, and Kubernetes
  • Experience with orchestration, deployment, and scheduling
  • Familiarity with Redis and S3-compatible storage
  • Nice to have: modern networking stacks including RDMA (RoCE, InfiniBand, NVLink)
  • Nice to have: high-performance large-scale ML systems (100+ GPUs)
  • Nice to have: CUDA, and FFmpeg or multimedia processing

Benefits

Comp & perks
  • Equal opportunity employer
  • Voluntary diversity and inclusion survey participation; refusal does not affect the job application
  • Remote work arrangement