FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in deploying and optimizing machine learning models at scale, with a strong focus on Python, system architecture, and high-performance computing. Proficient in managing inference workloads and ensuring reliability across diverse hardware environments.
Highest-signal resume keywords
Python ProgrammingModel Deployment with PyTorchKubernetes OrchestrationCI/CD for Model CheckpointsHigh-Performance ML Systems
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
System ArchitectureModel DeploymentTraffic ControlFleet ManagementLinuxDockerSchedulingCUDAFFmpegRedis
Tools & Technologies
Hugging FaceVLLMSGLangTensorRT-LLMS3-Compatible StorageRDMARoCEInfiniBandNVLink
Industry Keywords
Inference EngineGPU UtilizationModel EfficiencyInternal ToolingInference Workflows
Tech Stack
Tools & technologiesDockerFFmpegKubernetesLinuxPythonPyTorchRedis
About the role
Key responsibilities & impact- Own how Luma's models get served by integrating new architectures into the inference engine
- Scale deployments across thousands of machines
- Keep expensive GPU fleets utilized while meeting internal SLOs
- Collaborate across research, engineering, and infrastructure to optimize model efficiency and deployments
- Build internal tooling to measure, profile, and track inference jobs and workflows
- Automate, test, and maintain inference services for uptime and reliability
- Manage and optimize inference workloads across clusters and hardware providers
- Build scheduling systems that optimize GPU resources while meeting SLOs
- Maintain CI/CD for model checkpoints and SDKs
- Learn the inference stack, fleets, and reliability or utilization issues during the first 30 days
- Integrate a model or ship tooling/scheduling improvements during days 30–60
- Harden deployment pipelines and scheduling across clusters and providers during days 60–90
Requirements
What you’ll need- Strong Python and system-architecture skills
- Experience deploying models with PyTorch, Hugging Face, vLLM, SGLang, TensorRT-LLM, or similar
- Experience with queues, scheduling, traffic control, and fleet management at scale
- Experience with Linux, Docker, and Kubernetes
- Experience with orchestration, deployment, and scheduling
- Familiarity with Redis and S3-compatible storage
- Modern networking stacks including RDMA (RoCE, InfiniBand, NVLink) preferred
- Experience with high-performance large-scale ML systems involving 100+ GPUs preferred
- CUDA and FFmpeg or multimedia processing preferred
Benefits
Comp & perks- Equal opportunity employer
- Optional voluntary diversity and inclusion survey; participation does not affect the job application
- Global Equal Employment Opportunity protections
