FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in deploying and optimizing machine learning models at scale, with a strong focus on Python, system architecture, and GPU resource management. Proficient in building and maintaining reliable inference services and CI/CD pipelines across diverse hardware environments.
Highest-signal resume keywords
Python ProgrammingModel Deployment with PyTorchKubernetes OrchestrationCI/CD for Model CheckpointsGPU Resource Management
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
System ArchitectureModel IntegrationInference Engine OptimizationTraffic ControlScheduling SystemsLinuxDockerRedisS3-Compatible StorageCUDA
Tools & Technologies
Hugging FaceVLLMSGLangTensorRT-LLMFFmpeg
Industry Keywords
Inference WorkflowsHigh-Performance ML SystemsFleet ManagementDeployment ScalingModern Networking Stacks
Tech Stack
Tools & technologiesDockerFFmpegKubernetesLinuxPythonPyTorchRedis
About the role
Key responsibilities & impact- Own how Luma's models get served through inference-engine integration, deployment scaling, and GPU fleet utilization
- Integrate new model architectures into the inference engine
- Collaborate across research, engineering, and infrastructure to optimize model efficiency and deployments
- Build internal tooling to measure, profile, and track inference jobs and workflows
- Automate, test, and maintain inference services for uptime and reliability
- Manage and optimize inference workloads across clusters and hardware providers
- Scale deployments across thousands of machines
- Build scheduling systems that optimize GPU resources while meeting SLOs
- Maintain CI/CD for model checkpoints and SDKs
- Learn the inference stack and diagnose reliability or utilization issues during the first 30 days
- Integrate a model or ship tooling/scheduling improvements during days 30–60
- Harden deployment pipelines and scheduling across clusters and providers during days 60–90
Requirements
What you’ll need- Strong Python and system-architecture skills
- Experience deploying models with PyTorch, Hugging Face, vLLM, SGLang, TensorRT-LLM, or similar
- Experience with queues, scheduling, traffic control, and fleet management at scale
- Experience with Linux, Docker, and Kubernetes
- Experience with orchestration, deployment, and scheduling
- Familiarity with Redis and S3-compatible storage
- Nice to have: modern networking stacks including RDMA (RoCE, InfiniBand, NVLink)
- Nice to have: high-performance large-scale ML systems (100+ GPUs)
- Nice to have: CUDA, and FFmpeg or multimedia processing
Benefits
Comp & perks- Equal opportunity employer
- Voluntary diversity and inclusion survey participation; refusal does not affect the job application
- Remote work arrangement
