FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and operating production-grade online ML inference systems, with a strong focus on optimizing model performance and ensuring system reliability. Proficient in leveraging tools and frameworks for distributed training and model serving, while maintaining observability and cost efficiency.
Highest-signal resume keywords
Production-Grade Online ML Inference SystemsModel Serving FrameworksDistributed SystemsPython ProgrammingModel Deployment Workflows
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Model CompilationDynamic BatchingGPU AccelerationQuantizationKubernetesService ReliabilityModel PackagingArtifact ValidationA/B ExperimentationRuntime Tuning
Soft Skills
Systems ThinkingTechnical Leadership
Tools & Technologies
PyTorchRay DataRay TrainNVIDIA Triton Inference ServerTorchServeRay ServeTensorFlow ServingFlyteAirflow
Industry Keywords
ML InfrastructureModel Serving WorkflowsObservabilityScalabilityCost Efficiency
Tech Stack
Tools & technologiesAirflowDistributed SystemsKubernetesPythonPyTorchRayTensorflow
About the role
Key responsibilities & impact- Design and operate large-scale online inference infrastructure serving production ML models with low latency and high reliability
- Develop infrastructure supporting distributed training workflows using PyTorch, Ray Data, and Ray Train
- Integrate ML pipelines with workflow orchestration systems such as Flyte or Airflow
- Optimize model performance through model compilation, GPU/CPU utilization improvements, request scheduling, kernel fusion, and runtime-level tuning
- Improve ML systems observability through latency, throughput, error-rate, cost, saturation, and model-health monitoring
- Partner with ML engineers to support faster model iteration while maintaining production safety, scalability, and cost efficiency
- Improve reliability and reproducibility of model serving workflows, including model packaging, artifact validation, compatibility testing, and deployment automation
- Lead architectural improvements to make the online ML platform more robust, user-friendly, scalable, and cost-efficient
Requirements
What you’ll need- Experience building and operating production-grade online ML inference systems
- Experience with model serving frameworks such as NVIDIA Triton Inference Server, TorchServe, Ray Serve, TensorFlow Serving, or similar systems
- Experience optimizing inference workloads using dynamic batching, model compilation, quantization, GPU acceleration, GPU kernel optimization, caching, or runtime tuning
- Strong experience with distributed systems, Kubernetes, autoscaling, service reliability, and production observability
- Strong programming skills in Python, with practical experience working on production ML systems and high-scale services
- Experience with PyTorch and modern model deployment workflows, including model packaging, validation, and serving lifecycle management
- Experience designing infrastructure for safe model rollout, canary testing, A/B experimentation, and automated rollback
- Strong systems thinking and ability to reason about latency, throughput, reliability, scalability, and cost tradeoffs
- Proven ability to lead technical direction and influence architectural decisions across teams without formal authority
- Sufficient knowledge of English for professional verbal and written exchanges
Benefits
Comp & perks- Equity awards
- Participation in company incentive plans, such as annual discretionary bonuses or sales commissions
- Comprehensive health, life, and disability insurance
- Commute subsidy
- Employee stock ownership
- Competitive retirement/pension plans
- Generous vacation and personal days
- Leave and family-care programs for new parents
- Office food snacks
- Mental Health and Wellbeing programs and support
- Employee Resource Groups
- Global Employee Assistance Program
- Training and development programs
- Volunteering and donation matching program
