FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Machine Learning Platform Engineer – Model Hosting, MLOps
HP. Design and build hosting for custom ML and AI models across AWS and Azure, focusing on GPU-backed LLM inference and real-time endpoints .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and building hosting solutions for ML and AI models across AWS and Azure, with a strong focus on GPU-backed LLM inference and MLOps workflows. Proficient in infrastructure as code, Kubernetes provisioning, and developing automated deployment pipelines for scalable and reliable model serving.
Highest-signal resume keywords
GPU Infrastructure HostingAWS and Azure ML Infrastructure DevelopmentKubernetes Provisioning and ConfigurationInfrastructure as Code (Terraform)MLOps Workflow Development
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Python ProgrammingModel-Serving PlatformsCI/CD AutomationBatch Inference SupportDeployment PatternsVersion Control for ArtifactsMonitoring and Incident ResponseSecure Connection DesignModel Evaluation and Quality MonitoringAgent Workflow Development
Soft Skills
Cross-Team CollaborationProblem-SolvingDocumentation Skills
Tools & Technologies
AWS SageMakerAzure Machine LearningKubernetesTerraformEKSAKS
Industry Keywords
MLOpsLLM InferenceContainerized WorkloadsProduction EnvironmentHybrid Hosting
Tech Stack
Tools & technologiesAWSAzureCloudKubernetesPythonTerraform
About the role
Key responsibilities & impact- Design and build hosting for custom ML and AI models across AWS and Azure, focusing on GPU-backed LLM inference and real-time endpoints
- Support batch inference where appropriate
- Package models and dependencies into reproducible serving workloads
- Select and implement managed services, containers, or Kubernetes-based hosting patterns based on throughput, latency, security, reliability, and cost
- Provision and configure Kubernetes clusters or other hosting infrastructure, including compute, storage, API access, identity and access controls, secrets, networking, observability, and environment configuration
- Design secure connections and deployment patterns between cloud and on-premises environments
- Develop infrastructure as code and deployment automation for consistent provisioning, review, promotion, and maintenance
- Build MLOps workflows for model registration, versioning, validation, release, rollback, and retirement
- Connect training or model preparation to deployment through automated pipelines and quality gates
- Partner with application teams to design and prototype agent workflows involving model and tool orchestration, state handling, failure recovery, and evaluation
- Translate agent workload patterns into hosting decisions concerning model selection, context length, concurrency, latency, cost, tool access, and end-to-end tracing
- Establish production monitoring for service health, latency, throughput, errors, GPU and resource use, and model behavior
- Diagnose incidents and improve capacity, reliability, and cost
- Create reusable deployment templates, reference architectures, documentation, and onboarding paths
- Partner with model and application teams on inference, serving, scaling, evaluation, data handling, and operational ownership tradeoffs
Requirements
What you’ll need- Hands-on experience hosting LLM inference on GPU infrastructure in a production environment
- Experience building model-serving platforms for custom models, including inference runtimes, deployment patterns, endpoint access, scaling, and operational tooling
- Strong software engineering skills, especially Python, with experience building services, automation, and maintainable production code
- Experience developing ML infrastructure across AWS and Azure, with deep hands-on delivery in at least one and practical ability to work in the other
- Experience using infrastructure as code such as Terraform or an equivalent tool
- Experience provisioning, configuring, and maintaining Kubernetes or another production hosting platform for containerized inference workloads
- Experience building or operating ML deployment pipelines with versioned artifacts, automated validation, CI/CD, environment promotion, and rollback
- Familiarity with LLM agent patterns, including model invocation, tool calls, and multi-step workflows
- Working knowledge of production concerns for inference services, including scaling, latency, availability, logging and metrics, incident response, access control, and cost
- Ability to work across model development, application, platform, and security teams; turn ambiguous requirements into working designs; and document operational approaches
- Helpful: advanced GPU inference optimization, generative or compute-intensive custom model serving, AWS SageMaker or EKS, Azure Machine Learning or AKS, hybrid or on-premises hosting, model registries, experiment tracking, data or feature pipelines, scheduled retraining, model evaluation, drift or quality monitoring, secure enterprise deployment patterns, shared ML platform capabilities, and agent workflow development
- No travel required
- No relocation provided
Benefits
Comp & perks- Bonus and/or equity opportunities (United States of America candidates only)
- Health insurance
- Dental insurance
- Vision insurance
- Long term/short term disability insurance
- Employee assistance program
- Flexible spending account
- Life insurance
- 4-12 weeks fully paid parental leave based on tenure
- 11 paid holidays
- Additional flexible paid vacation and sick leave