Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
evoila

Senior AI Platform Engineer

evoila

. Plan, build, and operate highly available AI/ML platforms and services based on Kubernetes across on-premises, hybrid, and cloud environments .

Posted 9/30/2026full-timeMainz • GermanySeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in building and operating AI/ML platforms using Kubernetes, with a strong focus on scalable GPU infrastructures and MLOps workflows. Proficient in programming with Python and familiar with model-serving frameworks, while also possessing excellent communication skills in both German and English.

Highest-signal resume keywords
Kubernetes Platform EngineeringGPU Infrastructure DesignMLOps Workflow DevelopmentModel-Serving FrameworksTechnical Leadership

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Python ProgrammingGolang ProgrammingJava ProgrammingKubernetes ManagementMLOps ToolsInfrastructure-as-CodeCI/CD ToolsPerformance TuningAI GovernanceData Security
Soft Skills
Excellent CommunicationCustomer ConsultingMentoring
Tools & Technologies
NVIDIA TritonMLflowKubeflowClearMLAnsibleTerraformArgo CDFluxVLLMKubernetes Inference Gateway
Industry Keywords
AI/ML PlatformsLarge Language ModelsAI GovernanceDACH RegionEU AI Act

Tech Stack

Tools & technologies
AnsibleCloudFluxJavaKubernetesPythonTerraformGo

About the role

Key responsibilities & impact
  • Plan, build, and operate highly available AI/ML platforms and services based on Kubernetes across on-premises, hybrid, and cloud environments
  • Design and implement scalable GPU infrastructures in Kubernetes clusters, including GPU scheduling and sharing
  • Build and operate model-serving and inference services for traditional ML models and Large Language Models
  • Develop and operate MLOps/LLMOps workflows for production-grade GenAI services
  • Implement guardrails, policies, evaluations, regression testing, and cost and quality monitoring
  • Provide fine-tuning and re-training workflows
  • Monitor, troubleshoot, and performance-tune AI platforms
  • Advise customers on best practices, AI governance, data security, and data privacy
  • Collaborate with data science, data platform, and development teams to integrate AI solutions
  • In the Senior role: take ownership of architecture, make key technology decisions, provide technical leadership in customer projects, and mentor colleagues

Requirements

What you’ll need
  • At least 3 years of experience operating and optimizing platforms and services at scale in Kubernetes environments
  • Strong knowledge of the architecture and operation of scalable solutions using Kubernetes, ideally including GPU workloads
  • Programming skills in Python, ideally complemented by Golang or Java
  • Experience with at least one model-serving framework, such as vLLM, NVIDIA Triton, or KServe
  • Initial or advanced experience with MLOps tools, such as MLflow, Kubeflow, or ClearML
  • Experience with Infrastructure-as-Code tools such as Ansible and Terraform
  • Experience with CI/CD and GitOps tools, such as Argo CD or Flux
  • Excellent communication skills in German and English
  • Ideally, experience in customer consulting and communicating technical topics to different target audiences
  • For the Senior role: at least 5 years of relevant experience in platform engineering
  • For the Senior role: responsibility for the end-to-end design of AI platforms and key technology decisions
  • For the Senior role: technical leadership in customer projects and mentoring colleagues
  • Ideally, experience with cloud AI services and their integration with Kubernetes workloads
  • Knowledge of operating LLM-based architectures, RAG pipelines, vector databases, embedding services, agentic AI approaches, and MCP
  • Understanding of system-level fundamentals of LLM serving and LLM concepts such as reasoning, tool calling, and prompt templates
  • Experience with llm-d or the Kubernetes Inference Gateway
  • Knowledge of TLS, RBAC, and network policies within Kubernetes environments
  • Basic understanding of AI governance, such as the EU AI Act
  • Willingness to travel occasionally within the DACH region

Benefits

Comp & perks
  • Company pension scheme
  • Collaborative company culture with regular team events
  • Targeted support for your professional development
  • Corporate Benefits program
  • Flexible working hours
  • Modern workplace with up-to-date equipment
  • Remote work opportunities