FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and managing cloud-native infrastructure for large-scale machine learning workloads, with a strong focus on Kubernetes, CI/CD pipelines, and observability solutions. Proven ability to lead infrastructure modernization initiatives and optimize resource utilization across distributed systems.
Highest-signal resume keywords
Kubernetes ManagementCI/CD Pipeline DevelopmentInfrastructure-as-CodePython ProgrammingObservability Solutions
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
KubernetesPythonGolangTerraformHelmGitHub ActionsJenkinsGitLab CI/CDAWS EKSLinux Systems Administration
Tools & Technologies
VolcanoKueuePrometheusGrafanaAzure DevOps
Industry Keywords
Cloud-Native ArchitectureDistributed SystemsMicroservicesML Lifecycle ManagementInfrastructure Migration
Tech Stack
Tools & technologiesAWSAzureCloudDistributed SystemsGrafanaJenkinsKubernetesLinuxMicroservicesPrometheusPythonTerraformGo
About the role
Key responsibilities & impact- Build, operate, and evolve large-scale, business-critical infrastructure platforms
- Develop and manage cloud-native infrastructure supporting large-scale ML workloads on Kubernetes
- Implement and operate batch scheduling solutions such as Volcano and Kueue to optimize GPU and compute resource utilization
- Build and maintain CI/CD pipelines for ML services, infrastructure, and platform components
- Manage and optimize Kubernetes environments, including production workloads on AWS EKS and other cloud platforms
- Lead and support infrastructure modernization and migration initiatives across cloud and platform ecosystems
- Partner with Data Scientists, ML Engineers, and Software Engineers to productionize machine learning solutions
- Implement observability, monitoring, and alerting for distributed systems and GPU clusters
- Ensure platform reliability, security, scalability, and operational excellence
- Troubleshoot complex distributed systems and performance bottlenecks across infrastructure and ML workloads
- Improve developer productivity and drive innovation through automation and AI-assisted operations
Requirements
What you’ll need- Strong experience with Kubernetes and large-scale workload orchestration
- Hands-on experience with Kubernetes batch schedulers such as Volcano and Kueue
- Solid understanding of distributed systems, containerization technologies, cloud-native architectures, and microservices-based platforms
- Experience managing workloads on Kubernetes platforms such as AWS EKS
- Proven track record delivering and supporting infrastructure migration projects
- Strong programming skills in Python or Golang
- Experience with Infrastructure-as-Code and deployment tools including Terraform and Helm
- Experience designing and maintaining CI/CD pipelines using GitHub Actions, Jenkins, GitLab CI/CD, or Azure DevOps
- Strong Linux systems administration and troubleshooting skills
- Experience building observability solutions using Prometheus and Grafana
- Understanding of ML lifecycle management, model deployment, and production operations
Benefits
Comp & perks- Paid time off: Vacation (12–25 days depending on grade), company-paid holidays, personal days, and sick leave
- Medical, dental, and vision coverage, or provincial healthcare coordination in Canada
- Retirement savings plans, including RRSP in Canada
- Life and disability insurance
- Employee assistance programs
- Other benefits provided by local policy and eligibility
- Potential eligibility for variable incentives, bonuses, or commissions
