FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and operating production infrastructure, with a strong focus on automation for GPU clusters and cloud environments. Proficient in troubleshooting distributed systems and improving operational workflows through advanced tools and methodologies.
Highest-signal resume keywords
Python ProgrammingKubernetes ExperienceInfrastructure AutomationGitOps ImplementationIncident Response
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Production InfrastructureDistributed Systems TroubleshootingCloud InfrastructureLinux ProficiencyGPU InfrastructureKubernetes OperatorsTerraformArgoCDFleet AutomationSLOs
Soft Skills
Clear CommunicationTeam Collaboration
Tools & Technologies
APIsGitOpsAutomation ToolsMonitoring ToolsIncident Response Tools
Certifications & Qualifications
BS/MS in Computer Science
Industry Keywords
BMaaSVMaaSManaged KubernetesMulti-Cloud InfrastructureObservability Practices
Tech Stack
Tools & technologiesCloudDistributed SystemsKubernetesLinuxPythonTerraformGo
About the role
Key responsibilities & impact- Build and operate automation for large-scale GPU clusters across NVIDIA Cloud Partners and on-prem environments
- Develop tools and services for provisioning, validation, upgrades, monitoring, repair, and cluster lifecycle operations
- Improve Day 0 / Day 1 / Day 2 workflows for cluster bringup, handoff, and production operations
- Reduce manual production touches through APIs, GitOps, automation, and agent-assisted workflows
- Participate in on-call, incident response, debugging, and durable follow-up work
- Partner with platform, storage, networking, security, and workload teams to make infrastructure production-ready
Requirements
What you’ll need- 8+ years of experience building or operating production infrastructure
- Strong programming skills in Python, Go, or similar
- Experience with Linux, Kubernetes, containers, cloud infrastructure, or infrastructure automation
- Ability to troubleshoot distributed systems in production
- Clear communication and ability to work across teams
- BS/MS in Computer Science or equivalent experience
- Experience with GPU infrastructure, Kubernetes operators, GitOps, Terraform, ArgoCD, or fleet automation
- Experience with SLOs, on-call, incident response, observability, and reliability practices
- Exposure to BMaaS, VMaaS, managed Kubernetes, or multi-cloud infrastructure
Benefits
Comp & perks- Equity
- Benefits 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score
