FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in infrastructure operations and systems reliability, with a strong focus on performance troubleshooting and technical incident management. Proficient in managing high-density GPU fleets and optimizing bare-metal HPC environments.
Highest-signal resume keywords
Infrastructure OperationsPerformance TroubleshootingNVIDIA Software StackLinux System AdministrationContainerization Using Docker
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Performance TuningSystem-Level TroubleshootingPythonGo (Golang)BashRDMAInfiniBandRoCEMulti-Node Performance TuningGrafana
Soft Skills
Clear CommunicationTechnical Translation
Tools & Technologies
DockerPrometheusDatadog
Industry Keywords
Datacenter EngineeringHPC EnvironmentsFleet Health MonitoringTechnical Lifecycle ManagementAI/ML Workloads
Tech Stack
Tools & technologiesDockerGrafanaLinuxNode.jsPrometheusPythonGo
About the role
Key responsibilities & impact- Validate new hardware and ensure partner deployments meet Runpod specifications for distributed AI/ML workloads
- Monitor fleet health and identify performance degradation
- Audit downtime and provide technical data needed to protect customer SLAs
- Use LLMs and AI agents to automate network triage and generate dynamic fleet runbooks
- Coordinate technical incident communications and translate outages into actionable resolutions
- Support the growth of infrastructure partners
- Own the technical lifecycle and operational health of Runpod’s high-density GPU fleet
- Serve as infrastructure advisor, technical translator, adopter, and incident commander for hardware partners
- Bridge hardware partners and internal engineering teams
Requirements
What you’ll need- 3–5 years of experience in infrastructure operations, systems reliability, or datacenter engineering
- Strong proficiency in standard datacenter networking and performance troubleshooting
- Exposure to RDMA, InfiniBand, or RoCE highly preferred
- Hands-on experience with the NVIDIA Software Stack, including driver installation and performance utilities
- Understanding of multi-node performance tuning
- Solid Linux system administration skills
- Experience with containerization using Docker
- System-level troubleshooting and performance tuning at kernel and hardware interface layers
- Clear written and verbal communication skills
- Ability to explain hardware or networking issues to technical partners and internal leadership
- Willingness to participate in a future on-call rotation
- Experience in a fast-paced startup environment preferred
- Experience managing or optimizing bare-metal HPC environments at massive scale preferred
- Experience with Grafana, Prometheus, or Datadog preferred
- Proficiency in Python, Go (Golang), or Bash preferred
- Must be eligible to work in the United States; Runpod is currently unable to sponsor employment visas
Benefits
Comp & perks- Meaningful equity; everyone on the team receives stock options
- Generous medical, dental & vision plans; 100% coverage for employees and partial coverage for dependents
- Flexible PTO
- Remote-first work with inclusive, collaborative teams
- $1,200 Home Office & Equipment Stipend
- Culture, learning, and ownership opportunities
