FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Support Engineer, GPU Infrastructure
Hydra Host. Diagnose Linux server issues end to end, including boot and network boot failures, kernel and driver problems, filesystems, storage pressure, services, memory and CPU behavior, and instability .
Posted 9/15/2026full-timeRemote • Florida • United StatesMid-LevelSenior💰 $95,000 - $130,000 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates strong expertise in diagnosing and troubleshooting Linux server issues, including hardware and network problems, while effectively managing incidents and communicating with technical customers. Proficient in scripting and contributing to infrastructure management and operational improvements.
Highest-signal resume keywords
Linux TroubleshootingOut-Of-Band ManagementServer Hardware ExperienceIncident Management SystemsScripting with Python or Bash
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Linux Server DiagnosisTCP/IP KnowledgeHardware Failure DiagnosisNVIDIA GPU TroubleshootingRAID ConfigurationFirmware ManagementNetwork UtilitiesSystemd LogsStorage ToolingFault-Domain Reasoning
Soft Skills
Clear Written CommunicationProblem-SolvingCollaboration
Tools & Technologies
IPMIRedfishIDRACILOPrometheusGrafanaAnsibleTerraformNetBoxDistributed Storage
Industry Keywords
Data Center InfrastructureBare Metal EnvironmentsCloud EnvironmentsHPCAI Training Environments
Tech Stack
Tools & technologiesAnsibleCloudDACDNSGrafanaLinuxPrometheusPythonTCP/IPTerraform
About the role
Key responsibilities & impact- Diagnose Linux server issues end to end, including boot and network boot failures, kernel and driver problems, filesystems, storage pressure, services, memory and CPU behavior, and instability
- Diagnose hardware failures using out-of-band management, sensor data, POST and boot errors, SMART data, and vendor diagnostics
- Isolate server-side network problems involving NICs, drivers, VLANs, addressing, routing, MTU, DNS, DHCP, bonding, link state, and packet captures
- Troubleshoot NVIDIA GPU servers, including GPU availability, thermal throttling, driver and VBIOS mismatch, PCIe, XID errors, and host-level conditions
- Separate hardware, OS, network, application, and configuration issues before escalation
- Own incidents through resolution or clean handoff, set severity by blast radius, and identify related tickets
- Escalate to engineering with evidence packs and participate in root cause analysis and post-incident reviews
- Coordinate with data center partners on remote hands, reboots, cabling, optics checks, component replacement, and physical inspection
- Open and track hardware RMAs with OEMs through replacement and validate repaired or replaced equipment
- Maintain accurate asset records and support server turn-ups, migrations, and decommissions
- Communicate clear cause and timeline updates to technically sophisticated customers
- Write and improve runbooks, operational procedures, categories, and closure reasons
- Build Python, Bash, or similar scripts and tools for health checks, data collection, and routine operations
- Contribute to infrastructure-as-code, configuration management, monitoring, alerting, and support workflow improvements
- Work a defined shift, participate in an escalation rotation, and provide written shift handoffs
Requirements
What you’ll need- Three or more years supporting production servers, data center infrastructure, or bare metal and cloud environments
- Strong hands-on Linux troubleshooting, including logs, dmesg, systemd, storage tooling, and network utilities
- Experience with server hardware, including CPU and memory, storage and filesystems, RAID, PCIe, NICs, power, BIOS and UEFI, firmware, and drivers
- Out-of-band management experience with IPMI, Redfish, iDRAC, iLO, or similar
- Working TCP/IP knowledge and ability to determine whether a problem is on the host or network
- Fault-domain reasoning and production judgment
- Experience with ticketing, monitoring, incident management, or infrastructure management systems
- Clear written English
- Helpful but not required: NVIDIA GPU servers at scale; CUDA, NCCL, NVLink, or DCGM; HPC or AI training environments; InfiniBand or high-performance Ethernet; Dell, HPE, Supermicro, or Lenovo platforms; NVMe, ZFS, Ceph, or distributed storage; Prometheus, Grafana, or similar observability tooling; NetBox or another infrastructure and asset register; Ansible, Terraform, or configuration management; Git-based infrastructure workflows; optics, transceivers, DAC or AOC cabling; geographically distributed third-party facilities
Benefits
Comp & perks- Defined shift schedule agreed before starting
- Escalation rotation for high-severity issues outside shift hours
- Written handoff at the end of every shift
- Opportunity to influence the structure of a newly built support function