FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in sustaining engineering and failure analysis for data center products, with a strong focus on root-cause analysis, corrective actions, and collaboration with OEMs and ODMs. Proficient in debugging complex server platforms and managing RMA processes while delivering clear communication and project management.
Highest-signal resume keywords
Sustaining EngineeringFailure AnalysisRoot-Cause AnalysisServer ArchitectureNVIDIA GPU Support
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Failure AnalysisRoot-Cause AnalysisDebuggingSystem FirmwareThermal MeasurementsStatistical AnalysisHardware Bring-UpBoard ValidationDiagnostic AutomationField Application Engineering
Soft Skills
Project ManagementCommunicationTask Prioritization
Tools & Technologies
NVIDIA HGXNVIDIA DGXNVIDIA MGXPCIeBMC/IPMIRedfishBIOS/UEFILinuxPythonShell Scripting
Industry Keywords
OEMODMCloud Service ProviderHyperscale CustomerReliability Engineering
Tech Stack
Tools & technologiesCloudLinuxPythonShell Scripting
About the role
Key responsibilities & impact- Serve as the technical lead for sustaining engineering related to NVIDIA data center products used by OEM customers
- Manage complex issues from initial triage and containment to root-cause analysis, corrective action, and resolution
- Perform system- and board-level failure analysis to isolate hardware, firmware, software, thermal, power, and integration-related problems
- Facilitate the complete RMA process, including failure-data collection, return authorization, material tracking, engineering disposition, and communication of findings
- Reproduce field failures in laboratory environments, analyze failure trends, and drive preventive improvements in product quality and serviceability
- Collaborate with OEMs, ODMs, manufacturing partners, and NVIDIA engineering, quality, reliability, operations, and supply-chain teams to resolve critical issues
- Provide on-site technical support and deliver clear failure-analysis reports, corrective-action updates, troubleshooting procedures, and executive-level communications
Requirements
What you’ll need- BS or MS in Electrical Engineering, Computer Engineering, Computer Science, or a related discipline—or equivalent experience
- 5+ years of relevant experience in Field Application Engineering, sustaining engineering, failure analysis, systems engineering, or product engineering
- Hands-on experience debugging complex server or computing platforms at both the system and board levels
- Demonstrated experience managing field failures or RMA cases through root-cause analysis and corrective-action closure using methods such as 8D, FRACAS, or fault-tree analysis
- Strong understanding of server architecture, including CPUs, GPUs, memory, storage, power delivery, cooling, and high-speed interconnects
- Knowledge of PCIe and experience troubleshooting BMC/IPMI or Redfish, BIOS/UEFI, Linux, drivers, firmware, system logs, and hardware telemetry
- Experience with hardware bring-up, board validation, system qualification, and L1–L10 server integration and manufacturing test stages
- Ability to interpret schematics, block diagrams, diagnostic data, manufacturing test results, and thermal or electrical measurements
- Excellent project-management and communication skills
- Ability to prioritize tasks and work effectively with business and engineering teams
- Must operate from the Round Rock area and travel as required
- Direct experience supporting NVIDIA HGX, DGX, or MGX platforms or comparable GPU-accelerated data center systems
- Experience owning sustaining engineering, failure-analysis, or RMA programs for an OEM, ODM, cloud service provider, or hyperscale customer
- Familiarity with physical failure-analysis techniques, reliability engineering, statistical analysis, and manufacturing or field-quality trend analysis
- Experience with ARM-based systems, NVIDIA GPUs, system firmware, and diagnostic automation using Python or shell scripting
Benefits
Comp & perks- Equity
- Benefits 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score
