FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

GPU AI Solution Architect
Intel Corporation. Design and optimize AI accelerator systems, including GPU clusters and Gaudi platforms, for production machine learning workloads .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and optimizing AI accelerator systems, with a strong focus on GPU clusters and Gaudi platforms. Proficient in debugging system-level issues and developing automated testing frameworks for AI systems.
Highest-signal resume keywords
AI Accelerator System DesignSystem EngineeringPlatform ValidationPython ProficiencyAI Frameworks (PyTorch, TensorFlow)
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
AI Cluster DesignDebugging High-Performance AI ClustersPCIe ConnectivityMemory SubsystemsLinux/Unix AdministrationDockerShell ScriptingFull-Stack DebuggingArchitectural ImprovementsTest Plans Development
Soft Skills
Mentoring Junior EngineersCross-Functional Collaboration
Tools & Technologies
Gaudi PlatformsIntel Platforms (Xeon)OpenMPIVLLMRedfishIPMIBMC Management Protocols
Industry Keywords
AI/ML Workload OptimizationEnterprise Platform Security
Tech Stack
Tools & technologiesDockerLinuxPythonPyTorchShell ScriptingTensorflowUnix
About the role
Key responsibilities & impact- Design and optimize AI accelerator systems, including GPU clusters and Gaudi platforms, for production machine learning workloads
- Debug system-level issues involving PCIe connectivity, memory subsystems, and interconnects in AI clusters
- Lead platform bring-up and validation for next-generation AI hardware
- Develop and execute comprehensive test plans and validation strategies for AI systems
- Collaborate with OEM vendors on firmware integration and system-level optimizations
- Perform full-stack debugging across hardware, firmware, and software layers
- Develop automated testing frameworks, diagnostic tools, and monitoring solutions for AI systems
- Mentor junior engineers and contribute to cross-functional collaborations
- Drive architectural improvements and technical decisions to enhance AI infrastructure capabilities
Requirements
What you’ll need- Bachelor's degree and 6+ years of experience, or Master's degree and 4+ years of experience, or PhD and 2+ years of experience in Computer Science, Electrical Engineering, or a related field
- 5+ years of experience in system engineering, platform validation, or related roles
- 3+ years of experience bringing up and debugging high-performance AI clusters
- 3+ years of experience resolving complex system-level issues in production AI/ML environments
- 3+ years of experience in AI cluster design, validation, and production deployment
- Solid understanding of PCIe, memory subsystems, and AI accelerators
- Experience with Intel platforms (Xeon, Gaudi) or other GPU/AI accelerators
- Familiarity with AI frameworks such as PyTorch, TensorFlow, OpenMPI, and vLLM
- Proficiency in Python
- Expertise in Linux/Unix administration, Docker, and shell scripting
- Knowledge of Redfish, IPMI, and BMC management protocols
- Strong grasp of computer architecture, AI/ML workload optimization, and enterprise platform security
Benefits
Comp & perks- Competitive pay
- Stock bonuses
- Health benefits
- Retirement benefits
- Vacation benefits
- Hybrid work model allowing employees to split time between working on-site and off-site