FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in analyzing and optimizing AI compute systems, with a strong focus on performance testing, debugging, and collaboration across hardware and software teams. Proficient in utilizing GPU computing technologies and performance tools to enhance workload efficiency.
Highest-signal resume keywords
AI Workload EvaluationLinux System DebuggingGPU Computing TechnologiesPerformance Testing AutomationDistributed Workload Management
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
PythonCC++GoBashCUDAROCmNCCLNsightKubernetes
Soft Skills
Strong OwnershipDocumentation SkillsCross-Team Collaboration
Tools & Technologies
DCGMRocprofilerSlurmNVIDIA NVLinkNVSwitchAMD Infinity FabricXGMI
Industry Keywords
AI Compute SystemsHPC WorkloadsPerformance BottlenecksTelemetryWorkload Scaling
Tech Stack
Tools & technologiesKubernetesLinuxPythonC++Go
About the role
Key responsibilities & impact- Analyze and improve the performance of AI compute systems across hardware, system software, and application workloads
- Run and analyze AI training, inference, and HPC workloads on GPU-accelerated systems
- Investigate compute, memory, PCIe, storage, and workload scaling issues
- Use telemetry, logs, traces, and profiling tools to identify performance bottlenecks
- Develop automation for performance testing, regression detection, and failure triage
- Work with hardware and software teams to resolve system and workload performance issues
- Document findings, test procedures, and performance results
- Support Cisco Validated Infrastructure Services, which establishes and proves the performance of large AI clusters before customer handoff
Requirements
What you’ll need- Bachelor’s degree with 5+ years of related experience, or Master’s degree with 3+ years of related experience, or PhD with relevant research experience, or equivalent related work experience
- Experience debugging, testing, or tuning Linux-based systems
- Experience running or evaluating AI, machine learning, or HPC workloads
- 3+ years of experience with a programming or scripting language such as Python, C, C++, Go, or Bash
- Familiarity with GPU computing technologies such as CUDA, ROCm, NCCL, or RCCL
- Familiarity with performance tools such as Nsight, DCGM, rocprofiler, or similar tools
- Understanding of GPU system components such as memory, NUMA, PCIe, and scale-up interconnects
- Exposure to NVIDIA NVLink and NVSwitch or AMD Infinity Fabric/xGMI
- Experience with Kubernetes, Slurm, or other distributed workload environments
- Strong ownership, documentation, and cross-team collaboration skills
Benefits
Comp & perks- Medical, dental and vision insurance
- 401(k) plan with a Cisco matching contribution
- Paid parental leave
- Short- and long-term disability coverage
- Basic life insurance
- Cisco restricted stock unit grants may be available
- 10 paid holidays per full calendar year
- 1 floating holiday for non-exempt employees
- Paid employee birthday day off
- Paid year-end holiday shutdown
- 4 paid personal wellness days
- 16 days of paid vacation per full calendar year for non-exempt employees
- Flexible vacation time off program with no defined limit for eligible exempt employees
- 80 hours of sick time off provided on hire date and each January 1st thereafter
- Up to 80 hours of unused sick time carried forward annually
- Additional paid time away for critical or emergency family issues
- Optional 10 paid volunteer days per full calendar year
- Annual bonuses for non-sales roles, subject to Cisco policies
