FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in AI frameworks and communication library integration, with a strong focus on kernel authoring and performance optimization on NVIDIA platforms. Proficient in rapid prototyping and development using Python, C++, and CUDA, with a solid understanding of large-scale AI workloads and models.
Highest-signal resume keywords
Kernel AuthoringCUDA OptimizationAI FrameworksPython DevelopmentMulti-GPU Communication
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
CUDAC++PythonTritonCuTePyTorchJAXTRT-LLMVLLMSGLang
Soft Skills
AdaptabilityEffective CommunicationCollaboration
Tools & Technologies
NVIDIA PlatformsAI ModelsPerformance Analysis
Industry Keywords
AI WorkloadsLarge-Scale TrainingProduction InferenceDeep Learning Communication PatternsExpert Parallelism
Tech Stack
Tools & technologiesPythonPyTorchC++
About the role
Key responsibilities & impact- Integrate new communication library features into AI frameworks from proof of concept through performance analysis to production
- Analyze AI workloads and frameworks to identify multi-GPU communication requirements and opportunities
- Collaborate hands-on with teams working on the latest AI models
- Author custom communication or fused compute-communication kernels for maximum performance on NVIDIA platforms
- Conduct in-depth research to achieve SOL GPU performance
- Build fault-tolerant and elastic solutions for large-scale or dynamic AI workloads
- Collaborate with a dynamic team across multiple time zones
Requirements
What you’ll need- Pursuing an M.S. or Ph.D. in CE/CS/EE
- Strong background in communication, kernel authoring, and/or AI training/inference
- Rapid prototyping and development with Python, C++, CUDA, or related DSLs such as Triton and cuTe
- Solid understanding of LLM models and parallelisms
- Adaptability and passion to learn new areas and tools
- Flexibility to work and communicate effectively
- Development experience with PyTorch, JAX, TRT-LLM, vLLM, SGLang, or veRL is advantageous
- Experience with DL communication patterns such as Expert Parallelism (EP), TP, DP, and PP is advantageous
- Experience with CUDA kernel optimization and profiling is advantageous
- Experience with large-scale training or production inference stack is advantageous
Benefits
Comp & perks- Highly competitive salaries
- Comprehensive benefits package
- Benefits for employees and their families
