FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expert-level proficiency in Triton and CUDA programming for GPU optimization, with a strong focus on high-performance PyTorch development and transformer model optimization. Capable of building performance monitoring tools and implementing advanced optimization techniques to enhance model architectures and deployment efficiency.
Highest-signal resume keywords
Expert-Level Triton ProgrammingCUDA OptimizationHigh-Performance PyTorch DevelopmentProfiling Tools ProficiencyTransformer Architecture Understanding
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Triton ProgrammingCUDA ProgrammingPyTorch DevelopmentKernel DevelopmentCustom OperationsPerformance OptimizationModel Architecture OptimizationDistributed DeploymentProfiling TechniquesAdvanced CUDA Optimization
Tools & Technologies
NVIDIA NsightTorch ProfilerTensorRTONNXXLA
Industry Keywords
GPU OptimizationTransformer ModelsAttention MechanismsKernel Fusion TechniquesInference Workloads
Tech Stack
Tools & technologiesNode.jsPyTorch
About the role
Key responsibilities & impact- Profile and optimize GPU, CPU, and accelerator code for maximum utilization and minimal latency
- Write high-performance PyTorch, Triton, and CUDA, including custom operations when needed
- Develop fused kernels and leverage tensor cores and modern hardware features across platforms
- Optimize model architectures and implementations for distributed multi-node production deployment
- Build performance monitoring and analysis tools and automation
- Research and implement cutting-edge optimization techniques for transformer models
- Profile current training and inference paths during the first 30 days
- Ship and validate a kernel or architecture optimization during days 30–60
- Build monitoring and automation to prevent performance regressions during days 60–90
Requirements
What you’ll need- Expert-level Triton/CUDA programming and GPU optimization
- Strong PyTorch skills, including kernel development and custom operations
- Proficiency with profiling tools, including NVIDIA Nsight, torch profiler, and custom tooling
- Deep understanding of transformer architectures and attention mechanisms
- Experience with compilers and exporters such as torch.compile, TensorRT, ONNX, or XLA (nice to have)
- Experience optimizing inference workloads for latency and throughput (nice to have)
- Knowledge of Triton compiler and kernel fusion techniques (nice to have)
- Knowledge of warp-level intrinsics and advanced CUDA optimization (nice to have)
Benefits
Comp & perks- Equal opportunity employer
- Voluntary diversity and inclusion survey participation; refusal will not affect the job application
