FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Tech Stack
Tools & technologiesAssemblyDistributed SystemsPythonPyTorchC++
About the role
Key responsibilities & impact- Design, implement, and optimize custom GPU kernels using C++, PTX, CUDA, ROCm, Triton, and/or JAX Pallas
- Profile and optimize end-to-end performance of ML operations, focusing on large-scale LLM training and inference
- Integrate low-level GPU kernels into PyTorch, JAX, and custom internal runtimes
- Develop performance models, identify bottlenecks, and deliver kernel-level improvements that accelerate AI workloads
- Collaborate with ML researchers, distributed systems engineers, and model-serving teams to optimize compute performance across the stack
- Work with NVIDIA and AMD hardware vendors and stay current on GPU architecture and compiler/toolchain improvements
- Contribute to tooling, documentation, benchmarking suites, and testing frameworks for correctness and reproducibility
- Build full-stack infrastructure powering frontier AI models and real-time applications at Sciforium, an AI infrastructure company
Requirements
What you’ll need- 5+ years of industry or research experience in GPU kernel development or high-performance computing
- Bachelor’s, Master’s, or PhD in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or a related field
- Strong programming skills in C++, Python, and familiarity with ML frameworks
- Deep expertise in CUDA/ROCm, GPU memory models, and performance optimization strategies
- Hands-on experience with Triton and/or JAX Pallas for custom kernel development
- Strong understanding of PTX, GPU ASM, and low-level GPU execution
- Extensive experience writing and optimizing custom GPU kernels in C++ and PTX
- Proven ability to integrate low-level kernels into PyTorch, JAX, or similar frameworks
- Experience with large-scale LLM workloads, training or inference
- Experience with AMD GPUs and ROCm optimization
- Familiarity with JAX FFI and custom ML operator development
- Experience with efficient model serving frameworks such as vLLM or TensorRT
- Experience with TPUs, XLA, or similar accelerator programming environments
- Contributions to open-source ML systems, compilers, or GPU kernels
- Legally authorized to work in the United States
- Work visa sponsorship requirements must be disclosed
Benefits
Comp & perks- Medical, dental, and vision insurance
- 401k plan
- Daily lunch, snacks, and beverages
- Flexible time off
- Competitive salary
- Equity
