Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Sciforium

GPU Kernel Engineer

Sciforium

. Design, implement, and optimize custom GPU kernels using C++, PTX, CUDA, ROCm, Triton, and/or JAX Pallas .

Posted 9/26/2026full-timeSan Francisco • California • United StatesMid-LevelSenior💰 $190,000 - $250,000 per yearWebsite

Tech Stack

Tools & technologies
AssemblyDistributed SystemsPythonPyTorchC++

About the role

Key responsibilities & impact
  • Design, implement, and optimize custom GPU kernels using C++, PTX, CUDA, ROCm, Triton, and/or JAX Pallas
  • Profile and optimize end-to-end performance of ML operations, focusing on large-scale LLM training and inference
  • Integrate low-level GPU kernels into PyTorch, JAX, and custom internal runtimes
  • Develop performance models, identify bottlenecks, and deliver kernel-level improvements that accelerate AI workloads
  • Collaborate with ML researchers, distributed systems engineers, and model-serving teams to optimize compute performance across the stack
  • Work with NVIDIA and AMD hardware vendors and stay current on GPU architecture and compiler/toolchain improvements
  • Contribute to tooling, documentation, benchmarking suites, and testing frameworks for correctness and reproducibility
  • Build full-stack infrastructure powering frontier AI models and real-time applications at Sciforium, an AI infrastructure company

Requirements

What you’ll need
  • 5+ years of industry or research experience in GPU kernel development or high-performance computing
  • Bachelor’s, Master’s, or PhD in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or a related field
  • Strong programming skills in C++, Python, and familiarity with ML frameworks
  • Deep expertise in CUDA/ROCm, GPU memory models, and performance optimization strategies
  • Hands-on experience with Triton and/or JAX Pallas for custom kernel development
  • Strong understanding of PTX, GPU ASM, and low-level GPU execution
  • Extensive experience writing and optimizing custom GPU kernels in C++ and PTX
  • Proven ability to integrate low-level kernels into PyTorch, JAX, or similar frameworks
  • Experience with large-scale LLM workloads, training or inference
  • Experience with AMD GPUs and ROCm optimization
  • Familiarity with JAX FFI and custom ML operator development
  • Experience with efficient model serving frameworks such as vLLM or TensorRT
  • Experience with TPUs, XLA, or similar accelerator programming environments
  • Contributions to open-source ML systems, compilers, or GPU kernels
  • Legally authorized to work in the United States
  • Work visa sponsorship requirements must be disclosed

Benefits

Comp & perks
  • Medical, dental, and vision insurance
  • 401k plan
  • Daily lunch, snacks, and beverages
  • Flexible time off
  • Competitive salary
  • Equity