Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Luma AI

Research Scientist / Engineer – Performance Optimization

Luma AI

. Profile and optimize GPU, CPU, and accelerator code for maximum utilization and minimal latency .

Posted 10/8/2026full-timeLondon • United KingdomMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expert-level proficiency in Triton and CUDA programming for GPU optimization, with a strong focus on high-performance PyTorch development and transformer model optimization. Capable of building performance monitoring tools and implementing advanced optimization techniques to enhance model architectures and deployment efficiency.

Highest-signal resume keywords
Expert-Level Triton ProgrammingCUDA OptimizationHigh-Performance PyTorch DevelopmentProfiling Tools ProficiencyTransformer Architecture Understanding

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Triton ProgrammingCUDA ProgrammingPyTorch DevelopmentKernel DevelopmentCustom OperationsPerformance OptimizationModel Architecture OptimizationDistributed DeploymentProfiling TechniquesAdvanced CUDA Optimization
Tools & Technologies
NVIDIA NsightTorch ProfilerTensorRTONNXXLA
Industry Keywords
GPU OptimizationTransformer ModelsAttention MechanismsKernel Fusion TechniquesInference Workloads

Tech Stack

Tools & technologies
Node.jsPyTorch

About the role

Key responsibilities & impact
  • Profile and optimize GPU, CPU, and accelerator code for maximum utilization and minimal latency
  • Write high-performance PyTorch, Triton, and CUDA code, including custom operations when needed
  • Develop fused kernels and leverage tensor cores and modern hardware features across platforms
  • Optimize model architectures and implementations for distributed multi-node production deployment
  • Build performance monitoring and analysis tools and automation
  • Research and implement cutting-edge optimization techniques for transformer models
  • Profile current training and inference paths during the first 30 days
  • Deliver a kernel or architecture optimization that measurably improves utilization or latency during days 30–60
  • Build monitoring and automation to prevent performance regressions during days 60–90

Requirements

What you’ll need
  • Expert-level Triton/CUDA programming and GPU optimization
  • Strong PyTorch skills, including kernel development and custom operations
  • Proficiency with profiling tools, including NVIDIA Nsight, torch profiler, and custom tooling
  • Deep understanding of transformer architectures and attention mechanisms
  • Experience with compilers and exporters such as torch.compile, TensorRT, ONNX, or XLA (nice to have)
  • Experience optimizing inference workloads for latency and throughput (nice to have)
  • Knowledge of Triton compiler and kernel fusion techniques (nice to have)
  • Knowledge of warp-level intrinsics and advanced CUDA optimization (nice to have)

Benefits

Comp & perks
  • Equal opportunity employer
  • Hybrid work arrangement
  • Remote work option in the UK (listed as a secondary location)