Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Proxima

Principal ML Performance Engineer – GPU Optimization

Proxima

. Profile and optimize training and inference for structural and generative models, including transformers, diffusion, and geometric deep learning .

Posted 9/23/2026full-timeRemote • New York • United StatesLeadWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in optimizing machine learning systems and performance engineering, with a strong focus on distributed training and GPU cluster management. Proficient in developing benchmarks and profiling tools to enhance AI and data-generation platforms.

Highest-signal resume keywords
Machine Learning Systems OptimizationCUDA and Triton ProficiencyDistributed Training at Multi-Node ScaleDeep Knowledge of PyTorch InternalsPerformance Engineering

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
CUDATritonPythonC++Distributed TrainingProfiling and BenchmarkingMemory Bandwidth OptimizationMixed Precision TrainingTensor ParallelismPipeline Parallelism
Soft Skills
Technical Direction SettingMentoring EngineersInfluencing Research Teams
Tools & Technologies
GCPNsightTorch.compileTensorRTXLAFSDPDeepSpeed
Industry Keywords
High-Performance ComputingGenerative ModelsTransformersGeometric Deep LearningProximity TherapeuticsProtein-Interaction Discovery

Tech Stack

Tools & technologies
Google Cloud PlatformNode.jsPythonPyTorchC++

About the role

Key responsibilities & impact
  • Profile and optimize training and inference for structural and generative models, including transformers, diffusion, and geometric deep learning
  • Write and tune custom kernels using CUDA and Triton
  • Use compilers such as torch.compile, TensorRT, and XLA when beneficial
  • Scale distributed training across 32–64 nodes using FSDP, DeepSpeed, tensor parallelism, pipeline parallelism, and mixed precision
  • Reduce inference cost by optimizing memory scaling for large complexes, improving diffusion sampling efficiency, batching ragged inputs, and maximizing throughput across up to 1000 GPUs
  • Manage GPU cluster efficiency on GCP, focusing on scheduling, utilization, spot strategy, and cost reporting
  • Develop benchmarks and profiling tools for the research team
  • Help develop Proxima’s AI and data-generation platform for proximity therapeutics and protein-interaction discovery

Requirements

What you’ll need
  • Minimum of 6+ years experience in ML systems, HPC, or performance engineering
  • BS/MS/PhD in CS, EE, or related field
  • Ability to set technical direction beyond coding, select infrastructure, influence research teams, and mentor engineers
  • Deep knowledge of PyTorch internals with hands-on experience profiling and fixing real bottlenecks
  • Experience with CUDA and Triton
  • Skilled at reading Nsight output
  • Strong understanding of memory bandwidth and occupancy
  • Experience with distributed training at multi-node scale
  • Strong proficiency in Python and C++
  • Able to name a model they made materially faster and quantify the improvement