FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in optimizing machine learning systems and performance engineering, with a strong focus on distributed training and GPU cluster management. Proficient in developing benchmarks and profiling tools to enhance AI and data-generation platforms.
Highest-signal resume keywords
Machine Learning Systems OptimizationCUDA and Triton ProficiencyDistributed Training at Multi-Node ScaleDeep Knowledge of PyTorch InternalsPerformance Engineering
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
CUDATritonPythonC++Distributed TrainingProfiling and BenchmarkingMemory Bandwidth OptimizationMixed Precision TrainingTensor ParallelismPipeline Parallelism
Soft Skills
Technical Direction SettingMentoring EngineersInfluencing Research Teams
Tools & Technologies
GCPNsightTorch.compileTensorRTXLAFSDPDeepSpeed
Industry Keywords
High-Performance ComputingGenerative ModelsTransformersGeometric Deep LearningProximity TherapeuticsProtein-Interaction Discovery
Tech Stack
Tools & technologiesGoogle Cloud PlatformNode.jsPythonPyTorchC++
About the role
Key responsibilities & impact- Profile and optimize training and inference for structural and generative models, including transformers, diffusion, and geometric deep learning
- Write and tune custom kernels using CUDA and Triton
- Use compilers such as torch.compile, TensorRT, and XLA when beneficial
- Scale distributed training across 32–64 nodes using FSDP, DeepSpeed, tensor parallelism, pipeline parallelism, and mixed precision
- Reduce inference cost by optimizing memory scaling for large complexes, improving diffusion sampling efficiency, batching ragged inputs, and maximizing throughput across up to 1000 GPUs
- Manage GPU cluster efficiency on GCP, focusing on scheduling, utilization, spot strategy, and cost reporting
- Develop benchmarks and profiling tools for the research team
- Help develop Proxima’s AI and data-generation platform for proximity therapeutics and protein-interaction discovery
Requirements
What you’ll need- Minimum of 6+ years experience in ML systems, HPC, or performance engineering
- BS/MS/PhD in CS, EE, or related field
- Ability to set technical direction beyond coding, select infrastructure, influence research teams, and mentor engineers
- Deep knowledge of PyTorch internals with hands-on experience profiling and fixing real bottlenecks
- Experience with CUDA and Triton
- Skilled at reading Nsight output
- Strong understanding of memory bandwidth and occupancy
- Experience with distributed training at multi-node scale
- Strong proficiency in Python and C++
- Able to name a model they made materially faster and quantify the improvement
