FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and scaling distributed systems, with a strong foundation in GPU and LLM infrastructure. Proficient in designing fault-tolerant services and optimizing performance for machine learning applications.
Highest-signal resume keywords
Distributed SystemsKubernetesGPU InfrastructureFault ToleranceMachine Learning Systems
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Production ML SystemsScalable Services DesignDebugging Across LayersReliability EngineeringCost Efficiency Optimization
Tools & Technologies
PyTorchDeepSpeed/FSDPRayCUDA/NCCLVLLM
Certifications & Qualifications
BS in Computer Science
Industry Keywords
Serverless GPU ComputeMulti-Tenant SchedulingCapacity-Aware RoutingEnterprise Scale ComponentsEnd-to-End Performance
Tech Stack
Tools & technologiesDistributed SystemsKubernetesPyTorchRay
About the role
Key responsibilities & impact- Design and build across the full stack, from public training APIs and SDK through the control plane to the GPU data plane
- Scale distributed systems that make GPU compute serverless
- Build multi-tenant scheduling, placement, and capacity-aware routing across regional GPU pools
- Implement fault tolerance across distributed systems
- Drive end-to-end performance for training, inference, and RL loops
- Keep the data plane responsive under heavy concurrent load and keep GPUs saturated
- Partner with Snowflake Research to productionize state-of-the-art training and inference techniques
- Build reliable, composable components customers can run at enterprise scale
Requirements
What you’ll need- 3+ years of experience building and shipping production ML systems for the Intermediate level; 6+ years for the Senior level
- Strong distributed systems and infrastructure foundation
- Experience designing scalable, fault-tolerant services and operating them on Kubernetes in production
- Familiarity with GPU and LLM infrastructure, including PyTorch, DeepSpeed/FSDP, Ray, CUDA/NCCL, and vLLM
- Ability to debug across data, infrastructure, and GPU layers
- Demonstrated ability to harden complex systems for reliability, throughput, and cost efficiency
- BS in Computer Science or a related field
- Hands-on LLM post-training/modeling experience is a bonus
- Must be authorized to work in the country to which applying
- Must reside within commuting distance of, or be open to relocation to, the designated local Snowflake office
