Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Snowflake

Senior Software Engineer – Cortex LLM Training Platform

Snowflake

. Design and build across the full stack, from public training APIs and SDK through the control plane to the GPU data plane .

Posted 9/17/2026full-timeBellevue • Washington • United StatesSenior💰 $200,000 - $287,500 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in building and scaling distributed systems, with a strong foundation in GPU and LLM infrastructure. Proficient in designing fault-tolerant services and optimizing performance for machine learning applications.

Highest-signal resume keywords
Distributed SystemsKubernetesGPU InfrastructureFault ToleranceMachine Learning Systems

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Production ML SystemsScalable Services DesignDebugging Across LayersReliability EngineeringCost Efficiency Optimization
Tools & Technologies
PyTorchDeepSpeed/FSDPRayCUDA/NCCLVLLM
Certifications & Qualifications
BS in Computer Science
Industry Keywords
Serverless GPU ComputeMulti-Tenant SchedulingCapacity-Aware RoutingEnterprise Scale ComponentsEnd-to-End Performance

Tech Stack

Tools & technologies
Distributed SystemsKubernetesPyTorchRay

About the role

Key responsibilities & impact
  • Design and build across the full stack, from public training APIs and SDK through the control plane to the GPU data plane
  • Scale distributed systems that make GPU compute serverless
  • Build multi-tenant scheduling, placement, and capacity-aware routing across regional GPU pools
  • Implement fault tolerance across distributed systems
  • Drive end-to-end performance for training, inference, and RL loops
  • Keep the data plane responsive under heavy concurrent load and keep GPUs saturated
  • Partner with Snowflake Research to productionize state-of-the-art training and inference techniques
  • Build reliable, composable components customers can run at enterprise scale

Requirements

What you’ll need
  • 3+ years of experience building and shipping production ML systems for the Intermediate level; 6+ years for the Senior level
  • Strong distributed systems and infrastructure foundation
  • Experience designing scalable, fault-tolerant services and operating them on Kubernetes in production
  • Familiarity with GPU and LLM infrastructure, including PyTorch, DeepSpeed/FSDP, Ray, CUDA/NCCL, and vLLM
  • Ability to debug across data, infrastructure, and GPU layers
  • Demonstrated ability to harden complex systems for reliability, throughput, and cost efficiency
  • BS in Computer Science or a related field
  • Hands-on LLM post-training/modeling experience is a bonus
  • Must be authorized to work in the country to which applying
  • Must reside within commuting distance of, or be open to relocation to, the designated local Snowflake office