FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and optimizing infrastructure systems for machine learning workloads, with a strong focus on building scalable ML model training and inference systems. Proficient in programming with Python and Java, and experienced in deploying large-scale production machine learning systems.
Highest-signal resume keywords
Machine Learning Infrastructure DesignPython ProgrammingJava ProgrammingDistributed Systems ExperienceML Model Training and Inference Optimization
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Machine Learning SystemsDistributed SystemsBig Data ProcessingHigh-Performance Inference SystemsFeature Generation PipelinesData Management SystemsAI Model ServingCode Correctness StandardsSystem Performance OptimizationScalability and Efficiency
Soft Skills
Problem-SolvingCollaborationAdaptabilityProactive Learning
Tools & Technologies
SparkFlinkRayTensorFlowPyTorchCaffe2Spark MLScikit-Learn
Industry Keywords
Cloud InfrastructureAI ToolsProduction-Ready QualityMachine Learning WorkloadsData Collection and Processing
Tech Stack
Tools & technologiesCloudDistributed SystemsJavaPythonPyTorchRayScikit-LearnSparkTensorflow
About the role
Key responsibilities & impact- Design and optimize infrastructure systems for machine learning workloads at scale
- Drive reliability and efficiency improvements across Snapchat’s ML Infrastructure
- Build and enhance feature generation and serving pipelines for online inferencing and offline training data generation
- Develop high-performance inference systems for fast and efficient AI model serving
- Build infrastructure for scalable ML model training, evaluation, and inference in the cloud
- Build data management systems for scalable data collection, labeling, processing, and evaluation
- Work closely with ML engineers to deploy cutting-edge models into production
- Use AI tools and high-velocity engineering workflows to design and ship scalable services
- Uphold rigorous standards for code correctness, security, and production-ready quality
Requirements
What you’ll need- Bachelor’s degree in a technical field such as computer science or equivalent experience
- 2+ years of post-Bachelor’s software development experience; or Master’s degree in a technical field + 1+ year of post-grad software development experience; or PhD in a relevant technical field
- Experience building large scale production machine learning systems, distributed systems or big data processing
- Strong programming skills in Python and Java
- Strong problem-solving skills focused on system performance, scalability, and efficiency
- Good understanding of distributed systems and infrastructure components of large-scale ML
- Experience with Spark, Flink, or Ray
- Proven track record of operating highly-available systems at significant scale
- Experience working with ML training platforms or optimizing AI model inference
- Familiarity with TensorFlow, PyTorch, Caffe2, Spark ML, scikit-learn, or related frameworks
- Ability to collaborate and work well with others
- Ability to proactively learn new concepts and apply them at work
- Adaptability in learning and applying evolving AI systems and tools
Benefits
Comp & perks- Paid parental leave
- Comprehensive medical coverage
- Emotional and mental health support programs
- Compensation packages including equity in the form of RSUs
