FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and optimizing infrastructure systems for machine learning workloads, with a strong focus on building scalable data management systems and high-performance inference systems. Proficient in collaborating with ML engineers to deploy advanced models and improve system performance and efficiency.
Highest-signal resume keywords
Machine Learning Infrastructure DesignLarge-Scale Production SystemsPython ProgrammingBig Data Processing FrameworksDistributed Systems Understanding
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Machine Learning SystemsInfrastructure OptimizationHigh-Performance Inference SystemsData Management SystemsVector Search AlgorithmsPythonJavaScalaC++Problem-Solving
Soft Skills
CollaborationProactive Learning
Tools & Technologies
SparkFlinkRayTensorFlowPyTorchCaffe2Spark MLScikit-learn
Industry Keywords
Machine LearningDistributed SystemsBig Data ProcessingAI Model ServingScalability
Tech Stack
Tools & technologiesCloudDistributed SystemsJavaPythonPyTorchRayScalaScikit-LearnSparkTensorflowC++
About the role
Key responsibilities & impact- Design and optimize infrastructure systems for machine learning workloads at scale
- Drive reliability and efficiency improvements across Snapchat’s ML Infrastructure
- Develop high-performance inference systems for fast and efficient AI model serving
- Build cloud infrastructure for scalable ML model training, evaluation, and inference
- Build comprehensive data management systems for scalable data collection, labeling, processing, and evaluation
- Work on state-of-the-art vector search algorithms to improve retrieval-system precision, recall, and scalability
- Collaborate with ML engineers to deploy cutting-edge models into production
Requirements
What you’ll need- Bachelor’s degree in a technical field such as computer science or equivalent experience
- 6+ years of post-Bachelor’s software development experience; or Master’s degree in a technical field plus 5+ years of post-graduate software development experience; or PhD in a relevant technical field plus 2+ years of post-graduate software development experience
- Experience building large-scale production machine learning systems, distributed systems, or big data processing
- Strong programming skills in Python, Java, Scala, or C++
- Strong problem-solving skills focused on system performance, scalability, and efficiency
- Understanding of distributed systems and large-scale ML infrastructure components
- Experience with big data processing frameworks such as Spark, Flink, or Ray
- Proven track record operating highly available systems at significant scale
- Ability to collaborate and work well with others
- Ability to proactively learn new concepts and apply them at work
- Preferred: Master’s/PhD in a technical field or equivalent industry experience
- Preferred: Experience with ML training platforms or optimizing AI model inference
- Preferred: Familiarity with TensorFlow, PyTorch, Caffe2, Spark ML, scikit-learn, or related frameworks
Benefits
Comp & perks- Paid parental leave
- Comprehensive medical coverage
- Emotional and mental health support programs
- Compensation packages including equity in the form of RSUs
- Default together approach with in-office work 4+ days per week