FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and optimizing infrastructure systems for machine learning workloads, with a strong focus on building scalable data management systems and high-performance inference systems. Proficient in programming with Python and Java, and experienced in deploying large-scale production machine learning systems.
Highest-signal resume keywords
Machine Learning Infrastructure DesignPython ProgrammingJava ProgrammingBig Data Processing FrameworksDistributed Systems
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Machine Learning SystemsInfrastructure OptimizationFeature Generation PipelinesHigh-Performance Inference SystemsData Management SystemsScalable Model TrainingCode Correctness StandardsSystem Performance OptimizationAI Model ServingProduction-Ready Quality
Soft Skills
Problem-Solving Skills
Tools & Technologies
SparkFlinkRayTensorFlowPyTorchCaffe2Spark MLScikit-Learn
Certifications & Qualifications
Bachelor’s Degree in Technical FieldMaster’s Degree in Technical FieldPhD in Relevant Technical Field
Industry Keywords
Machine LearningDistributed SystemsBig Data ProcessingAI Model InferenceCloud Infrastructure
Tech Stack
Tools & technologiesCloudDistributed SystemsJavaPythonPyTorchRayScikit-LearnSparkTensorflow
About the role
Key responsibilities & impact- Design and optimize infrastructure systems for machine learning workloads at scale
- Drive reliability and efficiency improvements across Snapchat’s ML Infrastructure
- Build and enhance feature generation and serving pipelines for online inferencing and offline training data generation
- Develop high-performance inference systems for fast and efficient AI model serving
- Build cloud infrastructure for scalable ML model training, evaluation, and inference
- Build data management systems for scalable data collection, labeling, processing, and evaluation
- Work closely with ML engineers to deploy models into production
- Use AI tools and high-velocity engineering workflows to design and ship scalable services
- Uphold standards for code correctness, security, and production-ready quality
Requirements
What you’ll need- Bachelor’s degree in a technical field such as computer science or equivalent experience
- 2+ years of post-Bachelor’s software development experience; or Master’s degree in a technical field + 1+ year of post-grad software development experience; or PhD in a relevant technical field
- Experience building large scale production machine learning systems, distributed systems or big data processing
- Strong programming skills in Python and Java
- Strong problem-solving skills focused on system performance, scalability, and efficiency
- Good understanding of distributed systems and infrastructure components of large-scale ML
- Experience with big data processing frameworks such as Spark, Flink, or Ray
- Proven track record of operating highly-available systems at significant scale
- Preferred: Master’s/PhD in a technical field such as computer science or equivalent industry experience
- Preferred: Experience working with ML Training platforms or optimizing AI model inference
- Preferred: Familiarity with ML frameworks such as TensorFlow, PyTorch, Caffe2, Spark ML, scikit-learn, or related frameworks
Benefits
Comp & perks- Paid parental leave
- Comprehensive medical coverage
- Emotional and mental health support programs
- Compensation packages including equity in the form of RSUs
