FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and maintaining data pipelines for machine learning, with a strong foundation in distributed systems and large-scale data processing. Proficient in optimizing performance and reliability of machine learning infrastructure using tools like PyTorch, Ray, and workflow orchestration systems.
Highest-signal resume keywords
Data Pipeline DevelopmentMachine Learning SystemsPython ProgrammingWorkflow Orchestration (Airflow, Flyte)Distributed Systems (Ray, Spark)
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Data Pipeline DevelopmentMachine Learning SystemsPython ProgrammingDistributed SystemsData ProcessingDataset ValidationModel Training WorkflowsLarge DatasetsPerformance OptimizationReliability Improvement
Soft Skills
Problem-SolvingCollaboration
Tools & Technologies
PyTorchRayAirflowFlyteTensorFlowData LakesData WarehousesStreaming Systems
Industry Keywords
Machine LearningDistributed SystemsLarge-Scale Data ProcessingScalable InfrastructureResearch Publications
Tech Stack
Tools & technologiesAirflowDistributed SystemsPythonPyTorchRaySparkTensorflow
About the role
Key responsibilities & impact- Build and maintain data pipelines that generate training datasets for machine learning models and experimentation
- Contribute to infrastructure supporting distributed training workflows using tools such as PyTorch and Ray
- Work with workflow orchestration tools such as Airflow and Flyte to support multi-stage ML pipelines
- Improve reproducibility and reliability through dataset validation, monitoring, and testing
- Partner with ML engineers to support experimentation and model iteration
- Optimize performance and efficiency across data processing and training systems
- Contribute to the evolution of the offline ML platform architecture as it scales
Requirements
What you’ll need- Bachelor’s degree in Computer Science, Machine Learning, Systems, or a related field
- Strong foundation in machine learning systems, distributed systems, or large-scale data processing through research or projects
- Experience with Python and data-intensive workloads
- Familiarity with ML frameworks such as PyTorch and TensorFlow and/or distributed systems such as Ray and Spark
- Academic or applied experience with data pipelines, model training workflows, or large datasets
- Strong problem-solving skills and ability to translate research ideas into practical systems
- Interest in building scalable, reliable machine learning infrastructure
- English proficiency sufficient for professional verbal and written exchanges
- Nice to have: experience with workflow orchestration systems such as Airflow or Flyte
- Nice to have: exposure to large-scale data platforms, including data lakes, warehouses, and streaming systems
- Nice to have: publications or research in ML systems, distributed systems, or related areas
Benefits
Comp & perks- Equity awards
- Participation in company incentive plans, such as annual discretionary bonuses or sales commissions
- Comprehensive health, life, and disability insurance
- Commute subsidy
- Employee stock ownership
- Competitive retirement/pension plans
- Generous vacation and personal days
- New-parent leave and family-care programs
- Office food snacks
- Mental Health and Wellbeing programs and support
- Employee Resource Groups
- Global Employee Assistance Program
- Training and development programs
- Volunteering and donation matching program
