FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and managing datasets for large language models and vision models, with a strong focus on data processing, statistical analysis, and pipeline design. Proficient in high-performance Python programming and data-at-scale frameworks to optimize model performance and address multimodal challenges.
Highest-signal resume keywords
Expert-Level Python SkillsData Processing Pipelines DesignExperience With Petabyte-Scale DatasetsMastery Of Data-At-Scale FrameworksBuilding Datasets For RLHF And DPO
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Data CleaningFuzzy DeduplicationSignal ExtractionStatistical AnalysisSynthetic-Data GenerationHigh-Performance CodingEfficient Memory ManagementMultithreadingMultiprocessingDataset Management
Tools & Technologies
SparkRayWebDatasetParquet
Industry Keywords
Data ScienceMachine LearningLarge Language ModelsVision ModelsHuman-Labeling Workflows
Tech Stack
Tools & technologiesPythonRaySpark
About the role
Key responsibilities & impact- Own the end-to-end creation of pre-training datasets for LLMs
- Define the mix of web data, code, books, and technical papers to optimize downstream model performance
- Design and implement pipelines for data cleaning, exact and fuzzy deduplication, and signal extraction from petabytes of raw unstructured data
- Lead development of post-training datasets, including SFT instructions, multi-turn dialogues, and RLHF/DPO preference modeling data
- Drive acquisition and processing of vision and video data
- Address multimodal alignment, video compression, and temporal data consistency
- Develop high-throughput Python data-processing scripts using multiprocessing and multithreading
- Conduct statistical analysis of training corpora to identify biases, knowledge gaps, and quality regressions
- Design synthetic-data generation pipelines to augment gaps in natural datasets
Requirements
What you’ll need- 5+ years of industry experience in Data Science or Machine Learning
- Proven track record of building and managing datasets for foundation models
- Expert-level Python skills
- High-performance coding with multiprocessing and multithreading
- Efficient memory management for large-scale data tasks
- Experience working with petabyte-scale datasets directly used to train production-grade LLMs or Large Vision Models
- Experience building massive LLM training sets from scratch, including raw web crawls such as Common Crawl and specialized domain data
- Hands-on experience building datasets for RLHF, DPO, and multi-turn instruction following
- Experience managing human-labeling workflows and quality gold-sets
- Mastery of data-at-scale frameworks such as Spark or Ray
- Experience with high-performance data-loading formats such as WebDataset or Parquet
- Must be legally authorized to work in the United States
- A Master’s or PhD in a quantitative field focused on data-centric AI or information retrieval is a nice-to-have, not required
Benefits
Comp & perks- Medical, dental, and vision insurance
- 401k plan
- Daily lunch, snacks, and beverages
- Flexible time off
- Competitive salary
- Equity
