Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Sciforium

Foundation Model Data Engineer

Sciforium

. Own the end-to-end creation of pre-training datasets for LLMs .

Posted 9/16/2026full-timeSan Francisco • California • United StatesMid-LevelSenior💰 $155,000 - $210,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in building and managing datasets for large language models and vision models, with a strong focus on data processing, statistical analysis, and pipeline design. Proficient in high-performance Python programming and data-at-scale frameworks to optimize model performance and address multimodal challenges.

Highest-signal resume keywords
Expert-Level Python SkillsData Processing Pipelines DesignExperience With Petabyte-Scale DatasetsMastery Of Data-At-Scale FrameworksBuilding Datasets For RLHF And DPO

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Data CleaningFuzzy DeduplicationSignal ExtractionStatistical AnalysisSynthetic-Data GenerationHigh-Performance CodingEfficient Memory ManagementMultithreadingMultiprocessingDataset Management
Tools & Technologies
SparkRayWebDatasetParquet
Industry Keywords
Data ScienceMachine LearningLarge Language ModelsVision ModelsHuman-Labeling Workflows

Tech Stack

Tools & technologies
PythonRaySpark

About the role

Key responsibilities & impact
  • Own the end-to-end creation of pre-training datasets for LLMs
  • Define the mix of web data, code, books, and technical papers to optimize downstream model performance
  • Design and implement pipelines for data cleaning, exact and fuzzy deduplication, and signal extraction from petabytes of raw unstructured data
  • Lead development of post-training datasets, including SFT instructions, multi-turn dialogues, and RLHF/DPO preference modeling data
  • Drive acquisition and processing of vision and video data
  • Address multimodal alignment, video compression, and temporal data consistency
  • Develop high-throughput Python data-processing scripts using multiprocessing and multithreading
  • Conduct statistical analysis of training corpora to identify biases, knowledge gaps, and quality regressions
  • Design synthetic-data generation pipelines to augment gaps in natural datasets

Requirements

What you’ll need
  • 5+ years of industry experience in Data Science or Machine Learning
  • Proven track record of building and managing datasets for foundation models
  • Expert-level Python skills
  • High-performance coding with multiprocessing and multithreading
  • Efficient memory management for large-scale data tasks
  • Experience working with petabyte-scale datasets directly used to train production-grade LLMs or Large Vision Models
  • Experience building massive LLM training sets from scratch, including raw web crawls such as Common Crawl and specialized domain data
  • Hands-on experience building datasets for RLHF, DPO, and multi-turn instruction following
  • Experience managing human-labeling workflows and quality gold-sets
  • Mastery of data-at-scale frameworks such as Spark or Ray
  • Experience with high-performance data-loading formats such as WebDataset or Parquet
  • Must be legally authorized to work in the United States
  • A Master’s or PhD in a quantitative field focused on data-centric AI or information retrieval is a nice-to-have, not required

Benefits

Comp & perks
  • Medical, dental, and vision insurance
  • 401k plan
  • Daily lunch, snacks, and beverages
  • Flexible time off
  • Competitive salary
  • Equity