Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Pika

Research Scientist, Data

Pika

. Architect and implement large-scale data pipelines supporting model training and research workflows for text, image, audio, and video datasets .

Posted 9/24/2026full-timePalo Alto • California • United StatesMid-LevelSenior💰 $185,000 - $400,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in architecting and implementing large-scale data pipelines for machine learning applications, with a strong focus on data quality, compliance, and the integration of multimodal models. Proficient in developing scalable data ingestion and management tools while ensuring ethical considerations throughout the data lifecycle.

Highest-signal resume keywords
Data Pipeline ArchitectureMachine Learning Data CurationDistributed Data SystemsPython ProgrammingCloud Data Platforms

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Data Pipeline ArchitectureMachine Learning Data CurationDistributed Data SystemsPython ProgrammingSQLPySparkETL WorkflowsData Quality AssuranceDataset ManagementData Ingestion Tools
Soft Skills
Cross-Functional CollaborationProblem-SolvingCommunication
Tools & Technologies
SparkHadoopRayAWSGCPAzure
Industry Keywords
Data ComplianceData EthicsGenerative AIMultimodal ModelsData Lifecycle Management

Tech Stack

Tools & technologies
AWSAzureCloudETLGoogle Cloud PlatformHadoopPySparkPythonRaySparkSQL

About the role

Key responsibilities & impact
  • Architect and implement large-scale data pipelines supporting model training and research workflows for text, image, audio, and video datasets
  • Partner with research and engineering teams to curate, clean, and manage diverse sensory-rich datasets for pre-training and mid-training of multimodal models
  • Develop strategies and tools for scalable data ingestion, labeling, filtering, augmentation, and storage
  • Ensure data quality, reliability, and compliance, including privacy and ethical considerations throughout the data lifecycle
  • Optimize data processing, transformation, and delivery for large-scale distributed training pipelines
  • Prototype and productionize methods for dataset creation, management, and continuous improvement
  • Integrate research-driven data advancements into production-ready systems
  • Stay informed about emerging data engineering and ML data management developments and apply best practices

Requirements

What you’ll need
  • 5+ years of experience building and scaling data pipelines for machine learning applications at staff or lead engineer level, ideally in research or model training environments
  • Strong background in data engineering and ML data curation for LLMs, VLMs, or other large-scale multimodal models
  • Expertise in distributed data systems, such as Spark, Hadoop, Ray, or similar
  • Experience with efficient large dataset processing and ETL workflows
  • Proven ability to build robust, scalable, and production-grade data infrastructure for ML pipelines
  • Experience developing tools for data labeling, filtering, deduplication, quality assurance, and dataset management
  • Strong programming skills in Python, SQL, PySpark, or similar
  • Familiarity with cloud data platforms such as AWS, GCP, or Azure
  • Knowledge of privacy, compliance, ethics, and best practices in data collection and management
  • Excellent cross-functional collaboration, problem-solving, and communication skills
  • Passion for enabling cutting-edge generative AI and creative technology through data excellence
  • Ability to regularly work from the Palo Alto HQ
  • Authorization to work for any employer in the United States or ability to address US employment sponsorship requirements

Benefits

Comp & perks
  • Competitive salary and substantial equity in a high-growth startup
  • Full health benefits
  • 401k matching
  • Collaborative, mission-driven team environment with major growth opportunities
  • Flexible on-site/remote hybrid work arrangement