Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
InnoData

Research Scientist, Video & Multimodal

InnoData

. Partner directly with customers and frontier labs building video understanding, video-language, and video-generation models .

Posted 9/24/2026full-timeRemote • United StatesMid-LevelSenior💰 $160,000 - $185,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in video understanding and multimodal machine learning, with a strong foundation in PyTorch and experience in building evaluation methodologies for video models. Proficient in data specification, annotation, and the use of advanced tools for video data processing and evaluation.

Highest-signal resume keywords
Video UnderstandingMultimodal Machine LearningPyTorch FundamentalsHuggingFace TransformersEvaluation Methodologies

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Video Data StructuringTemporal GroundingLong-Context ReasoningVideo Generation EvaluationAnnotation Schema DesignExperimentation and AblationData SpecificationFine-Tuning ModelsStreaming Video ProcessingSynthetic Video Generation
Soft Skills
Clear CommunicationCollaboration with Research ScientistsRigorous Documentation
Tools & Technologies
FFmpegDecordWebDatasetParquetArrowHuggingFace Datasets
Industry Keywords
Computer ScienceElectrical EngineeringVideo-Language ModelsCOCO-Style AnnotationResponsible AI EvaluationFirst-Author Publications

Tech Stack

Tools & technologies
FFmpegPyTorch

About the role

Key responsibilities & impact
  • Partner directly with customers and frontier labs building video understanding, video-language, and video-generation models
  • Define how video data is designed, structured, and evaluated for video and multimodal models
  • Translate model requirements into data specifications covering modalities, annotation schemas, sampling, and evaluation criteria
  • Build evaluation methodologies for video understanding, including temporal grounding, long-context reasoning, multi-turn, cross-modal, and retrieval-and-grounding evaluation
  • Build evaluation methodologies for video generation, including fidelity, temporal coherence, and physical plausibility
  • Decide how video is structured, enriched, and sampled, including messy, domain-specific footage
  • Run experiments and ablations connecting data choices to measurable model improvements
  • Design adversarial and stumping evaluations to identify system failures and improve data
  • Publish benchmarks, methodology, and papers
  • Work with annotation teams, subject-matter experts, and synthetic-data pipelines on collection and labeling plans

Requirements

What you’ll need
  • Roughly 5+ years of hands-on industry experience in video understanding or multimodal ML
  • A Bachelor's degree in computer science, electrical engineering, or a related technical or quantitative field is required
  • An advanced degree (MS or PhD) in a relevant field is preferred
  • Trained and evaluated video or multimodal models, with strong PyTorch fundamentals
  • Fluency in ffmpeg and decord pipelines, temporal and COCO-style annotation, WebDataset, Parquet and Arrow, and HuggingFace datasets
  • Experience fine-tuning large video or vision-language models with HuggingFace transformers, PEFT, and efficient inference
  • Experience with long-form video, streaming, temporal segmentation, or synthetic video generation
  • Experience building evaluation sets, calibrating difficulty, and defining quality video data for objectives
  • First-author publications or strong open-source contributions at venues such as CVPR, ICCV, ECCV, NeurIPS, or ICLR
  • Ability to work with research scientists and explain data and modeling decisions clearly to expert and non-expert audiences
  • Rigorous, reproducible approach to experiments and documentation
  • Interest or hands-on experience in responsible-AI evaluation and red-teaming is a bonus