Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Firecrawl

Research Engineer – Benchmarks

Firecrawl

. Design and run head-to-head benchmarks of data providers against verified ground truth .

Posted 9/30/2026full-timeSan Francisco • California • United StatesMid-LevelSenior💰 $250,000 - $290,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in designing and executing benchmarks for data providers, with a strong focus on accuracy, coverage, and reproducibility. Proficient in Python and API integrations, with a solid understanding of test-set design and evaluation methodologies.

Highest-signal resume keywords
Machine LearningBenchmark DesignPython ProgrammingData EngineeringAPI Integration

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Machine LearningBenchmark DesignPython ProgrammingData EngineeringTest-Set DesignScoring Method DesignSamplingLabelingInter-Rater AgreementLLM Judges
Soft Skills
Clear Communication
Tools & Technologies
APIsBenchmark HarnessAutomated Testing
Industry Keywords
Data ProvidersPublic BenchmarksEvaluation WorkFrontier LabThird-Party Benchmark Organization

Tech Stack

Tools & technologies
Python

About the role

Key responsibilities & impact
  • Design and run head-to-head benchmarks of data providers against verified ground truth
  • Build and maintain test datasets and ground truth, preventing staleness and leakage
  • Own an automated, reproducible, versioned, and defensible benchmark harness and weekly release
  • Measure accuracy, coverage, freshness, speed, and cost
  • Work with marketing to publish public leaderboard pages
  • Feed benchmark results into Alexandria's provider-selection system
  • Expand benchmark categories within Alexandria and beyond
  • Publish rigorous automated public benchmarks weekly

Requirements

What you’ll need
  • 4+ years in ML, research engineering, or data engineering, with evaluation or benchmark work shipped
  • Shipped evals or benchmarks and ability to explain why results were trustworthy
  • Experience at a frontier lab, data company, third-party benchmark organization, or public benchmark project; OSS contributors welcome
  • Strong Python skills and API integrations
  • Comfortable with messy vendor APIs, rate limits, and inconsistent schemas
  • Knowledge of test-set and scoring-method design, including sampling, labeling, inter-rater agreement, and appropriate use of LLM judges
  • Heavy use of AI and ability to upgrade personal workflow
  • Ability to write findings clearly for engineers, marketers, and vendors
  • Must be authorized to work in the United States or Canada
  • Must be based in or willing to relocate to the Bay Area or Greater Toronto Area before starting
  • LinkedIn and GitHub profiles required in the application

Benefits

Comp & perks
  • Competitive equity
  • 15 days mandatory PTO; anything after 24 days with approval; holidays excluded
  • 12 weeks fully paid parental leave for all parents
  • $100 USD/month wellness stipend
  • Up to $1,000 USD/year learning and development expenses
  • Team offsites
  • 3 paid months sabbatical after 4 years
  • US-based full-time employees: medical, dental, and vision coverage; employer-paid basic life and AD&D, short-term disability, and long-term disability insurance; Teladoc; Rightway; Talkspace; Carrot fertility and family-building coverage; EAP counseling and legal/financial consultations; 401(k); HSA, FSA, commuter benefits; supplemental insurance and protection options through MetLife
  • Canada-based full-time employees: employer-paid extended health, dental, and vision coverage through Manulife; employer-paid life, AD&D, short-term disability, and long-term disability insurance; Dialogue Premium virtual care; Talkspace Elite; Carrot fertility and family-building coverage; group RRSP through Wealthsimple
  • SF-based employees: snacks, drinks, team lunches, ping pong, and loaner electric bike
  • Toronto-based employees: snacks, drinks, team lunches, PRESTO transit card, and station parking
  • Paid work trial, approximately 40 hours, at a contractor rate