Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
24-MAG

Member of Technical Staff, Frontier AI

24-MAG

. Own research and evaluation initiatives from problem framing through data design, quality calibration, and signal validation .

Posted 9/15/2026full-timeRemote • New York • United StatesLead💰 $600,000 - $2,000,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in designing ML-oriented datasets and evaluation frameworks while maintaining high research signal quality. Capable of translating complex real-world behaviors into structured research opportunities and communicating findings effectively to diverse stakeholders.

Highest-signal resume keywords
ML-Oriented Dataset DesignEvaluation Framework DevelopmentQuality Assurance ProcessesSystems-Level Understanding of AI PerformanceTechnical Communication Skills

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Data DesignSignal ValidationModel EvaluationAnnotation SystemsQuality CalibrationExperimental AnalysisReinforcement LearningFeedback-Driven TrainingResearch Signal Quality AssessmentEvaluation of Complex Tasks
Soft Skills
Strong Professional JudgementOwnership MindsetDecision-Making in UncertaintyCollaboration with Cross-Functional TeamsExcellent Written and Verbal Communication
Industry Keywords
Applied ResearchTechnical Research ProgrammesAI SystemsAgentic SystemsReal-World Workflows

About the role

Key responsibilities & impact
  • Own research and evaluation initiatives from problem framing through data design, quality calibration, and signal validation
  • Define rigorous approaches for determining whether experimental results provide reliable and defensible research signal
  • Analyse model and system failures to identify root causes, edge cases, and opportunities for improvement
  • Evaluate datasets, experiments, and conclusions against appropriate quality thresholds
  • Act as a quality gate when signal strength, data integrity, or supporting evidence is insufficient
  • Design ML-oriented data systems, including task definitions, annotation schemas, rubrics, incentives, and supporting pipelines
  • Structure data and evaluation workflows around downstream model-performance objectives
  • Translate ambiguous real-world behaviour into measurable evaluation frameworks and new data categories
  • Identify evaluation or dataset coverage gaps and recommend additional investment or iteration
  • Develop quality-assurance processes that maintain consistent research standards
  • Investigate model and system behaviour to identify recurring weaknesses and performance limitations
  • Iterate on evaluations, datasets, feedback loops, and quality standards
  • Use experimental findings to guide improvements in model or agent performance
  • Determine when research directions should be expanded, revised, paused, or discontinued based on evidence
  • Maintain a systems-level perspective focused on end-to-end AI performance
  • Collaborate with researchers, domain experts, operators, and cross-functional teams during project kickoff, calibration, and iteration
  • Communicate research findings, trade-offs, limitations, and signal strength to technical and non-technical stakeholders
  • Translate research progress into evidence-grounded narratives
  • Support alignment between experimental work and real-world system requirements
  • Contribute technical judgement in ambiguous, high-impact research environments

Requirements

What you’ll need
  • Experienced technical professional
  • Strong professional judgement regarding research signal quality
  • Experience designing ML-oriented datasets, evaluation frameworks, annotation systems, rubrics, or QA processes
  • Ability to translate complex and ambiguous real-world system behaviour into structured research and evaluation opportunities
  • Strong ownership mindset and comfort making decisions in uncertain or rapidly evolving environments
  • Excellent written and verbal communication skills
  • Ability to explain technical trade-offs, limitations, evidence quality, and research findings clearly
  • Proven experience working directly with researchers, technical experts, or domain specialists during project calibration and iteration
  • Systems-level understanding of model, agent, or AI-system performance
  • Experience with reinforcement-learning environments, simulators, or feedback-driven training systems is advantageous
  • Experience improving agentic systems or AI systems operating within real-world workflows is beneficial
  • Prior work within applied research or production environments with direct impact on deployed systems is advantageous
  • Experience designing evaluations for complex or real-world tasks is strongly valued
  • Familiarity with expert incentive design or high-stakes technical research programmes is beneficial
  • Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party

Benefits

Comp & perks
  • Fully remote work
  • Full-time engagement
  • Remote consulting opportunity
  • Compensation of $600,000–$2,000,000/year