Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
24-MAG

Member of Technical Staff, Coding Research

24-MAG

. Design and own evaluation frameworks for advanced coding agents .

Posted 9/15/2026full-timeRemote • New York • United StatesLead💰 $600,000 - $1,300,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in designing evaluation frameworks and methodologies for coding agents, with a strong focus on performance measurement, data generation, and technical assessment. Proficient in Python and C++, with a solid background in software engineering and AI research.

Highest-signal resume keywords
Python ProgrammingC++ ProgrammingMachine LearningBenchmark DesignAnalytical Skills

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Software EngineeringEvaluation MethodologiesData GenerationTechnical AssessmentsReinforcement LearningModel EvaluationCoding TasksExperiment DesignAutomationTooling Development
Soft Skills
Excellent Communication SkillsAnalytical ThinkingAdaptabilityCollaboration
Tools & Technologies
AI SystemsLarge Language ModelsEvaluation PipelinesResearch Documentation
Industry Keywords
Coding AgentsFrontier AI SystemsModel TrainingTechnical LeadershipOpen-Source Contributions

Tech Stack

Tools & technologies
PythonC++

About the role

Key responsibilities & impact
  • Design and own evaluation frameworks for advanced coding agents
  • Develop benchmark specifications, scoring methodologies, rubrics, and quality standards
  • Establish rigorous methods for measuring coding-model performance across diverse software-engineering tasks
  • Define objective criteria for correctness, reasoning quality, robustness, and task completion
  • Maintain methodological rigour, reproducibility, and consistency across evaluation workflows
  • Develop high-quality datasets, golden examples, and structured evaluation protocols
  • Design technical tasks for reliable assessment of frontier coding systems
  • Build data and evaluation workflows supporting model development and iterative improvement
  • Identify benchmark or dataset coverage gaps and develop new evaluation categories
  • Analyse coding-agent behaviour and identify systematic weaknesses, failure modes, and performance limitations
  • Investigate incorrect reasoning, implementation errors, tool-use failures, and incomplete task execution
  • Translate findings into recommendations for model training and evaluation
  • Design experiments testing hypotheses about coding-model capabilities
  • Build tooling and infrastructure for large-scale experimentation, data generation, review workflows, and evaluation pipelines
  • Automate technical processes to improve evaluation efficiency and research velocity
  • Collaborate with researchers, engineers, and applied AI teams
  • Contribute to technical reports, benchmark studies, research documentation, and external-facing research initiatives
  • Communicate complex technical findings to specialist and broader technical audiences

Requirements

What you’ll need
  • Strong software-engineering background with expertise in Python, C++, or comparable programming languages
  • Minimum of 3 years of experience in software engineering, machine learning, AI research, evaluation, or a related technical discipline
  • Experience designing, reviewing, or validating technical assessments, benchmarks, coding tasks, or evaluation methodologies
  • Familiarity with large language models, coding agents, reinforcement learning, model evaluation, or related AI systems
  • Proven ability to build tooling, automate workflows, and improve technical processes through systematic experimentation
  • Strong analytical skills and ability to investigate complex model behaviour and technical failure modes
  • Excellent written and verbal communication skills
  • Ability to operate effectively in fast-moving research environments with significant ambiguity and evolving priorities
  • Experience with frontier AI systems, coding agents, or model-evaluation research is advantageous
  • Experience designing benchmarks or datasets for machine-learning systems at scale is strongly valued
  • Familiarity with agentic workflows, tool use, reinforcement learning, or post-training methodologies is beneficial
  • Publications, open-source contributions, or demonstrated technical leadership in AI, machine learning, or software engineering are advantageous
  • Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party

Benefits

Comp & perks
  • Fully remote work
  • Full-time engagement
  • Opportunity to contribute to frontier AI research and development
  • Collaboration with researchers, engineers, and applied AI teams
  • Opportunity to contribute to technical reports, benchmark studies, research documentation, and external-facing research initiatives