FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

AI Quality & Evaluation Lead
Samsung Food. Build the experience-quality evaluation layer for AI-generated coaching in Samsung Health .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and validating failure taxonomies for AI-generated outputs, with a strong focus on creating binary criteria and annotation guidelines. Proficient in conducting error analysis and maintaining rigorous documentation for evaluation processes.
Highest-signal resume keywords
Failure Taxonomy DevelopmentBinary Criteria CreationLLM Judge ValidationError AnalysisAnnotation Guideline Writing
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Evaluation Loop ExecutionQualitative CodingData AnalysisTaxonomy CategorizationTrace Grading
Soft Skills
CollaborationDocumentationCritical Thinking
Tools & Technologies
NotebooksSpreadsheets
Industry Keywords
AI Product QualityConversation DesignHuman Data OperationsModel BehaviourContent Design
About the role
Key responsibilities & impact- Build the experience-quality evaluation layer for AI-generated coaching in Samsung Health
- Author a failure taxonomy for coaching outputs
- Hand-grade at least 150 traces with open-coded notes, including at least 40 sparse-data synthetic profiles
- Categorize failures into 5–10 categories and quantify their frequency
- Freeze at least 50 traces as a holdout set
- Create binary pass/fail criteria for each taxonomy category
- Write annotation guidelines with pass and fail examples
- Double-code at least 30 traces, resolve disagreements, and maintain a guideline revision log
- Create judge prompts for every criterion
- Validate the LLM judge using true-positive and true-negative rates on development and holdout sets
- Run weekly readouts on helpfulness, relevance, and tone
- Document a revalidation routine triggered by model or prompt changes and quarterly regardless
- Produce a complete playbook covering grading, taxonomy updates, rubric revisions, judge revalidation, and weekly readouts
- Conduct a handover test enabling the internal owner to rerun validation and a weekly readout independently
- Obtain Head of Product sign-off on the taxonomy and rubric
- Collaborate with and hand over the method to an internal owner
Requirements
What you’ll need- Experience running the full evaluation loop at least once on conversational or generated-text output
- Experience with error analysis on real traces
- Experience building a failure taxonomy
- Experience creating binary criteria and annotation guidelines
- Experience validating an LLM judge against personal labels
- Background in conversation design, AI product quality, model behaviour/policy, content design, human data operations, UX research with strong qualitative coding, RLHF, applied linguistics, or product management
- Ability to identify previously unnamed failure modes by reading output
- Ability to act as arbiter and document overrulings and rationale
- Understanding that rubrics are discovered through grading
- Ability to explain limitations of judge agreement rates
- Ability to write unambiguous annotation guidelines
- Comfort working in notebooks and spreadsheets
- No production coding required
- Nutrition, weight management, or behaviour change domain expertise is not required
- Must be comfortable being responsible for own taxes and having no paid time off
Benefits
Comp & perks- Remote-first work arrangement
- Independent contractor setup
- Flexible remote work
- Global team of more than 100 people in over 30 countries