Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Samsung Food

AI Quality & Evaluation Lead

Samsung Food

. Build the experience-quality evaluation layer for AI-generated coaching in Samsung Health .

Posted 10/5/2026contractRemote • PolandSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in building and validating failure taxonomies for AI-generated outputs, with a strong focus on creating binary criteria and annotation guidelines. Proficient in conducting error analysis and maintaining rigorous documentation for evaluation processes.

Highest-signal resume keywords
Failure Taxonomy DevelopmentBinary Criteria CreationLLM Judge ValidationError AnalysisAnnotation Guideline Writing

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Evaluation Loop ExecutionQualitative CodingData AnalysisTaxonomy CategorizationTrace Grading
Soft Skills
CollaborationDocumentationCritical Thinking
Tools & Technologies
NotebooksSpreadsheets
Industry Keywords
AI Product QualityConversation DesignHuman Data OperationsModel BehaviourContent Design

About the role

Key responsibilities & impact
  • Build the experience-quality evaluation layer for AI-generated coaching in Samsung Health
  • Author a failure taxonomy for coaching outputs
  • Hand-grade at least 150 traces with open-coded notes, including at least 40 sparse-data synthetic profiles
  • Categorize failures into 5–10 categories and quantify their frequency
  • Freeze at least 50 traces as a holdout set
  • Create binary pass/fail criteria for each taxonomy category
  • Write annotation guidelines with pass and fail examples
  • Double-code at least 30 traces, resolve disagreements, and maintain a guideline revision log
  • Create judge prompts for every criterion
  • Validate the LLM judge using true-positive and true-negative rates on development and holdout sets
  • Run weekly readouts on helpfulness, relevance, and tone
  • Document a revalidation routine triggered by model or prompt changes and quarterly regardless
  • Produce a complete playbook covering grading, taxonomy updates, rubric revisions, judge revalidation, and weekly readouts
  • Conduct a handover test enabling the internal owner to rerun validation and a weekly readout independently
  • Obtain Head of Product sign-off on the taxonomy and rubric
  • Collaborate with and hand over the method to an internal owner

Requirements

What you’ll need
  • Experience running the full evaluation loop at least once on conversational or generated-text output
  • Experience with error analysis on real traces
  • Experience building a failure taxonomy
  • Experience creating binary criteria and annotation guidelines
  • Experience validating an LLM judge against personal labels
  • Background in conversation design, AI product quality, model behaviour/policy, content design, human data operations, UX research with strong qualitative coding, RLHF, applied linguistics, or product management
  • Ability to identify previously unnamed failure modes by reading output
  • Ability to act as arbiter and document overrulings and rationale
  • Understanding that rubrics are discovered through grading
  • Ability to explain limitations of judge agreement rates
  • Ability to write unambiguous annotation guidelines
  • Comfort working in notebooks and spreadsheets
  • No production coding required
  • Nutrition, weight management, or behaviour change domain expertise is not required
  • Must be comfortable being responsible for own taxes and having no paid time off

Benefits

Comp & perks
  • Remote-first work arrangement
  • Independent contractor setup
  • Flexible remote work
  • Global team of more than 100 people in over 30 countries