Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Gramian Consulting

Technical AI Evaluation Analyst – LATAM

Gramian Consulting

. Validate task quality by checking instructions, source materials, reference solutions, and evaluation criteria for consistency and completeness.

Posted 9/22/2026contractRemote • Argentina, Brazil, Mexico, Uruguay, Peru, ChileMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in validating task quality and reviewing technical workflows, with a strong focus on analytical skills and attention to detail. Proficient in assessing technical deliverables and providing clear, evidence-backed feedback.

Highest-signal resume keywords
Python ProficiencySQL ProficiencyAnalytical SkillsTechnical Workflow ReviewAttention to Detail

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
PythonSQLShell ScriptsData AnalysisEvaluation LogicTechnical Deliverable AssessmentAutomated Quality ControlEvidence DocumentationGrading Logic AuditExecution Log Interpretation
Soft Skills
Strong Written EnglishClear CommunicationReproducible Feedback
Industry Keywords
AI Agent ExecutionModel PerformanceEvaluation CriteriaTask Quality ValidationDiscrepancy Investigation

Tech Stack

Tools & technologies
PythonSQL

About the role

Key responsibilities & impact
  • Validate task quality by checking instructions, source materials, reference solutions, and evaluation criteria for consistency and completeness.
  • Review AI agent execution traces, tool calls, and deliverables to determine whether outcomes are justified.
  • Audit grading logic to identify brittle checks, incorrect expected answers, unsupported rubric criteria, and unfair penalties for valid alternative solutions.
  • Investigate discrepancies between model performance, grader results, and expected outcomes.
  • Distinguish genuine model limitations from task defects, grader errors, and environment or tool failures.
  • Independently assess automated QC findings rather than accepting them without verification.
  • Document concise, evidence-backed findings and provide actionable, reproducible feedback.
  • Flag uncertainty and verify that implemented revisions resolve previously identified issues.

Requirements

What you’ll need
  • Comfortable reading and interpreting Python, SQL, shell scripts, structured data, and execution logs.
  • Experience reviewing technical workflows, software behavior, data outputs, or evaluation logic.
  • Strong analytical skills, including the ability to verify calculations and reconcile conflicting evidence.
  • Demonstrated ability to assess the correctness and completeness of technical deliverables.
  • Strong written English with experience providing clear, specific, and reproducible feedback.
  • High attention to detail when identifying inconsistencies, missing information, and evaluation defects.