FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and operating evaluation systems for Generative AI and complex machine learning systems, with a strong focus on architectural decisions that enhance user experience and reliability. Proficient in Python programming and familiar with evaluation methodologies, distributed systems, and risk management in AI.
Highest-signal resume keywords
Python ProgrammingEvaluation System DesignGenAI Evaluation FrameworksDistributed Systems PrinciplesScalable Platform Development
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Evaluation System DesignBenchmark DesignAutomated MetricsModel-Based GradingFailure AnalysisData VersioningOrchestrationRelease ManagementProduction MonitoringHuman Evaluation
Soft Skills
Judgment Around GenAI RisksCross-Team CollaborationTechnical Leadership
Tools & Technologies
AutoEvalsOpenEvalsRagasAI Tools
Industry Keywords
Generative AIMachine LearningEvaluation MethodologiesUser ExperienceReliability
Tech Stack
Tools & technologiesDistributed SystemsPython
About the role
Key responsibilities & impact- Identify modernization opportunities and scale the roadmap for a company-wide evaluation system in partnership with technical leadership
- Define cross-team technical strategy, align partner teams, and drive broad adoption of shared evaluation capabilities
- Make durable architectural decisions that improve the user experience and reliability of Pinterest’s AI/ML evaluation platform
- Establish evaluation methodologies and standards balancing automated metrics, expert judgment, human feedback, safety, and business outcomes
- Lead complex cross-team workstreams, resolve technical tradeoffs and risks, and deliver impact
- Shape and scale Pinterest’s AI/ML evaluation platform into a unified, reliable system adopted company-wide
- Partner across engineering, research, product, and safety to improve quality, launch decisions, and trust in AI-powered experiences
Requirements
What you’ll need- 3+ years of deep, hands-on experience designing and operating evaluation systems for GenAI and complex ML systems
- Experience with benchmark design, golden datasets, automated metrics, human evaluation, model-based grading, and failure analysis
- Strong Python programming skills
- Solid grasp of distributed systems principles
- Comfortable iterating with AI tools on coding and design
- Deep familiarity with GenAI and RecSys evaluation libraries and frameworks, including AutoEvals, OpenEvals, and Ragas
- 5+ years of experience building scalable platforms for distributed experimentation, data and model versioning, orchestration, observability, release management, and production monitoring
- Strong judgment around GenAI risks, including robustness, bias, privacy, security, and limitations of automated evaluation signals
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent experience
Benefits
Comp & perks- Equity eligibility
- Flexible work model / PinFlex
- In-person collaboration 1–2 times per quarter
- Equal opportunity workplace
- Medical or religious accommodation support during the application process
