FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Technical Product Manager, Data Science
CENTRL Inc. Own the accuracy, reliability, and cost-efficiency of the AI behind CentrlX .
Posted 9/21/2026full-timeRemote • California • United StatesMid-LevelSenior💰 $115,000 - $130,000 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and evaluating AI models, particularly in LLM-based products, with a strong focus on quality standards, dataset management, and cost optimization. Proficient in Python and SQL, with the ability to analyze data and implement effective evaluation strategies.
Highest-signal resume keywords
LLM-Based Product ManagementPython ProgrammingSQL ProficiencyEvaluation Dataset DevelopmentStatistical Literacy
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Data ScienceModel EvaluationDocument DigitizationRegression TestingQuality Standards DefinitionCost OptimizationBatch InferenceUser Story WritingAcceptance Criteria DevelopmentAgile Engineering Process
Soft Skills
Strong Written Communication
Tools & Technologies
BraintrustLangSmithLangfuseArize PhoenixW&B WeaveInspectPromptfoo
Industry Keywords
AI Quality EscalationsDocument AIAgentic SystemsFinancial ServicesInvestment Management
Tech Stack
Tools & technologiesPythonSQL
About the role
Key responsibilities & impact- Own the accuracy, reliability, and cost-efficiency of the AI behind CentrlX
- Build and own the evaluation foundation, including golden datasets, grading rubrics, LLM-as-judge pipelines calibrated against human labels, and regression suites
- Define measurable quality standards for document digitization, extraction, retrieval, groundedness, summaries, responses, evaluations, and multi-step agent runs
- Evaluate agent behavior, including tool selection, retrieval quality, step sequencing, and finished deliverable quality
- Convert client failures into permanent evaluation cases
- Benchmark models across providers on accuracy, latency, and cost
- Own model migrations and determine model deployment by workflow
- Track AI spend and optimize cost through routing, model tiering, caching, and context strategy
- Define logging and tracing requirements for prompts, retrieved context, outputs, tool calls, token counts, and latency
- Build in-app feedback capture with Product and Design
- Build, label, curate, and maintain evaluation datasets with holdouts and rotation
- Serve as the point of contact for AI quality escalations; triage, reproduce, identify root causes, and close the loop
- Write user stories and acceptance criteria and implement prompt, configuration, and model changes
- Publish regular quality and cost readouts
- Work with domain practitioners to encode industry judgment into evaluation rubrics
Requirements
What you’ll need- Must have work authorization in the USA
- 3+ years across product management, data science, or applied AI, including at least 2 years working on LLM-based products in production
- Hands-on Python and SQL
- Comfortable in a notebook pulling data, running batch inference, and computing metrics
- Demonstrated experience building evaluation datasets and harnesses for LLM systems: golden sets, rubric design, LLM-as-judge with human calibration, and regression testing against prompt and model changes
- Working fluency with at least one eval or LLM observability platform: Braintrust, LangSmith, Langfuse, Arize Phoenix, W&B Weave, Inspect, Promptfoo, or a comparable in-house harness
- Practical understanding of RAG systems: retrieval quality, groundedness and faithfulness, hallucination detection, and chunking and context strategy
- Statistical literacy: ability to size a comparison, judge significance, and identify when a difference is not real
- Ability to write clear user stories and acceptance criteria and work inside an agile engineering process
- Strong written communication
- Preferred: experience with document AI, agentic systems, financial services or investment management, inference cost reduction, in-product feedback mechanisms, agent frameworks and MCP
- Degree in a quantitative or technical field is preferred