FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing evaluation frameworks for LLM outputs and agentic workflows, with a strong focus on test automation, quality criteria definition, and integration into CI/CD processes. Proficient in API testing and building custom test harnesses while collaborating effectively with cross-functional teams.
Highest-signal resume keywords
6+ Years In QA, SDET, Or Test AutomationHands-On Experience Testing LLM-Based Or Agentic SystemsDeep Experience With Test Frameworks Such As Vitest, Jest, Or PytestAPI-First Testing Mindset, Including REST And PostmanWorking Knowledge Of CI/CD Pipelines
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Test AutomationLLM TestingPrompt Regression TestingCustom Test Harness DevelopmentProbabilistic Reasoning
Soft Skills
CoachingCollaboration
Tools & Technologies
PlaywrightCypressTypeScriptPythonGitHub Actions
Industry Keywords
Evaluation FrameworksRegression SuitesGolden DatasetsScoring RubricsShift-Left QA Model
Tech Stack
Tools & technologiesCloudCypressJestPythonTypeScript
About the role
Key responsibilities & impact- Design and maintain evaluation frameworks for LLM outputs and agentic workflows, including regression suites, golden datasets, and scoring rubrics
- Build test harnesses that catch hallucinations, tool-calling failures, prompt regressions, and unsafe or off-policy behavior before they reach production
- Define measurable quality criteria for agent reliability: task completion, factual grounding, latency, cost, and reasoning quality
- Integrate evaluation runs into CI/CD so model, prompt, and agent changes are gated the same way code changes are
- Partner with engineers on observability and tracing for agent runs, so failures are diagnosable rather than mysterious
- Contribute to conventional API and end-to-end automation where AI features sit inside larger product flows
- Help shape the team's shared playbook for testing AI features, and coach other QA engineers as agentic work spreads across scrum teams
- Work hands-on as an individual contributor reporting to the QA Manager
- Partner with product, engineering, and the broader QA team on LLM-driven and agentic workflows
Requirements
What you’ll need- 6+ years in QA, SDET, or test automation, with real production automation shipping
- Hands-on experience testing LLM-based or agentic systems: building evals, working with LLM-as-judge patterns, prompt regression testing, or agent trajectory analysis
- Prior experience in a shift-left, embedded QA model
- Comfort with at least one modern automation stack (Playwright, Cypress, or similar) and a typed language, TypeScript preferred, Python fine
- Deep experience with test frameworks such as vitest, jest, or pytest, and comfort building custom test harnesses rather than only running off-the-shelf suites
- API-first testing mindset, including REST and Postman or equivalent
- Fluency in HTTP-level API testing, including recording proxies and observing service-to-service traffic
- Working knowledge of CI/CD pipelines, GitHub Actions a plus, and how to plug evals into them
- Familiarity with cloud secret managers and disciplined handling of sensitive test data in restore-from-prod environments
- Ability to reason clearly about probabilistic systems: variance, sample sizes, confidence, and when a flaky result is signal rather than noise
Benefits
Comp & perks- Medical
- Dental
- Vision
- Flexible Spending Accounts
- PTO
- Paid and Floating Holidays
- 401k with Company match and immediate vesting
- Company-funded Life Insurance
- Employee Assistance Programs
- No-cost Mental Health Benefits
- Annual Bonus Opportunity of 10%
