Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Impiricus

Director, AI Automation Engineering

Impiricus

. Own infrastructure that continuously verifies whether autonomous agents take intended actions before and after production .

Posted 9/22/2026full-timeRemote • New York • United StatesLead💰 $185,000 - $210,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in building and evaluating AI/ML systems, particularly in agentic AI and LLM-based automation. Proficient in establishing quality standards and implementing automated evaluation frameworks to ensure software quality across AI initiatives.

Highest-signal resume keywords
AI/ML Systems EvaluationRAG EvaluationGitHub Actions CI/CDPython ProficiencyAgent Testing

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
AI/ML SystemsAgentic AILLM-Based AutomationAutomated Evaluation FrameworksQuality MetricsCI/CD PipelinesTesting FrameworksBot Lifecycle ManagementRegression TestingBenchmarking
Soft Skills
Independent OperationCommunication
Tools & Technologies
GitHub ActionsLangSmithBraintrustSage
Industry Keywords
HealthcarePharmaRegulated Environments

Tech Stack

Tools & technologies
Python

About the role

Key responsibilities & impact
  • Own infrastructure that continuously verifies whether autonomous agents take intended actions before and after production
  • Stage, test, and lifecycle-manage bots and agent workflows
  • Establish baseline quality metrics and coverage standards for AI initiatives
  • Build infrastructure to run automated research chains and agent evaluations at scale
  • Create automated evaluation frameworks for LLM outputs covering accuracy, safety, bias, and regression
  • Implement RAG evaluation, prompt regression testing, and golden-set benchmarking
  • Define and enforce quality thresholds before release
  • Partner with AI Researchers and Engineers on evaluation design for new model and agent patterns
  • Build release-gating infrastructure, including staged rollouts and rollback authority
  • Develop AI automation solutions to maintain software quality across the product portfolio
  • Integrate testing into GitHub Actions or equivalent for AI model and agent releases
  • Maintain test suites as models, prompts, and tools change
  • Collaborate with AI Solutions Architects, AI Engineers, Principal AI Product, and product teams on evaluation, launch readiness, QA, and feature validation
  • Document standards, runbooks, and patterns so non-engineers building in Sage can self-serve basic testing
  • Set technical standards and grow into leading a small automation team

Requirements

What you’ll need
  • 5+ years building, evaluating, or governing production AI/ML systems, with direct experience in agentic AI, LLM-based automation, or autonomous system safety
  • Direct experience evaluating LLM outputs, including RAG evaluation, prompt regression, safety, and accuracy benchmarking
  • Experience building CI/CD pipelines with integrated test gates using GitHub Actions or similar
  • Comfortable working across AI platform infrastructure and product-facing AI features
  • Ability to operate independently as a senior IC and communicate quality standards to engineers and non-engineers
  • Experience with agent testing, bot lifecycle management, or AI agent orchestration platforms is a strong plus
  • Familiarity with LLM observability and evaluation tooling such as LangSmith, Braintrust, or custom harnesses is a strong plus
  • Background in healthcare, pharma, or regulated environments is a strong plus
  • Prior experience establishing a QA or automation function from scratch, with appetite to grow into leading a team
  • Proficiency in Python and standard testing frameworks such as pytest or equivalent

Benefits

Comp & perks
  • Medical, dental, and vision coverage for you and your dependents
  • On-demand healthcare concierge
  • Pre-tax HSA, FSA, and DCFSA savings options
  • Monthly employer HSA contributions if enrolled in a high-deductible plan
  • 100% paid short- and long-term disability insurance
  • Life and AD&D insurance
  • Flexible vacation policy
  • Paid parental leave after 6 months
  • Remote work option
  • 401(k) with company match