FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Applied AI Engineer, Agent Quality & Evaluations
Flodesk. Manage member-facing AI behavior across prompt-to-create, agentic editing and future jobs such as segmentation, scheduling, analytics and recommendations .
Posted 9/23/2026full-timeRemote • California • United StatesSenior💰 $150,000 - $250,000 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in managing AI behavior and evaluation practices, with strong software engineering skills in Python and TypeScript. Capable of translating member feedback into actionable system changes while ensuring quality through effective communication and systems thinking.
Highest-signal resume keywords
AI Evaluation PracticeSoftware Engineering Skills in PythonLLM Evaluation and ObservabilitySystems ThinkingClear Communication
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
PythonTypeScriptLLM EvaluationBehavioral ScenariosRegression SuitesAutomated ChecksExperimental DesignQuality BaselinesAgent LoopsContext Understanding
Soft Skills
Clear CommunicationProduct SenseCollaboration
Tools & Technologies
LangfuseBraintrustEvaluation Platforms
Industry Keywords
AI Behavior ManagementContent-Generation ToolsCreative ToolsMarketing ToolsMulti-Turn Agents
Tech Stack
Tools & technologiesPythonTypeScript
About the role
Key responsibilities & impact- Manage member-facing AI behavior across prompt-to-create, agentic editing and future jobs such as segmentation, scheduling, analytics and recommendations
- Create, test, version and document prompts and agent instructions
- Define system use of context, member data, tools and product state
- Build and prototype agent loops, tools and structured interfaces
- Own AI evaluation practice, including representative datasets, behavioral scenarios, scoring rubrics, regression suites and human review
- Build and operate evaluation harnesses, monitoring and quality dashboards
- Diagnose failures across prompts, context, orchestration, models, tools, data and product code
- Set quality baselines and release gates for prompt, model and agent changes
- Turn production failures, traces and member feedback into durable regression cases
- Partner with product, design, marketing and copy experts to encode Flodesk's point of view into the system
Requirements
What you’ll need- A track record of building or meaningfully improving production LLM or agentic products, not just prototypes
- Strong software engineering skills in Python, TypeScript or a similar language
- Understanding of the interaction between prompts, context, tools, agent loops, orchestration and product state
- Direct experience with LLM evaluation, observability and experimental design
- Ability to combine automated checks with human judgment
- Strong product sense for translating ambiguous member feedback into testable quality definitions and concrete system changes
- Systems thinking and reusable patterns
- Clear communication across technical and nontechnical disciplines
- Bonus: experience building content-generation, creative or marketing tools
- Bonus: experience building multi-turn agents that use tools and preserve user intent across edits
- Bonus: experience with Langfuse, Braintrust or comparable evaluation and observability platforms
- Bonus: experience working on products for small businesses, creators or marketers
Benefits
Comp & perks- Fully paid health insurance for individual coverage
- 16 weeks paid parental leave for non-birthing parents
- 22 weeks paid maternity leave for birthing parents
- Unlimited flexible time off
- 401(k) match (US employees only)
- $1,000 annual stipend for learning and development