FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Tech Stack
Tools & technologiesAWSAzureCloudJavaScriptPythonTypeScriptGo
About the role
Key responsibilities & impact- Design and build LLM pipelines and agents for survey workflows
- Build AI-assisted data-to-report pipelines
- Enable natural-language analysis of survey data
- Develop multilingual questionnaire translation tools
- Build field-support tools for interviewers and supervisors
- Build evaluation harnesses with acceptance criteria, expert-produced ground truth, test sets, automated and human grading, regression testing, and ship/no-ship decisions
- Route work across models based on reasoning complexity, volume, cost, latency, and data-residency requirements
- Manage inference costs through token-usage tracking, context limits, prompt caching, and batch processing
- Connect models to data services and full-stack applications
- Use agentic coding tools to produce, review, test, and own production software
- Monitor production systems, conduct error analysis, and report positive and negative results
- Document methods and results for public release
- Transfer tools to partner-country institutions for independent operation
- Collaborate with the Technical Lead, survey experts, funders, and partner-government staff
Requirements
What you’ll need- Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent years of experience
- Minimum 4 years of professional software development experience, including building LLM-powered systems relied on by real users in production
- U.S. citizenship required by federal government contract
- Must be eligible to work in the United States without sponsorship
- Proficiency in Python or another modern programming language, such as JavaScript/TypeScript or Go
- Hands-on experience with frontier-model APIs, including tool use and retrieval-augmented generation
- Experience designing LLM evaluations, including test-set construction, automated and human grading, regression testing, and ship/no-ship decisions
- Working knowledge of LLM cost, latency, and model-selection trade-offs
- Experience building and deploying on AWS, Azure, Google Cloud, or another major cloud platform
- Experience using agentic coding tools to specify, review, test, and own production code
- Experience with agent frameworks and AI evaluation tooling
- Preferred: experience with subject-matter-expert ground truth, structured or tabular data, multilingual or low-resource-language NLP, open-weight models, sensitive or regulated data, model governance, global health, public work, or technical writing
- Ability to explain AI trade-offs clearly to technical and non-technical audiences
- Ability to collaborate across disciplines, cultures, time zones, and organizational levels
- Ability to communicate uncertainty, limitations, and negative results candidly
- Self-motivated and able to add value with limited direct supervision
Benefits
Comp & perks- Occasional travel
- Reasonable accommodations for disability, religious purposes, disabled veterans, individuals with disabilities, and sincerely held religious beliefs
- Equal opportunity employment
- Benefit offerings referenced under the Transparency in Coverage Act
