FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior ML / AI Engineer
Nonstop Administration & Insurance Services. Design and ship agentic systems using dynamic tool-calling agents, structured Pydantic outputs, governed tool catalogs, per-step verification, grounding checks, LLM-as-judge, human-in-the-loop, and bounded auditable loops .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and shipping ML/AI systems in production, with a strong focus on LLM applications, MLOps governance, and cloud infrastructure management. Proficient in Python programming, system optimization, and ensuring reliability and compliance in regulated environments.
Highest-signal resume keywords
ML/AI Systems DevelopmentPython ProgrammingMLOps on AWSLLM Application ExperienceRegulated Data Compliance
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
ML/AI SystemsPythonLLM Application WorkPromptingEmbeddingsRAGKubernetesTerraformPostgresMongoDB
Soft Skills
Clear Written CommunicationJudgment Under Ambiguity
Tools & Technologies
EKSOpenTelemetryPrometheusGrafanaReactTypeScriptVLLMSQSCognitoAmazon Bedrock
Industry Keywords
HIPAASOC 2ISO 27001PHIPIIAudit TrailsHealthcareInsuranceClaimsX12 835
Tech Stack
Tools & technologiesAWSCloudGrafanaKubernetesMongoDBPostgresPrometheusPythonReactTerraformTypeScript
About the role
Key responsibilities & impact- Design and ship agentic systems using dynamic tool-calling agents, structured Pydantic outputs, governed tool catalogs, per-step verification, grounding checks, LLM-as-judge, human-in-the-loop, and bounded auditable loops
- Build the RAG layer, including document ingestion, chunking, embeddings, hybrid retrieval, reranking, access scoping, and injection/poisoning defenses
- Operate self-hosted inference at scale with vLLM and embedding/rerank services on EKS GPU nodes
- Optimize inference throughput using continuous batching, prefix caching, quantization, and KV-cache tuning
- Enforce fair-share concurrency across tenants
- Own the MLOps and governance plane, including offline evaluation harnesses, golden sets, quality scoring, canary/gray releases, auto-rollback, cost/budget governance, and observability
- Make systems production-safe through durable state, circuit breakers, retries, PHI-safe logging/redaction, RBAC, fail-closed defaults, and horizontal AWS scalability via Terraform
- Extend the platform to new domains from problem framing through governed, evaluated, deployed services
- Mentor engineers on building initiatives on the shared platform foundation
- Maintain clean, typed, tested code and participate in thoughtful design reviews
- Balance autonomy and determinism for high-stakes tasks
Requirements
What you’ll need- 5+ years building and shipping ML/AI systems in production, including hands-on LLM application work in the last 1–2 years
- Strong Python (typed, tested, production-grade) and solid software-engineering fundamentals
- Comfortable across an async web service, a data layer, and infrastructure
- Practical depth in prompting, structured/function-calling outputs, RAG, embeddings, vector search, retrieval quality, agent/tool-use loops, grounding, hallucination control, eval sets, and guardrails
- Cloud and MLOps experience on AWS or equivalent, including containers, Kubernetes, Terraform, CI/CD, observability, and model-serving cost/performance tuning
- Track record of owning reliability, including state durability, failure handling, scaling, and debugging production incidents
- Clear written communication and judgment under ambiguity
- Experience self-hosting/optimizing open-weight models such as vLLM or TGI on GPUs
- Experience with embeddings/rerank serving, such as bge or Infinity
- Experience with LangGraph, LangChain, or comparable agent frameworks and multi-agent orchestration
- Regulated-data experience with HIPAA, SOC 2, ISO 27001, PHI/PII handling, RBAC, and audit trails
- pgvector/Postgres, MongoDB/DocumentDB, SQS, and Cognito/OIDC
- OpenTelemetry, Prometheus/Grafana, and frontend development with React/TypeScript
- Amazon Bedrock or a multi-provider abstraction
- On-premises or air-gapped deployment experience
- Healthcare/insurance/claims domain experience, including X12 835, EOB, or benefits, is a strong plus
- Must be authorized to work for any employer in the United States; this job does not offer visa sponsorship
Benefits
Comp & perks- 100% coverage of medical, dental and vision benefits for employees and dependents
- 401k match up to 4%