FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Software Engineer – Artificial Intelligence/LLM
Beacon Venture Capital. Design and implement retrieval-augmented generation and tool-calling flows .
Posted 10/3/2026full-timeSan Carlos • California • United StatesSenior💰 $173,000 - $225,000 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and implementing retrieval-augmented generation systems, with a strong focus on integrating LLM features and optimizing performance metrics. Proficient in production coding, testing, and documentation, while ensuring compliance with safety-critical domain standards.
Highest-signal resume keywords
Production ML/LLM Systems ExperienceLLM Feature Shipping and ImprovementEmbeddings and Vector Search UnderstandingA/B Testing and Metrics EvaluationDevOps Basics for CI/CD and IaC
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
PythonTypeScriptJSONEmbeddingsChunkingFunction CallingPrompt VersioningLatency TrackingData SynchronizationContent and Policy Checks
Soft Skills
CollaborationProblem-SolvingOwnershipAttention to DetailAdaptability
Tools & Technologies
AWS BedrockOpenSearchPgvectorPineconeDynamoDBS3TritonTensorRT-LLMCI/CDIaC
Industry Keywords
AviationSafety-Critical DomainMultimodal WorkGPU InferenceData Services
Tech Stack
Tools & technologiesAWSDynamoDBPythonTypeScript
About the role
Key responsibilities & impact- Design and implement retrieval-augmented generation and tool-calling flows
- Deliver robust JSON and schema-bound outputs with validation, retries, and fallbacks
- Integrate function calling with internal tools, search, routing, and data services
- Ship APIs and workers in Python or TypeScript with clear contracts, streaming, and backoff
- Add caching, request shaping, prompt templates, and context packing to control latency and cost
- Integrate with AWS Bedrock, OpenAI, Anthropic, or self-hosted endpoints
- Collaborate on chunking, embeddings, and indexing for documents, time series, and multimedia
- Choose and tune vector backends such as OpenSearch, pgvector, or Pinecone
- Maintain knowledge bases through data synchronization from S3, Aurora, DynamoDB, and external sources
- Create offline evaluations and golden sets for prompts, retrievers, and tools
- Establish online metrics for task success, hallucination rate, retrieval precision/recall, p95 latency, and cost per request
- Run A/B tests and prompt/version rollouts with guardrails and canaries
- Implement content and policy checks, PII detection and redaction, access controls, and auditing
- Design human-in-the-loop paths for sensitive actions
- Add tracing, logs, and dashboards for model calls, token usage, errors, and saturation
- Debug failures across retrieval, prompts, tools, and providers
- Partner with product, infrastructure, ML, and security teammates to deliver reliable aviation systems
- Build user-facing LLM features for Beacon AI’s platform that makes flying safer, more efficient, and more capable
Requirements
What you’ll need- 5–8 years of experience, including some experience in production ML/LLM systems
- Experience shipping LLM features to users and improving them with data
- Production coding, testing, and documentation skills
- Understanding of embeddings, chunking, vector search tradeoffs, and function calling
- Experience designing evaluations, defining success metrics, and iterating based on evidence
- Ability to track p95 latency, meet SLAs, and reduce cost without compromising quality
- Ability to own features from design through production with minimal oversight
- Experience with Bedrock, OpenSearch Serverless, pgvector, Pinecone, or Weaviate preferred
- Experience with prompt versioning, guardrails, and provider routing preferred
- Multimodal work with time series or video preferred
- Familiarity with GPU inference, Triton, or TensorRT-LLM preferred
- Aviation or other safety-critical domain exposure preferred
- DevOps basics for CI/CD, IaC, and secure secrets handling preferred
- Must be a U.S. Person
- Must be authorized to work lawfully in the United States
- No visa sponsorship or visa transfers available
- All work must be performed in the United States
- Must be based in or willing to relocate to the San Francisco Bay Area
- Must work onsite in San Carlos, CA, 3+ days per week
- Full-time availability
Benefits
Comp & perks- Offers equity
- 100% of employee medical premiums covered
- 25% of dependent medical premiums covered
- 3 weeks PTO
- 13+ paid company holidays
- 401(k) offered
- Flexible remote work on remaining days of the hybrid schedule