FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Staff AI Platform Engineer
Code Metal. Set the technical direction and architecture for Code Metal's AI platform and lead the four-person team building it .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and operating AI platforms, with a strong focus on production-grade Python, API design, and LLM systems. Proven ability to lead technical teams, mentor engineers, and align platform development with organizational needs.
Highest-signal resume keywords
Production-Grade PythonAPI And Service DesignLLM Inference EnginesOpenTelemetry TracingStaff-Level Technical Leadership
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Distributed SystemsContainers And KubernetesCI/CDTestingContext EngineeringTransformersModel EvaluationFine-Tuning Language ModelsExperiment DesignBenchmarking
Soft Skills
MentoringTechnical DirectionCollaboration
Tools & Technologies
VLLMSGLangTensorRT-LLMPyTorchHugging FaceMLflowWeights & Biases
Certifications & Qualifications
U.S. CitizenshipEligibility For U.S. Security Clearance
Industry Keywords
AI SystemsLLM GatewayMulti-TenancyRegulated DefenseAerospace Domains
Tech Stack
Tools & technologiesDistributed SystemsKubernetesPythonPyTorch
About the role
Key responsibilities & impact- Set the technical direction and architecture for Code Metal's AI platform and lead the four-person team building it
- Own design documents and RFCs, make build-vs-buy decisions, and mentor the team
- Deploy, benchmark, and tune production inference for open-weight models on vLLM, SGLang, and TensorRT-LLM
- Own the model gateway for self-hosted and commercial models, including authentication, routing, failover, quotas, and cost attribution
- Design reusable agent harnesses and orchestration primitives for reliable, verifiable workflows
- Build context-engineering services for memory, retrieval, and data discovery
- Instrument the stack with OpenTelemetry traces and service metrics
- Build experiment-tracking and artifact infrastructure for reproducing and comparing results
- Design for multi-tenancy, versioned APIs, security, and deployment in customer and air-gapped environments
- Partner with Applied AI Research, product teams, and DevOps to align the platform with organizational needs
- Run experiments when platform decisions require evidence
- Build and operate production systems as the technical lead of Code Metal's AI Platform team
Requirements
What you’ll need- Production-grade Python and strong platform engineering fundamentals: API and service design, distributed systems, containers and Kubernetes, CI/CD, and testing
- Shipped production agentic systems, with a clear sense of where they break and how to make them reliable
- Experience with context engineering: retrieval-augmented generation, embeddings, vector or hybrid search, and memory for agents, ideally over code or large technical corpora
- Experience instrumenting services with OpenTelemetry tracing and metrics and operating AI services against SLOs
- Solid data science and AI research fundamentals, including transformers, LLM inference, experiment design, benchmarking, and model evaluation
- Working familiarity with PyTorch and Hugging Face
- Experience fine-tuning, evaluating, or serving language models
- Staff-level technical leadership across multiple systems or teams
- Experience writing design docs and RFCs, converting ambiguous internal-customer needs into a roadmap, and mentoring engineers
- Typically 8+ years of software engineering experience
- Typically 4+ years building and operating ML or LLM systems in production
- Demonstrated Staff-level scope and impact
- Production experience running an LLM gateway or proxy such as SMG or Bifrost, or equivalent API gateway experience, including routing, auth, rate limiting, quotas, failover, and cost attribution
- Hands-on experience deploying and tuning LLM inference engines such as vLLM, SGLang, or TensorRT-LLM on GPU infrastructure
- Familiarity with speculative decoding, prefix caching, tensor/pipeline/expert parallelism, disaggregated prefill and decode, and GPU profiling
- Experience building evaluation harnesses for LLMs and agents and experiment-tracking or artifact systems such as MLflow or Weights & Biases
- Experience taking an internal platform to an external product, including multi-tenancy, SDKs, versioned APIs, and documentation
- Experience deploying AI systems on-premises, in air-gapped or classified environments, or in regulated defense or aerospace domains
- U.S. Citizenship may be required for certain project assignments involving security clearance
- May require eligibility to obtain and maintain a U.S. security clearance
- Legally authorized to work in the United States
Benefits
Comp & perks- Offers equity
- Health care plan with 100% premium coverage, including medical, dental, and vision
- 401k with 5% matching
- Paid Time Off (uncapped vacation, plus sick and public holidays)
- Flexible hybrid or remote work arrangement
- Relocation assistance for qualifying employees