Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Capital One

Director, AI Engineering

Capital One

. Partner with engineers, research scientists, technical program managers, and product managers to deliver AI-powered products .

Posted 10/6/2026full-timeRemote • United StatesLead💰 $244,700 - $279,200 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in AI and ML algorithm development, large-scale distributed training infrastructure, and GPU capacity planning. Proven ability to lead engineering teams, establish Responsible AI standards, and execute enterprise-level AI strategies.

Highest-signal resume keywords
AI And ML Algorithm DevelopmentLarge-Scale Distributed Training InfrastructureGPU Capacity PlanningPeople Leadership ExperiencePython Proficiency

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
AI TechnologiesML AlgorithmsDistributed TrainingGPU SchedulingPythonGoC++CUDAML OrchestrationParallelism Strategies
Soft Skills
Excellent CommunicationPresentation SkillsMentoring
Tools & Technologies
AWS UltraclustersHugging FaceKubernetesKubeflowSlurmRayKServeVLLMDeepSpeedMegatron
Industry Keywords
Responsible AI StandardsEthical StandardsRegulatory FrameworksAI StrategyCost Efficiency

Tech Stack

Tools & technologies
AWSCloudKubernetesNode.jsPythonPyTorchRayC++Go

About the role

Key responsibilities & impact
  • Partner with engineers, research scientists, technical program managers, and product managers to deliver AI-powered products
  • Oversee design, development, testing, deployment, and operation of distributed training, fine-tuning, reinforcement learning, GPU scheduling, fault tolerance, utilization, and model experimentation systems
  • Make build-versus-buy decisions across open-source and SaaS AI technologies including AWS Ultraclusters, Hugging Face, vector databases, and PyTorch
  • Introduce techniques improving scalability, cost, throughput, and reliability of large-scale distributed training and fine-tuning
  • Own GPU capacity planning and cost governance across teams
  • Contribute to the technical vision and long-term roadmap of foundational AI systems
  • Attract, retain, mentor, and develop AI engineering talent
  • Translate enterprise AI strategy into portfolio-level execution plans across product areas
  • Scale AI engineering practices through shared infrastructure, reusable components, observability, and governance frameworks
  • Establish enterprise Responsible AI standards covering fairness metrics, model evaluation, documentation, and audit readiness
  • Partner with research, compliance, and enterprise risk teams on ethical and regulatory standards

Requirements

What you’ll need
  • Bachelor's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 8 years of experience developing AI and ML algorithms or technologies; or Master's degree in these fields plus at least 6 years of such experience
  • At least 3 years of people leadership experience
  • 5+ years of experience managing and leading an engineering team (preferred)
  • 7+ years of experience building and operating large-scale ML or GPU training infrastructure on cloud platforms (preferred)
  • Hands-on experience with distributed training at scale, including multi-node, multi-GPU jobs and parallelism strategies such as PyTorch FSDP, DeepSpeed, and Megatron
  • Proficiency in Python, Go, C++, or CUDA
  • Experience with ML orchestration and scheduling tools such as Kubernetes, Kubeflow, Kueue, Slurm, Ray, KServe, and vLLM
  • Experience operating large GPU fleets with focus on reliability, fault tolerance, utilization, and cost efficiency
  • Experience right-sizing GPU clusters, instance types, interconnect, and quotas
  • Passion for current AI and ML-systems research and applying novel training and optimization techniques
  • Excellent communication and presentation skills
  • Experience building and leading a multi-team AI organization
  • Ability to execute long-term AI platform strategies aligned with enterprise priorities and regulatory frameworks
  • Experience establishing cross-functional operating rhythms and review cadences
  • Capital One will consider sponsoring a new qualified applicant for employment authorization

Benefits

Comp & perks
  • Performance-based incentive compensation, which may include cash bonuses and/or long-term incentives (LTI)
  • Comprehensive, competitive, and inclusive health, financial, and other benefits supporting total well-being
  • Employment authorization sponsorship may be considered for a new qualified applicant
  • Reasonable accommodations for applicants with disabilities