FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and leading the architecture of enterprise-scale AI and ML platforms, focusing on operational excellence, scalability, and responsible AI practices. Proven ability to mentor teams and drive technical decisions that enhance system performance and reliability.
Highest-signal resume keywords
AI And ML System DesignMLOps Best PracticesTechnical LeadershipObservability And MonitoringLarge Language Models
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Software EngineeringMachine Learning EngineeringAI EngineeringReliability EngineeringCloud TechnologiesModel ServingInference OptimizationAI InfrastructureArchitectural StandardsIncident Management
Soft Skills
MentoringCoachingProblem-SolvingCollaborationTechnical Guidance
Industry Keywords
Trustworthy AIResponsible AIAI GovernanceModel Risk ManagementProduction EnvironmentsOperational ReadinessScalable AI DeploymentMulti-Agent WorkflowsPerformance DegradationCapacity Planning
Tech Stack
Tools & technologiesCloud
About the role
Key responsibilities & impact- Define and lead technical architecture for enterprise-scale AI and ML platforms
- Design scalable, resilient, and reusable AI systems for mission-critical workloads
- Establish architectural standards, engineering patterns, and best practices for AI deployment and operations
- Drive technical decisions around model serving, inference optimization, agent architectures, orchestration frameworks, observability, and AI infrastructure
- Partner with AI researchers to transform prototypes into production-grade solutions
- Operationalize advanced AI capabilities including LLMs, Trustworthy and Responsible AI, and agentic AI systems
- Establish repeatable pathways to accelerate innovation-to-production cycles
- Solve complex AI engineering and scalability challenges
- Design systems balancing performance, latency, governance, security, and cost
- Drive adoption of MLOps, LLMOps, and AI platform engineering best practices
- Improve robustness, maintainability, observability, and operational readiness of AI products
- Identify and eliminate architectural bottlenecks affecting scale, reliability, or client experience
- Provide coaching, architecture reviews, design guidance, and technical leadership
- Own operational excellence, reliability, performance, and availability of products
- Lead response and resolution for production incidents, performance degradation, model failures, and system outages
- Serve as senior technical escalation point for challenging production issues
- Establish practices for monitoring, observability, alerting, incident management, capacity planning, and SLOs
- Mentor junior engineers in troubleshooting, root cause analysis, operational decision-making, and incident response
- Lead post-incident reviews and long-term corrective actions
- Develop operational processes that keep AI solutions secure, scalable, performant, and reliable
- Partner with product, infrastructure, security, and support teams to identify operational risks and improve service reliability
- Mentor AI and ML engineers and foster technical excellence and operational ownership
- Represent the team as a thought leader in scalable AI deployment, operational excellence, and responsible AI practices
Requirements
What you’ll need- 10+ years of experience in software engineering, machine learning engineering, AI engineering, or related technical disciplines
- Deep expertise designing, deploying, and supporting large-scale AI and ML systems in production environments
- Demonstrated success leading complex technical initiatives from concept through deployment and ongoing operations
- Strong knowledge of software architecture, reliability engineering, observability, ML Ops, DevOps, and cloud technologies
- Proven ability to mentor engineers and lead teams through highly complex technical and operational challenges
- Preferred: Experience with foundation models, Large Language Models, and agentic AI architectures
- Preferred: Experience deploying agentic AI systems and multi-agent workflows
- Preferred: Experience with Trustworthy AI, Responsible AI, AI governance, or model risk management frameworks
- Preferred: Experience optimizing large-scale inference systems and AI infrastructure
- Preferred: Experience working in highly regulated environments and mission-critical production systems
- Must be able to work without employer-sponsored visa; Vanguard is not offering visa sponsorship
Benefits
Comp & perks- Hybrid working model enabling flexibility, in-person learning, collaboration, and connection
- Opportunities to work directly with world-class AI researchers
- Opportunity to develop deep expertise in operating advanced AI systems at scale
- Collaboration with leaders across research, product, and engineering
