Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
BMO U.S.

Principal Cloud Engineer – AI

BMO U.S.

. Define reference and production-grade AI/GenAI solutions on cloud .

Posted 9/23/2026full-timeUnited StatesLead💰 $120,000 - $250,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in building and managing cloud-native AI infrastructure, with a strong focus on observability, security, and operational excellence. Proficient in implementing infrastructure as code and driving FinOps for AI platforms to optimize performance and cost.

Highest-signal resume keywords
Cloud Infrastructure ManagementAI/ML Infrastructure ExpertiseInfrastructure As Code (IaC)CI/CD Pipeline DevelopmentKubernetes and Containerization

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Python ProgrammingTerraform/BicepKubernetesGPU OptimizationCI/CDObservabilityNetworkingSecurityAI/ML InfrastructureMLOps/LLMOps
Soft Skills
MentoringCollaborationStakeholder Engagement
Tools & Technologies
AzureAWSPrometheusGrafanaOpenTelemetryKafkaAzure Event HubsMLflowKServeHugging Face
Industry Keywords
Cloud-Native PatternsService MeshServerlessFeature StoresVector DatabasesResponsible AIData ResidencyAudit ReadinessError BudgetsChaos Testing

Tech Stack

Tools & technologies
AWSAzureCloudGrafanaKafkaKubernetesMicroservicesPrometheusPythonTerraformTypeScriptGo

About the role

Key responsibilities & impact
  • Define reference and production-grade AI/GenAI solutions on cloud
  • Build reusable, secure, and observable APIs, SDKs, microservices, and pipelines
  • Operationalize LLMs and RAG with Responsible AI guardrails
  • Drive platform roadmaps that enable faster delivery, lower risk, and measurable business outcomes
  • Design, build, and operate cloud-native AI infrastructure for ML/GenAI workloads
  • Manage GPU/CPU clusters, autoscaling, spot instance strategies, networking, storage, databases, IAM, secrets, encryption, and policy-as-code
  • Implement observability and reliability for AI infrastructure using metrics, logging, tracing, SLOs, and SLIs
  • Build CI/CD and GitOps pipelines for infrastructure-as-code and AI platform components
  • Drive FinOps for AI infrastructure through GPU rightsizing, caching, inference optimization, and cost governance
  • Enable secure APIs, microservices, event-driven architectures, model runtimes, and RAG retrieval pipelines
  • Define and evolve AI infrastructure reference architecture, including Kubernetes, service mesh, ingress, serverless, multi-region, HA/DR, and compliance-ready designs
  • Establish standards for containerization, IaC, secure networking, and AI systems
  • Implement defense-in-depth security and Responsible AI controls, including data residency, encryption, lineage, and audit readiness
  • Lead infrastructure discovery and solution design with stakeholders
  • Operate platforms with SRE principles, including error budgets, incident response, and chaos testing
  • Mentor engineers and create reusable IaC modules, templates, and golden paths

Requirements

What you’ll need
  • Bachelor’s/Master’s/PhD in CS, Engineering, or related field
  • 7+ years building large-scale distributed cloud infrastructure
  • 5+ years hands-on with Azure/AWS
  • Proven experience with AI/ML infrastructure, including GPU clusters, Kubernetes, CI/CD, and observability
  • Strong in infrastructure as code, Terraform/Bicep, Kubernetes, networking, and security
  • Expertise in cloud-native patterns, including containers, service mesh, and serverless
  • Familiarity with MLOps/LLMOps infrastructure, model serving, feature stores, and vector databases
  • Programming in Python and one of Go/TypeScript for tooling
  • Understanding of frontend/backend integration for AI services
  • GPU optimization knowledge, including CUDA/NCCL and TensorRT-LLM
  • Familiarity with Prometheus, Grafana, and OpenTelemetry
  • Experience with Kafka/Azure Event Hubs and real-time systems
  • Experience with AI platform products such as Azure ML, MLflow, KServe, or Hugging Face

Benefits

Comp & perks
  • Performance-based incentives
  • Discretionary bonuses
  • Health insurance
  • Tuition reimbursement
  • Accident and life insurance
  • Retirement savings plans
  • In-depth training and coaching
  • Manager support
  • Network-building opportunities
  • Tools and resources to reach new milestones
  • Equal employment opportunity
  • Reasonable accommodations for individuals with disabilities