FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Lead AI Platform Engineer
EQ Bank | Equitable Bank. Lead the engineering, configuration, and operation of enterprise AI platforms to ensure availability, performance, resilience, and scalability .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in platform engineering and site reliability engineering, with a strong focus on cloud operations, observability, and automation. Capable of leading technical delivery, incident response, and operational reporting while ensuring compliance with engineering and security standards.
Highest-signal resume keywords
Platform EngineeringSite Reliability EngineeringCloud OperationsObservability DesignAutomation Initiatives
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Incident ResponseRoot-Cause AnalysisService RestorationOperational ReportingConfiguration ManagementCI/CDInfrastructure-as-CodeTechnical DocumentationAI/ML Operational ConceptsRisk Assessment
Soft Skills
Technical LeadershipMentoringPrioritizationContinuous Improvement
Tools & Technologies
Azure AutomationAzure MonitorGrafanaAzure DevOpsTerraformBicepApplication InsightsLog AnalyticsAzure AICosmos DB
Industry Keywords
ITILITSMGovernance ControlsHuman-in-the-Loop PracticesResponsible AI
Tech Stack
Tools & technologiesApacheAzureCloudGrafanaITSMKafkaTerraformVault
About the role
Key responsibilities & impact- Lead the engineering, configuration, and operation of enterprise AI platforms to ensure availability, performance, resilience, and scalability
- Define and implement platform engineering patterns, standards, reusable components, and operational guardrails
- Lead platform triage, incident resolution, escalation coordination, and post-incident reviews
- Track and report service reliability indicators, incident trends, engineering risks, and operational performance improvements
- Enable approved AI use cases in non-production and production environments through environment readiness, dependency validation, release readiness, operational supportability, and service transition planning
- Partner with architecture, security, cloud, infrastructure, delivery, and application teams on secure, supportable, scalable platform implementations
- Guide release coordination, change readiness, maintenance planning, capacity planning, and technical risk mitigation
- Ensure AI platform changes meet engineering, operational, security, and control readiness criteria
- Design and improve observability capabilities including telemetry, logging, metrics, traces, dashboards, and alerting
- Lead automation initiatives to reduce manual effort, improve reliability, and standardize operational activities
- Analyze operational data for anomalies, recurring issues, root-cause patterns, performance bottlenecks, and service improvement opportunities
- Implement AI Ops use cases including alert correlation, anomaly detection, forecasting, root-cause support, knowledge retrieval, and repetitive-task automation
- Mentor engineers on observability, automation, troubleshooting, and service reliability
- Embed governance, security, privacy, auditability, traceability, and human oversight into AI platform engineering
- Assess implementation risks, close control gaps, maintain audit and governance evidence, and escalate technical and control risks
- Maintain visibility of AI platform assets, validate ownership and configuration integrity, and promote engineering standards and reusable patterns
Requirements
What you’ll need- University degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience
- 7+ years of experience in platform engineering, site reliability engineering, DevOps, cloud operations, enterprise IT operations, or production platform support
- Experience leading technical delivery, engineering standards, production readiness, incident response, problem management, service restoration, and operational reporting for enterprise platforms
- Advanced experience with cloud platforms, observability, automation, configuration management, and integration patterns
- Experience with Azure Automation runbooks, Azure AI, Copilot integrations, AKS, virtual networks, App Service, and supporting Azure services
- Expertise with Azure Monitor, Application Insights, Log Analytics, Grafana, dashboards, alerting, and operational telemetry design
- Experience with CI/CD, automation, and infrastructure-as-code tools including Azure DevOps, GitHub Actions, Logic Apps, Bicep, Terraform, Azure Policy, and Key Vault
- Knowledge of API Management, open-source API tools, Service Bus, Event Grid, and Apache Kafka
- Working knowledge of Elastic, Azure AI Search, Cosmos DB, and related data platform capabilities
- Knowledge of enterprise network, edge security, identity, access management, DNA, Fortinet, and Akamai is an asset
- Working knowledge of AI/ML operational concepts, model lifecycle support, platform telemetry, governance controls, human-in-the-loop practices, responsible AI, and production monitoring
- Understanding of ITIL/ITSM processes including change, release, incident, problem, configuration, service reporting, and operational risk practices
- Technical leadership, mentoring, standards development, implementation guidance, and complex cross-functional delivery coordination
- Advanced troubleshooting, root-cause analysis, prioritization, risk assessment, and continuous improvement skills
- Experience creating technical documentation, engineering patterns, operational procedures, support playbooks, dashboards, and user guidance materials
Benefits
Comp & perks- Full-time employment
- Hybrid work arrangement