FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and implementing enterprise-scale observability frameworks for AI systems, with a strong focus on telemetry, monitoring, and operational reporting. Proven ability to lead teams, manage projects, and translate complex operational requirements into scalable technical solutions.
Highest-signal resume keywords
AI EngineeringObservability FrameworksTelemetry CollectionAzure Kubernetes ServicePython Development
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Observability PrinciplesTelemetry PipelinesMonitoring ArchitecturesOperational AnalyticsDistributed Systems MonitoringAgentic AI ImplementationLangChainLangGraphFastAPIFiddler
Soft Skills
Problem-SolvingCommunicationPresentationLeadershipMentoring
Tools & Technologies
Microsoft Azure AI ServicesMCP IntegrationsEvent-Driven ArchitecturesMQTTAI Lifecycle Monitoring
Industry Keywords
Operational IntelligenceGovernanceCompliance MonitoringAuditability SolutionsPerformance Optimization
Tech Stack
Tools & technologiesAzureDistributed SystemsKubernetesPython
About the role
Key responsibilities & impact- Lead the design and implementation of enterprise-scale observability frameworks for AI, GenAI, Agentic AI, and Agentic RAG systems
- Define enterprise standards for AI telemetry, tracing, monitoring, logging, evaluation, alerting, governance, and operational reporting
- Design traceability frameworks for agent reasoning, tool usage, retrieval paths, model interactions, and workflow execution
- Build end-to-end observability solutions for AI agents, retrieval systems, APIs, and distributed AI applications
- Develop telemetry pipelines to monitor model performance, latency, cost, quality, reliability, and user interactions
- Develop monitoring and evaluation frameworks using LangGraph, LangChain, MCP integrations, and Azure AI services
- Build production-grade observability services, APIs, and monitoring components using Python and FastAPI
- Design operational dashboards, alerting frameworks, and real-time monitoring solutions
- Establish evaluation, benchmarking, experimentation, and continuous improvement processes
- Deploy and manage scalable AI observability solutions on Azure Kubernetes Service
- Integrate telemetry from AI workloads, enterprise applications, APIs, databases, and distributed systems
- Design pipelines for collection, processing, and enrichment of observability data
- Collaborate with AI engineers, platform teams, data engineers, business stakeholders, clients, risk, and governance teams
- Manage and mentor AI engineers, platform engineers, and observability specialists
- Communicate operational health, risks, performance trends, and observability insights to clients and leadership
- Support proposals, AI governance, platform modernization, and operational transformation programs
- Drive reusable frameworks, accelerators, engineering standards, thought leadership, talent development, coaching, and technical mentorship
Requirements
What you’ll need- 8–12+ years of total professional experience
- 4+ years of directly relevant experience in AI engineering, observability, telemetry, distributed systems monitoring, or AI operations
- 2+ years of people, technical, or delivery leadership experience
- Proven experience delivering enterprise-scale monitoring, observability, or operational intelligence platforms
- Strong depth in at least 3 listed areas, with direct ownership in at least 2
- Ability to translate operational and governance requirements into scalable technical solutions
- Bachelor’s/master’s degree in computer science, Data Science, Engineering, or related field
- Deep expertise in Agentic AI, Agentic AI implementation, and Agentic RAG systems
- Experience with observability principles, telemetry collection, traceability frameworks, monitoring architectures, and operational analytics
- Experience with LangChain, LangGraph, Model Context Protocol (MCP), Microsoft Azure AI services, and Azure Kubernetes Service (AKS)
- Experience developing monitoring services, telemetry collectors, and operational APIs using Python and FastAPI
- Experience with distributed tracing, logging frameworks, alerting, anomaly detection, dashboards, and operational reporting
- Experience with AI observability and evaluation platforms including Fiddler or similar technologies
- Familiarity with MQTT, event-driven architectures, messaging systems, MLOps, LLMOps, and AI lifecycle monitoring
- Experience implementing Responsible AI controls, governance, compliance monitoring, and auditability solutions
- Strong understanding of security, reliability, scalability, performance optimization, and operational support for enterprise AI systems
- Excellent problem-solving, communication, and presentation skills
- Ability to clearly explain at least two relevant professional AI implementations
Benefits
Comp & perks- Continuous learning and development opportunities
- Tools and flexibility to make a meaningful impact
- Coaching and leadership development
- Diverse and inclusive culture
- Opportunity to work with global clients and AI experts
- Work on AI-driven innovation across multiple client sectors
