FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Lead AI Operations Engineer
Johnson & Johnson. Design and implement end-to-end observability for AI applications and agents .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and implementing observability for AI applications, with strong capabilities in Python, CI/CD, and cloud-native platforms. Proficient in operational tooling, incident management, and compliance within AI environments.
Highest-signal resume keywords
Python ProgrammingObservability ImplementationCloud-Native PlatformsIncident ManagementStakeholder Management
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Platform EngineeringSite Reliability EngineeringDevOpsMLOpsInfrastructure as CodeCI/CDDistributed TracingService-Level ObjectivesMachine Learning OperationsGenerative AI
Soft Skills
Technical CommunicationDocumentation SkillsStakeholder Management
Tools & Technologies
TerraformGitHub ActionsAzure DevOpsOpenTelemetryAzure MonitorApplication InsightsMicrosoft FabricLangfuseLangSmithMLflow
Industry Keywords
AI ApplicationsOperational ReadinessPrivacy by DesignAudit ControlsRegulated IndustriesFinOps Practices
Tech Stack
Tools & technologiesAzureCloudCyber SecurityPythonTerraform
About the role
Key responsibilities & impact- Design and implement end-to-end observability for AI applications and agents
- Establish common telemetry and distributed tracing across agent workflows, APIs, data services, vector stores, model endpoints, and external tools
- Build dashboards and alerts for availability, latency, errors, reliability, quality, policy violations, consumption, and service health
- Define service-level indicators, objectives, error budgets, alert thresholds, and operational readiness criteria
- Enable trace-based debugging, incident reconstruction, and controlled session replay while protecting sensitive information
- Monitor retrieval quality, data freshness, model and prompt regressions, anomalous agent loops, degraded tool performance, and unexpected runtime behavior
- Lead root-cause analysis for AI platform and runtime incidents and implement preventive controls and improvements
- Engineer reusable pipelines, templates, and controls for configuration, prompt, model, and agent-component versioning
- Implement automated deployment-readiness gates, integration tests, regression checks, security checks, and observability coverage
- Enable canary releases, feature flags, model/provider routing, fallback strategies, and rollback mechanisms
- Provide operational tooling supporting multiple models, frameworks, clouds, and agent patterns
- Maintain runbooks, reference implementations, engineering standards, and production support patterns
- Create metering, allocation, showback, budgets, thresholds, anomaly alerts, and runtime guardrails for AI consumption
- Partner with product teams to optimize model selection, routing, context size, caching, batching, retries, and tool usage
- Embed security, privacy, Responsible AI, compliance, identity, access control, secrets management, and auditability into the AI runtime
- Engineer controls for prompt injection, jailbreaks, unauthorized tool use, excessive permissions, data leakage, unsafe execution paths, and anomalous access
- Define operating processes for monitoring, support, incident, problem, change, escalation, and service recovery management
- Create production-readiness checklists, acceptance criteria, on-call runbooks, escalation paths, recovery procedures, and business-continuity requirements
- Drive automation to reduce reconstruction effort and shorten detection, diagnosis, containment, and recovery times
- Produce technical documentation and coach engineering teams on approved operational patterns
- Act as technical bridge across AI, platform, cloud, data, architecture, cybersecurity, privacy, compliance, and service management functions
- Define and monitor KPIs for reliability, observability, incident performance, cost efficiency, control coverage, audit evidence, and platform adoption
- Use operational evidence to prioritize automation, reliability improvements, cost optimization, and risk reduction
Requirements
What you’ll need- Bachelor's or Master's degree in Computer Science, Engineering, Artificial Intelligence, or a related discipline, or equivalent relevant experience
- 5+ years of hands-on experience in Platform Engineering, Site Reliability Engineering, DevOps, DevSecOps, MLOps, or production software engineering, including responsibility for live services
- Practical experience operating Machine Learning, Generative AI, or agentic systems in production
- Hands-on capability in Python, APIs, infrastructure as code, and CI/CD using technologies such as Terraform, GitHub Actions, or Azure DevOps
- Experience with cloud-native platforms, containers, orchestration, and enterprise cloud services; Azure experience strongly valued
- Experience implementing observability using logs, metrics, distributed traces, dashboards, and alerts; OpenTelemetry or equivalent standards ideally
- Strong understanding of LLM application patterns including RAG, vector stores, model gateways, tool calling, agent memory, and multi-agent orchestration
- Working knowledge of identity and access management, secrets management, secure logging, threat modelling, privacy by design, and audit controls
- Ability to turn ambiguous operational and control requirements into reusable technical capabilities, standards, automation, and runbooks
- Strong stakeholder management, technical communication, and documentation skills
- Preferred: experience in regulated industries with formal security, privacy, quality, validation, risk, or audit requirements
- Preferred: hands-on experience with Azure OpenAI, Azure AI Foundry, Azure Monitor, Application Insights, Microsoft Fabric, or comparable cloud services
- Preferred: experience with AI observability or evaluation platforms such as Langfuse, LangSmith, Arize, MLflow, or equivalent technologies
- Preferred: experience operating heterogeneous model providers, multi-cloud services, or multiple agent frameworks
- Preferred: knowledge of FinOps practices, cloud cost allocation, budgeting, and unit-economics measurement for AI workloads
- Preferred: experience defining service-level objectives, on-call models, incident management, and operational-readiness gates for enterprise platforms
Benefits
Comp & perks- Annual bonus with set target depending on pay grade/location, based on employee and company performance
- Vacation days
- Parental leave for a minimum of 12 weeks
- Bereavement leave
- Caregiver leave
- Volunteer leave
- Well-being reimbursement
- Financial, physical, and mental health programs
- Service anniversary and recognition awards
- Insurance plans for employees and, in some locations, eligible dependents