FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates strong operational ownership and production engineering capabilities, with hands-on experience in cloud environments such as AWS and Azure. Proficient in CI/CD practices, infrastructure-as-code, and automation of operational tasks, while maintaining high standards of security and documentation.
Highest-signal resume keywords
Production Engineering ExperienceAWS/Azure Cloud OperationsCI/CD and Git-Based DeliveryInfrastructure-as-Code (Terraform, CloudFormation)Automation of Operational Tasks
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Production SupportTroubleshootingConfiguration ManagementScripting (Python, TypeScript, Bash, PowerShell)API ManagementMonitoring and MetricsRoot-Cause AnalysisOperational PlaybooksIncident ResponseSecurity and Access Management
Soft Skills
Strong Written Documentation SkillsEffective Communication with StakeholdersCoordination Across TeamsProblem-SolvingTechnical Risk Explanation
Tools & Technologies
AWSAzureTerraformCloudFormationCI/CD ToolsObservability ToolingKubernetesMicrosoft EntraSalesforceDistributed Tracing
Industry Keywords
AI OperationsHyperautomationOperational StandardsService DependenciesSecure ConfigurationAuditabilityOperational ControlsIncident ManagementCloud NetworkingReusable Operational Tooling
Tech Stack
Tools & technologiesAWSAzureCloudJavaScriptKubernetesPythonTerraformTypeScript
About the role
Key responsibilities & impact- Own the operational health of AI-powered workflows, applications and agents
- Establish support arrangements, dependencies and operational standards
- Make new and changed services production-ready before go-live
- Build and maintain repeatable deployment and environment patterns using CI/CD, infrastructure-as-code and configuration management
- Configure cloud services, identity, secrets and connectivity securely
- Create and maintain logs, metrics, dashboards, alerts and health checks
- Use telemetry, incidents and recurring issues to improve resilience, error handling, recovery and automation
- Act as a technical responder for production incidents
- Troubleshoot issues across cloud, application, identity, integration and workflow layers
- Coordinate service restoration with internal teams and vendors
- Contribute to root-cause analysis and corrective actions
- Maintain visibility over APIs, connectors, authentication and external platform dependencies
- Produce and maintain technical documentation, runbooks, support procedures, recovery processes and operational playbooks
- Implement security, access, auditability, data-handling and AI operational controls
- Coordinate technical support and escalations with AWS, Microsoft, Salesforce, Anthropic and Mindflow
- Assess platform changes affecting reliability, security or supportability
- Feed operational learning back to the Hyperautomation Lead, AI Engineers and wider Technology teams
- Improve reusable operational patterns and reduce recurring support effort and production risk
Requirements
What you’ll need- Strong production engineering and operational ownership experience
- Hands-on experience operating applications or services in AWS, Azure or a comparable cloud environment, including configuration, troubleshooting, access, monitoring and production support
- Practical experience with CI/CD and Git-based delivery practices
- Experience with infrastructure-as-code such as Terraform, CloudFormation, CDK or equivalent
- Experience with environment and configuration management
- Experience automating repeatable operational tasks
- Practical programming or scripting experience using Python, TypeScript/JavaScript, Bash, PowerShell or similar
- Ability to automate tasks and diagnose production issues
- Good working knowledge of APIs, HTTP, JSON, authentication and system integrations
- Experience working with production logs, metrics, dashboards and alerts
- Experience troubleshooting live services
- Experience responding to and coordinating production incidents
- Experience contributing to root-cause analysis and corrective actions
- Understanding of identity and access management, secrets and credential management, least-privilege access, secure configuration, auditability and operational controls
- Experience taking responsibility for how production services are operated and supported
- Experience creating and maintaining runbooks, operational playbooks, recovery and troubleshooting procedures, deployment and support documentation, service dependencies and escalation paths
- Strong written documentation skills
- Ability to work effectively with engineers, infrastructure and security teams, business stakeholders and external technology providers
- Ability to coordinate technical resolution across multiple teams
- Ability to work with vendors during incidents or technical escalations
- Ability to explain technical risks and operational issues clearly
- Ability to follow issues and corrective actions through to resolution
- AI operations experience, workflow platforms, model platforms, Microsoft Entra, AWS IAM, Salesforce, observability tooling, containers/Kubernetes, distributed tracing, SRE practices, MCP, RAG, agentic AI, secure cloud networking or reusable operational tooling are nice to have, not required
Benefits
Comp & perks- Equal opportunity employer welcoming differences and people from all backgrounds, cultures and experiences
- Environment where employees can achieve their full potential and do interesting and meaningful work
- Support available throughout the interview process
- Opportunity to share in the company’s success
- Purpose-driven, high-performing culture
- Interesting and meaningful work
