FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in managing production operations with a focus on AWS environments, incident management, and automation strategies. Proven ability to lead technical teams and drive systemic improvements in operational health and reliability.
Highest-signal resume keywords
AWS Production EnvironmentsIncident ManagementTerraform Infrastructure-as-CodeAutomation/ScriptingObservability and Monitoring
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Production OperationsDevOpsSREAPIsETL/Data PipelinesSQLJob OrchestrationOperational TroubleshootingMTTA/MTTR ManagementService-Level Agreements
Soft Skills
LeadershipCollaborationCommunicationAccountabilityProblem-Solving
Tools & Technologies
TerraformCI/CD ToolingCloud Observability PlatformsWorkflow OrchestrationAI-Assisted Tooling
Industry Keywords
Healthcare DataFinancial ServicesSaaSRegulated Production DataCMS Submissions
Tech Stack
Tools & technologiesAWSCloudETLSQLTerraform
About the role
Key responsibilities & impact- Own the 24x7 operational health and reliability of CDM-managed production data services and workflows
- Serve as accountable leader and escalation owner for Sev-1 and Sev-2 production incidents
- Define SLAs/SLOs, escalation paths, on-call/coverage models, service-health measures, and operational performance expectations
- Drive service restoration, communication coordination, and corrective-action follow-through
- Own CDM incident and problem-management disciplines, including severity definitions, incident command, escalation, RCA, post-incident review, and corrective actions
- Track recurring failures and use MTTA, MTTR, availability, incident volume, and recurrence metrics to drive systemic improvement
- Provide technical leadership for production services operating in AWS and related enterprise environments
- Partner with Engineering on infrastructure-as-code using Terraform or comparable tooling, including repeatable configuration, deployment, and recovery
- Guide operational issues involving IAM, networking/connectivity, secure file transfer, storage, compute, logging, monitoring, and cloud dependencies
- Lead an automation-first strategy to eliminate repetitive manual work, fragile handoffs, and key-person dependencies
- Drive scripting, orchestration, automated validation, job recovery, exception handling, and self-healing patterns where appropriate
- Partner with Data Engineering on CI/CD, APIs, ETL/data pipelines, file movement, deployment/support patterns, and production automation
- Apply AI-assisted monitoring, troubleshooting, documentation, and workflow automation where appropriate
- Establish monitoring, logging, alerting, and operational dashboards that provide actionable visibility into production health
- Define production-readiness gates for workflows transitioning from implementation or engineering into Live Operations
- Require current runbooks, SOPs, recovery procedures, escalation paths, dependency maps, ownership, and cross-trained coverage before production handoff
- Define clear operating boundaries among Live Operations, Data Engineering, Data Management, Clinical Intelligence, Product, and IT
- Coordinate technical response across teams during production incidents and complex operational issues
- Build and lead a geographically distributed technical operations team with a culture of urgency, transparency, documentation, collaboration, and measurable improvement
Requirements
What you’ll need- 8+ years in production operations, cloud/platform operations, DevOps, SRE, operations engineering, data platform operations, or a comparable high-availability technical environment
- 4+ years leading technical production operations, DevOps, SRE, platform-support, or operations-engineering teams
- Demonstrated accountability for business-critical production systems operating under extended-hours or 24x7 support models
- Demonstrated leadership of Sev-1/Sev-2 or equivalent major incidents, including incident command, restoration, RCA, problem management, and corrective-action follow-through
- Strong working knowledge of AWS production environments, including IAM, networking/connectivity, logging/monitoring, cloud dependencies, and operational troubleshooting
- Demonstrated experience with Terraform or comparable infrastructure-as-code technologies and repeatable infrastructure deployment/recovery practices
- Strong technical understanding of APIs, ETL/data pipelines, secure file transfer, workflow/job orchestration, SQL, automation/scripting, and production integration patterns
- Demonstrated experience establishing observability, monitoring, alerting, runbooks, operational dashboards, and measurable reliability practices
- Demonstrated success automating manual production processes and reducing key-person dependencies through tooling, scripting, orchestration, or platform improvements
- Experience managing availability, SLA/SLO attainment, MTTA, MTTR, incident volume, recurring failures, and automation coverage
- Experience leading geographically distributed technical teams and effective on-call, escalation, and coverage models
- Ability to lead across Engineering, Product, IT, implementation, and customer-facing organizations during production incidents and reliability initiatives
- Preferred: Experience operating high-volume healthcare, financial-services, SaaS, or other regulated production data environments
- Preferred: Experience with MuleSoft, Flatfile, or comparable enterprise integration and data-ingestion platforms
- Preferred: Experience with CI/CD tooling, cloud observability platforms, workflow orchestration, and automated recovery patterns
- Preferred: Experience applying AI-assisted tooling to production support, incident analysis, documentation, or operational automation
- Preferred: Healthcare payer data, CMS submissions, EDI/X12, enrollment, claims, risk-adjustment, or related domain knowledge
Benefits
Comp & perks- Competitive salary
- Medical, Dental and Vision benefits
- 401k match
- Generous PTO plan
