Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
American Bureau of Shipping (ABS)

Senior Manager, SRE, Operations, Product Support

American Bureau of Shipping (ABS)

. Define service health measures and reliability objectives for the Fleet Management System (FMS) .

Posted 10/2/2026full-timeUnited StatesSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in Site Reliability Engineering (SRE) practices, including incident response, operational readiness, and automation of operational workflows. Proficient in leveraging Azure cloud services and AI-assisted tools to enhance service reliability and operational efficiency.

Highest-signal resume keywords
Site Reliability Engineering (SRE)Incident Response LeadershipAzure Cloud OperationsOperational AutomationObservability and Monitoring

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Site Reliability EngineeringIncident ResponseOperational ReadinessAutomationAI-Assisted ToolsScriptingCI/CDInfrastructure as CodeRoot-Cause AnalysisDisaster Recovery
Soft Skills
Clear CommunicationCollaborative Leadership
Tools & Technologies
Azure MonitorApplication InsightsLog AnalyticsKey Vault
Industry Keywords
Fleet Management System (FMS)Production Web ServicesSaaS ApplicationsService Health MeasuresOperational Efficiency

Tech Stack

Tools & technologies
AzureCloudVault

About the role

Key responsibilities & impact
  • Define service health measures and reliability objectives for the Fleet Management System (FMS)
  • Establish end-to-end observability across applications, infrastructure, integrations, and AI-enabled services
  • Lead technical incident response, including severity definitions, on-call and escalation practices, incident coordination, recovery procedures, and post-incident reviews
  • Design for resilience, performance, capacity, backup and recovery, and safe operation under customer and data growth
  • Define and implement operational-readiness criteria for beta and production releases
  • Establish the technical support escalation model and partner with customer support, product, and engineering to resolve issues
  • Support customer migrations, go-lives, and post-launch stabilization
  • Automate routine operational tasks, health checks, deployment verification, incident triage, and recovery
  • Track reliability, incident, supportability, and operational-efficiency trends and communicate risks and progress to leadership
  • Build and mentor an SRE/production-operations capability as FMS scales
  • Report to the Sr Director, Platform Engineering & Cloud Architecture
  • Lead cross-functional operational practices; direct-report scope will expand as the capability scales

Requirements

What you’ll need
  • Bachelor’s degree in Computer Science, Engineering, Information Systems, or a related field, or equivalent relevant experience
  • 10+ years of relevant experience in site reliability engineering, production engineering, cloud operations, or software operations
  • Experience leading incident response or operational improvement across teams
  • Experience operating production web or SaaS services and improving reliability through software engineering and automation
  • Experience establishing observability, on-call practices, runbooks, and service-health or reliability measures for production systems
  • Experience partnering with software engineering, platform engineering, and customer-facing support teams during releases, incidents, and customer go-lives
  • Experience applying AI-assisted tools or workflows to technical operations, incident triage, monitoring analysis, support knowledge retrieval, or operational automation
  • Understanding of how to validate AI-assisted results before production use
  • Experience with Azure cloud preferred
  • Strong command of SRE practices, including SLIs/SLOs, incident response, root-cause analysis, performance, capacity, resilience, and disaster recovery
  • Experience operating production SaaS applications on Azure, including compute, networking, identity, storage, containers, databases, integrations, and security
  • Experience with Azure Monitor, Application Insights, and Log Analytics
  • Ability to automate secure deployments and operational workflows using scripting, APIs, CI/CD, infrastructure as code, managed identities, and Key Vault
  • Sound judgment on release risk, rollback, customer impact, responsible AI-assisted operations, data protection, and human oversight of production-impacting actions
  • Clear communication and collaborative leadership across engineering, support, and business stakeholders

Benefits

Comp & perks
  • Medical insurance (PPO and HD)
  • Dental and vision insurance
  • Health Savings Account (HSA)
  • Flexible Savings Account (FSA)
  • Life insurance
  • Accidental death and dismemberment insurance
  • Disability leave programs
  • Parental leave program
  • Paid holidays
  • Paid vacation time
  • Employee Assistance Plan (EAP) with personal wellness and work-life services
  • 401K plan with a generous company match, subject to plan requirements