Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Peraton

Technical Enterprise Incident Manager

Peraton

. Lead and coordinate incident bridge calls involving infrastructure, application, network, cloud, security, and vendor teams .

Posted 9/29/2026full-timeRemote • United StatesMid-LevelSenior💰 $86,000 - $138,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in Cloud Incident Management and Operations Engineering, with a strong focus on ITIL Incident and Problem Management processes. Proficient in utilizing monitoring tools like Datadog and ServiceNow to drive service restoration and improve operational efficiency.

Highest-signal resume keywords
Cloud Incident ManagementITIL Incident ManagementDatadog MonitoringServiceNow ITSMAWS Cloud Services

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Incident ManagementRoot Cause AnalysisTroubleshootingAutomation of Incident WorkflowsMonitoring and ObservabilityCapacity PlanningPerformance OptimizationSLA Compliance MonitoringAnalytical SkillsTechnical Documentation
Soft Skills
Excellent Communication SkillsOrganizational SkillsLeadershipFacilitation SkillsCollaboration
Tools & Technologies
DatadogServiceNowAWSCloudcraftWindows ServersLinux Servers
Certifications & Qualifications
Public Trust Clearance
Industry Keywords
Site Reliability EngineeringNOCProduction SupportFederal ComplianceHealthcare ComplianceFinancial Compliance

Tech Stack

Tools & technologies
AWSAzureCloudDNSFirewallsGoogle Cloud PlatformITSMLinuxServiceNow

About the role

Key responsibilities & impact
  • Lead and coordinate incident bridge calls involving infrastructure, application, network, cloud, security, and vendor teams
  • Drive rapid service restoration while maintaining accurate timelines, communications, and executive updates
  • Prioritize incidents based on business impact and operational risk
  • Manage escalation procedures and engage leadership when required
  • Monitor SLA compliance and incident response metrics
  • Facilitate Post-Incident Reviews and ensure Root Cause Analysis documentation and corrective-action follow-through
  • Drive automation of incident detection, triage, and response workflows
  • Maintain structured stakeholder communications during major incidents
  • Validate runbooks, service dependency maps, and technical documentation
  • Use log aggregation and APM tools for troubleshooting and incident analysis
  • Work with application teams to resolve issues and implement root-cause remediations
  • Develop and enhance monitoring, alerting, and dashboarding capabilities
  • Analyze trends, KPIs, and operational metrics to identify reliability risks
  • Support resiliency strategies including redundancy, failover, capacity planning, and performance optimization
  • Use Datadog for monitoring, alert correlation, dashboards, incident investigation, and performance analysis
  • Participate in after-hours on-call incident management rotation
  • Develop and maintain incident management procedures, runbooks, and knowledge articles
  • Ensure accurate ticket documentation in ServiceNow
  • Drive continual service improvement aligned with ITIL and SRE best practices
  • Collaborate with cross-functional teams to improve communication, escalation paths, and workflows
  • Support audit, compliance, and operational reporting requirements

Requirements

What you’ll need
  • Must be a U.S. citizen
  • Ability to obtain and maintain the required Public Trust level clearance
  • Bachelor's Degree and 5 years of experience, or a High School diploma or equivalent and 9 years of experience
  • 5+ years of experience in Cloud Incident Management, Operations Engineering, NOC, SRE, Application or Production Support environments
  • Experience leading enterprise Major Incident response efforts in a 24x7 operational environment
  • Strong understanding of ITIL Incident and Problem Management processes
  • 3+ years of experience working with AWS cloud services
  • Experience with monitoring and observability platforms such as Datadog, Cloudcraft, or similar
  • Experience using ServiceNow or similar ITSM platforms
  • Strong analytical, troubleshooting, and organizational skills
  • Excellent written and verbal communication skills with ability to facilitate meetings and brief technical teams and executive leadership
  • Preferred: Experience in a Site Reliability Engineering (SRE) or Cloud Platform DevOps environment
  • Preferred: Experience supporting federal, healthcare, financial, or other highly regulated environments
  • Preferred: Hands-on experience with Windows/Linux Servers, networking concepts, cloud platforms (AWS, Azure, or GCP), load balancers, proxies, DNS, and firewalls

Benefits

Comp & perks
  • Potential eligibility for overtime
  • Shift differential may be available
  • Discretionary bonus may be available