FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in incident management and site reliability engineering, with a strong focus on performance monitoring, root cause analysis, and effective communication under pressure. Proficient in utilizing observability tools and managing high-pressure situations to ensure service reliability and operational efficiency.
Highest-signal resume keywords
Incident ManagementSite Reliability Engineering (SRE)Datadog MonitoringRoot Cause AnalysisAzure Cloud Platform
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Incident ManagementRoot Cause AnalysisPerformance MonitoringProblem-SolvingTask ExecutionLinux EnvironmentsNetworking ConceptsSystem ArchitectureCloud PlatformsServiceNow
Soft Skills
Strong CommunicationNegotiation SkillsAttention to DetailAnalytical MindsetAbility to Manage High-Pressure Situations
Tools & Technologies
DatadogGrafanaPagerDutyJIRADatabricksGitHub ActionsDockerKubernetes
Certifications & Qualifications
Bachelor's Degree
Industry Keywords
Incident ResponseService Level Agreements (SLAs)Mean Time to MitigateOperational BottlenecksPost-Incident Reviews
Tech Stack
Tools & technologiesAzureCloudDockerGoogle Cloud PlatformGrafanaKubernetesLinuxServiceNow
About the role
Key responsibilities & impact- Own the end-to-end lifecycle of major P1 and P2 incidents within a 16x7 shift window
- Ensure incident response and resolution milestones adhere to corporate SLAs
- Drive technical triage using telemetry data, system metrics, monitors, and logs
- Analyze telemetry data, consumer lags, and pipeline bottlenecks to guide engineering teams toward fixes
- Monitor performance anomalies, queue lags, and throughput drops to mitigate service degradation
- Use observability and monitoring platforms such as Datadog and Grafana
- Maintain clear, concise, and assertive communication under pressure
- Broadcast business-focused impact statements and progress metrics to executive leadership and client-facing teams
- Keep technical communication channels separate from high-level notification streams
- Perform Root Cause Analysis and facilitate blameless Post-Incident Reviews
- Identify operational bottlenecks and optimize playbooks to improve Mean Time to Mitigate
- Maintain structured handoffs between EMEA and APAC regions
Requirements
What you’ll need- Experience in SRE / Incident Management / Production Support
- Strong communication and negotiation skills
- Ability to manage high-pressure situations confidently
- Strong problem-solving and analytical mindset
- Technical knowledge of task execution
- Attention to detail and understanding of workflows
- Ability to identify risks using Datadog and Grafana monitoring tools
- Experience restoring services through restart, patching, or remediation of live issues
- Experience supporting on-premises Linux environments and cloud platforms, primarily Azure with some GCP exposure
- Understanding of networking concepts and system architecture
- Understanding of incident impact and ability to analyze and make decisions accordingly
- Good understanding of ServiceNow, PagerDuty, JIRA, Databricks, GitHub Actions, Docker, and Kubernetes
- Understanding of monitoring systems and ability to troubleshoot root causes
- Bachelor's Degree
