FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in observability and monitoring strategies, including the implementation of SLIs, SLOs, and automated incident response workflows. Proficient in managing observability stacks and ensuring compliance with monitoring data security standards.
Highest-signal resume keywords
Observability EngineeringDatadog ManagementIncident ManagementPython ProgrammingSRE Best Practices
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Distributed TracingSLI/SLO/SLASQLError BudgetsAnomaly DetectionMonitoring ToolsAutomationTrend AnalysisService Health ViewsRunbooks
Soft Skills
Clear CommunicationCollaborationCalm During OutagesBias To Action
Tools & Technologies
AWSITSM ToolsMonitoring FrameworksLogging Tools
Industry Keywords
SaaS ApplicationsEnterprise ApplicationsIncident FrameworksPerformance MetricsGovernance
Tech Stack
Tools & technologiesAWSDistributed SystemsITSMPythonSQL
About the role
Key responsibilities & impact- Define and implement standards for metrics, logs, traces, and profiling
- Establish golden signals, SLIs/SLOs, and health checks for priority services
- Automate baselining and anomaly detection
- Create executive and on-call dashboards, service health views, and dependency maps
- Develop alerting policy as code and reduce false positives through suppression and deduplication
- Ensure monitoring data security and compliance, including role-based access and guardrails
- Develop strategies to reduce high-priority incidents through trend analysis and proactive stability improvements
- Engineer automated response workflows and self-healing capabilities
- Design and coordinate comprehensive observability strategies across distributed systems
- Manage Datadog observability stacks and integrate them with ITSM tools and applications
- Onboard observability and monitoring for new applications handed over to operations
- Oversee monitoring tickets and lead monthly performance metric analysis and reporting
- Ensure SLA compliance and improve application monitoring maturity
- Automate incident reporting and alert correlation rules
- Manage L1 SRE Engineers handling automated monitoring incidents
- Manage governance calls, lead RCA reviews, and maintain the Known Error Database
- Partner with domain leads to define SLOs/SLIs, align capacity planning, and maintain runbooks and technical documentation
Requirements
What you’ll need- 6+ years of industry experience in Observability/SRE/Platform/Monitoring roles supporting SaaS or enterprise applications
- Hands-on experience with monitoring and logging tools
- Strong grasp of distributed tracing, RED/USE/golden signals, SLI/SLO/SLA, and error budgets
- Proficiency in common programming languages such as Python
- Experience with AWS
- Comfortable operating within Incident/Problem/Change frameworks
- Adept at runbooks, RCAs, and post-incident reviews
- SQL or log query language experience
- Clear communication, collaboration, bias to action, and calm during outages
- Hybrid role in India with core hours aligned to IST
- Occasional off-hours participation for major incidents or change windows
- Must be physically located and plan to work from Pune, India
- Must attend the local office for part of the week
Benefits
Comp & perks- Hybrid work model with flexibility to work remotely for part of the week
- On-call rotation with follow-the-sun support
- Opportunity to work with global teams
- Inclusive and diverse workplace
- Reasonable accommodations for applicants with disabilities and disabled veterans
