FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates advanced expertise in Datadog for observability strategy, including APM, Infrastructure Monitoring, and application instrumentation with OpenTelemetry. Capable of developing proactive monitoring strategies and collaborating effectively with cross-functional teams to enhance operational excellence.
Highest-signal resume keywords
Datadog Platform ImplementationAPM and Infrastructure MonitoringOpenTelemetry Application InstrumentationAWS Environment ExperienceAutomation Using Python or Shell Scripting
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Datadog APMInfrastructure MonitoringLogs ManagementDashboardsMonitoring Service CatalogDistributed TracingKubernetesDockerCI/CDDevSecOps
Soft Skills
Excellent Communication SkillsStakeholder Collaboration
Tools & Technologies
DatadogDynatraceGrafanaPrometheusElastic StackZabbix
Certifications & Qualifications
Datadog Certified Associate
Industry Keywords
ObservabilityMicroservices ArchitecturesSite Reliability EngineeringOperational GovernanceIncident Management
Tech Stack
Tools & technologiesAWSCloudDockerGrafanaKubernetesPrometheusPython
About the role
Key responsibilities & impact- Lead the evolution of the organization’s observability strategy, using Datadog as the primary enterprise platform
- Serve as the technical and functional subject-matter expert for the Datadog platform
- Design, implement, and enhance solutions using APM, Infrastructure Monitoring, Logs Management, Dashboards, RUM, Synthetic Monitoring, and Continuous Testing
- Define enterprise observability standards for applications, APIs, microservices, and cloud workloads
- Build and maintain executive, operational, and analytical dashboards
- Develop proactive monitoring strategies based on SLIs, SLOs, SLAs, and business indicators
- Create, review, and optimize intelligent monitors and alerts to reduce operational noise and false positives
- Support application instrumentation with OpenTelemetry and native Datadog integrations
- Lead root cause analyses of critical incidents and recommend preventive actions
- Identify opportunities for automation, failure prediction, and self-healing
- Develop training, playbooks, standards, and documentation related to Datadog
- Conduct capacity, performance, availability, and end-user experience analyses
- Promote a culture of observability and operational excellence
- Collaborate with Engineering, Architecture, Development, SRE, and Operations teams
Requirements
What you’ll need- Bachelor’s degree in Computer Science, Engineering, Information Systems, or a related field
- Advanced experience with Datadog, including platform implementation, administration, and enhancement
- Hands-on experience with Datadog APM, Infrastructure Monitoring, Logs Management, Dashboards, Monitors, and Monitoring Service Catalog
- Advanced knowledge of logs, metrics, and distributed tracing
- Experience with observability for APIs and microservices architectures
- Experience working with AWS environments
- Hands-on experience troubleshooting and investigating complex incidents
- Knowledge of OpenTelemetry and application instrumentation
- Knowledge of automation using Python, Shell scripting, or PowerShell
- Experience with Kubernetes, Docker, and cloud-native ecosystems
- Solid understanding of system availability, performance, scalability, and reliability
- Ability to translate technical indicators into business impact
- Excellent communication skills and the ability to work with multiple stakeholders
- Datadog Certified Associate certification or higher preferred
- Experience with Site Reliability Engineering (SRE) practices preferred
- Experience implementing observability strategies for large-scale distributed environments preferred
- Knowledge of CI/CD and DevSecOps preferred
- Experience with automated incident response and self-healing processes preferred
- Experience with operational governance, incident management, Problem Management, and ITIL processes preferred
- Experience with Dynatrace, Grafana, Prometheus, Elastic Stack, or Zabbix preferred
Benefits
Comp & perks- Remote work
- Inclusive, purpose-driven environment
- Opportunities to balance your career with personal commitments and interests
- Well-being initiatives
- Career opportunities and professional experiences
- Great Place To Work™ certification in 24 countries
- International Top Employers certification