Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
GrooveTech

Senior SRE Engineer – Observability & Reliability

GrooveTech

. Lead the administration and evolution of Datadog as the central observability platform, including APM, traces, log pipelines, metrics, RUM, Synthetics, SLOs, monitors, and Cloud SIEM .

Posted 10/7/2026contractRemote • BrazilSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates advanced expertise in Datadog and AWS architecture, ensuring comprehensive observability and compliance in production environments. Capable of leading incident response, creating technical documentation, and managing observability disciplines effectively.

Highest-signal resume keywords
Datadog Platform ExpertiseAWS Architecture KnowledgeIncident Management ExperienceObservability As Code PracticesPCI DSS Compliance Support

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Site Reliability Engineering (SRE)Application Performance Monitoring (APM)Log ManagementStructured Technical DocumentationError Budget TrackingSignal-Focused Alerts DesignCloud SIEM IntegrationRunbooks and Playbooks CreationObservability FinOps ManagementGo and/or TypeScript/Node.js Code Analysis
Soft Skills
Excellent AdaptabilityInterpersonal SkillsTransparent CollaborationFocus on Continuous ImprovementAbility to Execute Quickly
Tools & Technologies
AWS CloudWatchCloudTrailGuardDutyInspectorMacieTerraformJira Service ManagementEventBridgeSQSDynamoDB
Industry Keywords
FintechPaymentsPCI DSSOperational TroubleshootingChange Management

Tech Stack

Tools & technologies
AWSCloudDynamoDBJavaScriptNode.jsTerraformTypeScriptGo

About the role

Key responsibilities & impact
  • Lead the administration and evolution of Datadog as the central observability platform, including APM, traces, log pipelines, metrics, RUM, Synthetics, SLOs, monitors, and Cloud SIEM
  • Ensure comprehensive observability coverage for production environments on AWS, including ECS, Lambda, API Gateway, DynamoDB, EventBridge/SQS, CloudFront, and WAF
  • Define and track SLIs and SLOs for critical journeys, monitoring error budget consumption
  • Design signal-focused alerts by eliminating false positives, duplicate notifications, and inefficient monitors
  • Integrate and audit security logs, including CloudTrail, GuardDuty, Inspector, Macie, and VPC Flow Logs
  • Lead critical incident response, including triage, event correlation, transparent communication, and leadership of RCAs and postmortems
  • Maintain technical governance by creating runbooks and playbooks and supporting Change Management processes in Jira Service Management
  • Support PCI DSS compliance by producing monitoring, logging, and FIM evidence, as well as supporting audits and penetration testing
  • Manage Observability FinOps by optimizing Datadog costs related to retention, log indexing, and metric cardinality
  • Serve as the technical reference and owner for platform observability and reliability

Requirements

What you’ll need
  • Strong, hands-on track record (8+ years) in SRE, Observability, or the support of critical environments
  • Deep, advanced expertise in the Datadog platform, including APM, Traces, Logs, Synthetics, SLOs, and Dashboards
  • Strong knowledge of AWS architecture focused on operational troubleshooting, including CloudWatch, CloudTrail, ECS, Lambda, queues, and events
  • Experience managing incidents, problems, and changes (ITIL or equivalent)
  • Proven experience creating structured technical documentation, including runbooks, playbooks, and postmortems
  • Ability to execute, troubleshoot, and deliver quickly, without remaining limited to the theoretical or architectural level
  • Ability to independently manage the observability discipline and prioritize backlogs without the need for micromanagement
  • Excellent adaptability to dynamic routines, openness to feedback, and a focus on continuous improvement
  • Excellent interpersonal skills for aligned and transparent collaboration with engineering, QA, and product teams
  • Advanced to fluent English
  • Experience in payments, fintech, or PCI DSS-regulated environments
  • Knowledge of Cloud SIEM and AWS security tools, including GuardDuty, Inspector, and Macie
  • Observability as Code practices using Terraform and the Datadog API
  • Ability to read and analyze Go and/or TypeScript/Node.js code to support diagnostics

Benefits

Comp & perks
  • Life insurance
  • Access to the Wellhub ecosystem (formerly Gympass)
  • Support from technical and management back-office teams to ensure accelerated onboarding and continued support throughout your journey