Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Runtalent

Senior Site Reliability Engineer

Runtalent

. Operate and evolve post-Go-Live environments .

Posted 10/8/2026full-timeRemote • BrazilSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in Site Reliability Engineering (SRE) practices, including monitoring, observability, and incident management, while ensuring high availability and performance of production services. Proficient in AWS and capable of automating operational processes to enhance system resilience.

Highest-signal resume keywords
Site Reliability Engineering (SRE)AWSMonitoring and ObservabilityIncident ManagementTroubleshooting and Root Cause Analysis

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Site Reliability Engineering (SRE)AWSMonitoringObservabilityIncident ManagementTroubleshootingRoot Cause AnalysisAutomationMetricsDistributed Tracing
Soft Skills
Analytical MindsetProblem-SolvingAdvanced English Proficiency
Industry Keywords
Production EnvironmentsAvailabilityPerformanceResilienceSLISLOSLA

Tech Stack

Tools & technologies
AWSGo

About the role

Key responsibilities & impact
  • Operate and evolve post-Go-Live environments
  • Monitor the availability, performance, and health of production services
  • Implement and enhance observability practices
  • Ensure that the platform is understandable, available, and recoverable

Requirements

What you’ll need
  • Solid experience as a Site Reliability Engineer (SRE), DevOps Engineer, or in a similar role
  • Proven experience in highly available production environments
  • Strong knowledge of AWS
  • Hands-on experience with monitoring, observability, and alerting
  • Knowledge of metrics, logs, and distributed tracing
  • Experience with incident management and response
  • Knowledge of capacity, performance, availability, and resilience
  • Experience automating operational processes
  • Knowledge of SRE practices, SLI, SLO, and SLA
  • Ability to perform troubleshooting and root cause analysis
  • Experience with applications and services handling real-time traffic
  • Hands-on, analytical, and problem-solving mindset
  • Advanced English proficiency