Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Verity Group

SRE Engineer

Verity Group

. Define and track SLIs, SLOs, SLAs, MTTR, and MTTD .

Posted 9/26/2026full-timeRemote • BrazilMid-LevelSeniorWebsite

Tech Stack

Tools & technologies
AnsibleAWSAzureCloudDockerElasticSearchGoogle Cloud PlatformGrafanaKubernetesLinuxPrometheusTerraform

About the role

Key responsibilities & impact
  • Define and track SLIs, SLOs, SLAs, MTTR, and MTTD
  • Implement observability, monitoring, alerting, and APM
  • Monitor latency, traffic, errors, saturation, availability, and performance
  • Work on incident prevention, identification, and resolution
  • Lead root cause analyses and define actions to prevent recurrence
  • Identify risks, bottlenecks, and single points of failure
  • Support the design of resilient, scalable, and highly available solutions
  • Automate operational activities and reduce manual tasks
  • Operate and evolve Kubernetes and Docker environments
  • Support capacity planning, business continuity, and disaster recovery strategies
  • Participate in deployments and support application stabilization
  • Collaborate with teams to improve reliability from the solution design stage
  • Create and maintain dashboards, alerts, procedures, and operational documentation
  • Promote a culture of reliability, observability, and continuous improvement

Requirements

What you’ll need
  • Experience as a Site Reliability Engineer, SRE, or in an equivalent role
  • Hands-on experience with cloud environments using GCP, AWS, and/or Azure
  • Knowledge of Kubernetes and Docker
  • Experience with observability, monitoring, alerting, and APM
  • Knowledge of SRE metrics and practices, such as SLI, SLO, SLA, MTTR, and MTTD
  • Experience managing, investigating, and resolving incidents
  • Knowledge of application and infrastructure troubleshooting
  • Experience administering Linux environments
  • Knowledge of networking, security, performance, and high availability
  • Experience with automation and Infrastructure as Code
  • Experience with CI/CD pipelines
  • Strong communication skills and the ability to work with cross-functional teams
  • Analytical, proactive, collaborative, and prevention-oriented mindset
  • Nice to have: experience with GKE, EKS, or AKS
  • Nice to have: knowledge of Dynatrace, Datadog, Grafana, Prometheus, or similar tools
  • Nice to have: experience with the ELK Stack, Elasticsearch, and Kibana
  • Nice to have: knowledge of Terraform and Ansible
  • Nice to have: experience with mission-critical environments and distributed systems
  • Nice to have: experience in financial institutions or regulated environments
  • Nice to have: experience with capacity management and cloud cost optimization
  • Nice to have: knowledge of disaster recovery and business continuity
  • Nice to have: experience defining and managing error budgets
  • Nice to have: certifications in Cloud, Kubernetes, or SRE

Benefits

Comp & perks
  • Meal voucher
  • Food allowance
  • Home office allowance
  • Health insurance
  • Dental insurance
  • Life insurance
  • Birthday Day Off
  • Total Pass / Wellhub app
  • Boon Saúde
  • Discount partnerships
  • Agreements with businesses and educational institutions
  • Welcome kit
  • Verity onboarding program
  • Verity Learning Interval
  • Great Place to Work certification and workplace improvement initiatives