Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
CI&T

Site Reliability Engineer

CI&T

. Own the day-to-day operation of a monitoring platform, including dashboards, monitors, log pipelines, APM instrumentation, synthetic tests, and Real User Monitoring .

Posted 10/5/2026full-timeRemote • BrazilMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in site reliability engineering and DevOps practices, focusing on operational excellence, incident response, and system observability. Proficient in cloud infrastructure management and automation, with strong analytical skills for troubleshooting and performance optimization.

Highest-signal resume keywords
Site Reliability EngineeringCloud Infrastructure ManagementMonitoring Platforms AdministrationScripting Languages (Python, TypeScript, Go, Bash)Incident Response and Post-Incident Review

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Site Reliability EngineeringCloud Infrastructure ManagementMonitoring Platforms AdministrationScripting Languages (Python, TypeScript, Go, Bash)Troubleshooting and Root-Cause AnalysisInfrastructure as Code (Terraform, CloudFormation, Pulumi)CI/CD Tooling (GitHub Actions)Networking FundamentalsContainersLinux Fundamentals
Soft Skills
Clear Written CommunicationVerbal Communication
Tools & Technologies
New RelicGrafanaSplunkDynatraceAWSGCPAzure
Industry Keywords
APM InstrumentationLog ManagementError TrackingService Level IndicatorsCybersecurity Risk Remediation

Tech Stack

Tools & technologies
AWSAzureCloudCyber SecurityDNSGoogle Cloud PlatformGrafanaJavaScriptLinuxPythonSplunkTerraformTypeScriptGo

About the role

Key responsibilities & impact
  • Own the day-to-day operation of a monitoring platform, including dashboards, monitors, log pipelines, APM instrumentation, synthetic tests, and Real User Monitoring
  • Proactively analyze logs, error tracking, traces, and metrics to find failures, regressions, and anomalies; triage, reproduce, and drive issues to resolution
  • Tune alert thresholds, monitor logic, and notification routing to reduce noise and false positives while ensuring customer-impacting issues reach the right people
  • Instrument services with meaningful metrics, structured logs, and distributed traces
  • Partner with engineers to improve code observability
  • Write post-incident reviews, track remediation items, and feed lessons learned into monitors, runbooks, and system design
  • Define, measure, and report service level indicators for key customer-facing services
  • Improve reliability, scalability, and cost efficiency of cloud infrastructure, CI/CD pipelines, and release processes
  • Automate repetitive operational work
  • Maintain and improve runbooks, escalation paths, and operational documentation
  • Partner with enterprise InfoSec to remediate cybersecurity risks, including SSL/TLS cleanup, security headers, and DNS configuration hygiene
  • Track security findings from scans and audits through verified closure
  • Collaborate with engineering, QA, product, and security teams to build reliability and observability into new features

Requirements

What you’ll need
  • 3+ years of experience in site reliability engineering, DevOps, platform engineering, or a production-focused software engineering role
  • Hands-on experience administering and building in platforms such as New Relic, Grafana, Splunk, or Dynatrace, including dashboards, monitors, log management, and APM
  • Strong troubleshooting and root-cause analysis skills
  • Working proficiency in at least one scripting or programming language such as Python, TypeScript/JavaScript, Go, or Bash
  • Comfort reading application code to understand failures
  • Experience operating services in AWS, GCP, or Azure
  • Solid grasp of networking, containers, and Linux fundamentals
  • Familiarity with infrastructure as code such as Terraform, CloudFormation, or Pulumi
  • Familiarity with CI/CD tooling such as GitHub Actions
  • Experience with on-call responsibilities, incident response, and post-incident review processes
  • Clear written and verbal communication, including explaining reliability concerns and trade-offs to non-technical stakeholders

Benefits

Comp & perks
  • Health and dental insurance
  • Meal and food allowance
  • Childcare assistance
  • Extended paternity leave
  • Partnership with gyms and health and wellness professionals via Wellhub (Gympass) TotalPass
  • Profit Sharing and Results Participation (PLR)
  • Life insurance
  • Continuous learning platform (CI&T University)
  • Discount club
  • Free online platform dedicated to physical, mental, and overall well-being
  • Pregnancy and responsible parenting course
  • Partnerships with online learning platforms
  • Language learning platform
  • Inclusion support and accommodations during the selection process
  • Dedicated Health and Well-being team, inclusion specialists, and affinity groups