FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

DevOps/SRE Tech Lead – Observability
BigDataCorp. Design, deploy, and maintain telemetry infrastructure, including OpenTelemetry Collectors, metrics/logs/traces agents, aggregators, and visualization tools .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and maintaining observability infrastructure, including proficiency in AWS, Kubernetes, and IaC tools. Capable of leading technical teams, optimizing CI/CD pipelines, and translating technical requirements into business priorities.
Highest-signal resume keywords
AWS Infrastructure DesignKubernetes AdministrationObservability Stack DeploymentIaC ProficiencyCI/CD Pipeline Optimization
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
OpenTelemetry CollectorsPrometheus OperatorGrafana StackTerraformPythonGoBashGitLab CIGitHub ActionsIncident Response
Soft Skills
Servant LeadershipMentorshipNegotiationResilienceAgile Facilitation
Tools & Technologies
DockerHelmDatadogNew RelicDynatraceArgo CDFluxIstioLinkerdFinOps
Certifications & Qualifications
CKAAWS Solutions Architect
Industry Keywords
Observability-as-a-ServiceTelemetry Data ManagementSDLCCloud ArchitectureDevSecOps
Tech Stack
Tools & technologiesAWSDockerFluxGrafanaJenkinsKubernetesPrometheusPythonSDLCTerraformGo
About the role
Key responsibilities & impact- Design, deploy, and maintain telemetry infrastructure, including OpenTelemetry Collectors, metrics/logs/traces agents, aggregators, and visualization tools
- Ensure high availability of the observability platform
- Configure application and infrastructure data collection and define actionable alerting rules with Alertmanager/Incident.io
- Reduce noise by operating an Observability-as-a-Service model for development teams
- Design, operate, and evolve AWS infrastructure using IaC
- Administer production Kubernetes clusters, ensuring scalability, security, and cost control
- Build and optimize continuous delivery pipelines and eliminate repetitive operational work
- Lead the infrastructure/platform team, organize the backlog, assign tasks, and mentor engineers
- Negotiate technical priorities with development and product teams
Requirements
What you’ll need- Deep experience with public cloud providers, preferably AWS
- Advanced expertise in Kubernetes and the container ecosystem (Docker/containerd)
- Mature experience operating and supporting distributed, highly available production systems
- Proven experience deploying and configuring the observability stack through code (IaC/Helm)
- Strong command of Prometheus Operator, OpenTelemetry Collector, Grafana Stack (Loki, Tempo, Mimir), or Datadog, New Relic, and Dynatrace
- Experience managing telemetry data retention and costs
- Ability to design resilient and scalable cloud architectures
- Systems-level understanding of the SDLC to improve the developer experience (DevEx)
- Deep understanding of observability data generation, collection, processing, and storage
- Experience with incident response protocols
- Experience or strong aptitude for facilitating agile ceremonies and managing delivery for a small technical team
- Proficiency in IaC tools such as Terraform, Pulumi, or equivalent
- Demonstrated ability to code automation using Python, Go, or advanced Bash
- Experience building and maintaining CI/CD pipelines, such as GitLab CI, GitHub Actions, or Jenkins
- Ability to translate technical debt and infrastructure risks into business language and negotiate priorities
- Resilience and composure to make rapid decisions under pressure during incidents
- Mentorship-oriented profile (Servant Leadership)
- Pragmatic approach to balancing architecture with business needs
- Nice to have: Argo CD, Flux, Istio, or Linkerd; FinOps; eBPF (Cilium, Pixie); log sampling and aggregation/discarding; DevSecOps, OPA/Gatekeeper, and supply chain security; CKA or AWS Solutions Architect certifications, or equivalent
- Advanced technical English for reading, conversation, and leading forums with vendors and global teams
Benefits
Comp & perks- 100% remote work (work from wherever you want)
- Offices in Rio de Janeiro and São Paulo with a relaxed work environment (for those who wish to use them)
- Growth and development opportunities
- Flexible working hours (yes, even when working from home)
- Flexible vacation (in addition to statutory CLT leave)
- Bradesco health and dental insurance
- Group life insurance
- Meal benefit via Caju card
- Monthly home-office allowance (furniture and internet)
- Home-office benefit
- OrienteMe (psychological support)
- Conexa Saúde (psychological and nutritional support)
- Wellhub
- Birthday Day Off
- Six months of maternity leave
- One month of paternity leave
- Childcare allowance
- SESC membership
- Game night (in person at the Rio de Janeiro office and online via Discord)
- BDC Learning and BDC Decola (professional development and growth programs that subsidize various types of training, such as courses, talks, books, and more)
- A close-knit team that is always ready to help