FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in architecting and governing observability frameworks, including metrics, logs, traces, and events, while leading complex implementations and ensuring secure telemetry pipelines. Proficient in automating observability processes and integrating platforms within multi-cloud environments.
Highest-signal resume keywords
Observability Framework ArchitectureTelemetry Pipeline DevelopmentKubernetes and Docker ManagementIncident Management and RCACloud Platform Integration
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Telemetry Pipeline DevelopmentRoot Cause Analysis (RCA)Observability-as-CodeEvent CorrelationDynamic ThresholdsInfrastructure AutomationTime-Series Database ManagementProactive MonitoringHigh-Cardinality ControlsService Level Indicators (SLIs)
Soft Skills
Mentoring Engineering TeamsLeadership in Incident ManagementCross-Functional Collaboration
Tools & Technologies
IBM InstanaGrafanaPrometheusOpenTelemetryTelegrafInfluxDBServiceNowNetcoolCI/CD PipelinesAnsible
Certifications & Qualifications
CKA (Certified Kubernetes Administrator)Cloud Architect (AWS/Azure)
Industry Keywords
ObservabilityITSMMajor Incident ManagementCloud EnvironmentsLegacy Monitoring Migration
Tech Stack
Tools & technologiesAnsibleAWSAzureCitrixCloudDockerGoogle Cloud PlatformGrafanaITSMJenkinsKubernetesLinuxMicroservicesOpenShiftPrometheusPythonServiceNowSplunkTerraformVMware
About the role
Key responsibilities & impact- Architect and govern a unified observability framework covering metrics, logs, traces, and events
- Lead First-of-a-Kind (FOAK) implementations and convert new observability technologies into secure, repeatable, production-ready patterns
- Define enterprise standards for telemetry pipelines, data retention, high-cardinality controls, and observability cost management
- Define and govern Service Level Indicators (SLIs), Objectives (SLOs), and error budgets
- Serve as the senior technical escalation point, lead major P1/P2 incident war rooms, and conduct evidence-based Root Cause Analysis (RCA)
- Reduce MTTD/MTTR and alert noise through event correlation, dynamic thresholds, and dependency mapping
- Drive Observability-as-Code and infrastructure automation
- Automate deployment, configuration, and self-healing workflows for monitoring agents and telemetry collectors
- Integrate observability platforms with ITSM (ServiceNow), Netcool, and CI/CD pipelines
- Design observability for Docker, Kubernetes, microservices, and multi-cloud environments (Azure/AWS/GCP)
- Correlate application APM telemetry with Kubernetes control planes, pods, nodes, and infrastructure dependencies
- Ensure secure-by-design telemetry pipelines, including RBAC, TLS, secrets management, and image scanning
- Lead Knowledge Transfer (KT) programs, vendor transitions, and operational readiness handovers for global 24x7 teams
- Mentor cross-functional engineering teams and influence enterprise technology roadmaps
Requirements
What you’ll need- 12+ years of total IT experience, with a minimum of 5 to 7 years functioning as a Lead Architect, SRE, or Principal Observability Engineer in a massive enterprise environment
- Proven track record of migrating organizations from legacy monitoring to proactive, automated observability platforms
- Hands-on expertise in building scalable, secure telemetry pipelines and time-series databases
- Extensive experience leading FOAK rollouts and complex vendor/operations transition (KT) programs
- IBM Instana, Grafana (Enterprise & Alloy), Prometheus, OpenTelemetry, Telegraf, and InfluxDB
- SolarWinds, Netcool, Elastic/Splunk
- Kubernetes, Docker, OpenShift, AWS/Azure/GCP
- Linux (RHEL), Windows Server, VMware, Citrix VDI, load balancers, and edge proxies
- Ansible, Terraform, Python, Bash, Webhooks, CI/CD (GitHub Actions/GitLab/Jenkins)
- ServiceNow, ITIL 4, and advanced Major Incident Management
- Preferred certifications: CKA (Certified Kubernetes Administrator), Cloud Architect (AWS/Azure), or specific APM/Observability vendor certifications
Benefits
Comp & perks- CLT hiring - 40 hours weekly workload
- Work remotely
- Health Insurance - SulAmérica Prestige (available for legal dependents)
- Dental Insurance - SulAmérica (available for legal dependents)
- Life Insurance - Prudential (24x salary)
- Private Pension - Metlife (up to 6% of company match)
- Meal Voucher and Internet allowance - Flash (R$1.000,00 per month)
- Employee Assistance Program
- Wellness program
- Individual performance compensation program, depending on eligibility
- Equity grant under the Associate Equity Appreciation Program, depending on eligibility
