FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Site Reliability Engineering Team Lead – Principal SRE
Cerence Inc.. Lead Cerence's Site Reliability Engineering team and own the reliability, availability, and operational health of its cloud-native automotive AI platform .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in Site Reliability Engineering, leading teams to ensure the reliability and operational health of cloud-native platforms. Proficient in implementing SLI/SLO/SLA frameworks, CI/CD automation, and observability practices while effectively communicating with both technical and non-technical stakeholders.
Highest-signal resume keywords
Site Reliability Engineering LeadershipCloud Platform Experience (Azure, AWS, Google Cloud)Container Orchestration (Kubernetes, Docker, Istio)CI/CD Automation (Terraform, Flux)Observability Tooling (Zabbix, Prometheus, Grafana)
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Site Reliability EngineeringCloud Platform ManagementContainer OrchestrationCI/CD Pipeline DevelopmentScripting (Python, Go, Shell)UNIX/Linux System ConfigurationHigh-Availability Service DesignLog Aggregation and AnalyticsInfrastructure-as-Code PracticesPerformance Debugging
Soft Skills
Excellent Communication SkillsTeam LeadershipMentoringCollaborationProblem-Solving
Tools & Technologies
KubernetesDockerAzureAWSGoogle CloudZabbixPrometheusGrafanaTerraformJira
Industry Keywords
AutomotiveEmbedded SystemsLatency-Sensitive EnvironmentsITSMProject Management
Tech Stack
Tools & technologiesAWSAzureCloudDNSDockerFluxGrafanaITSMKubernetesLinuxPrometheusPythonSDLCTerraformUnixGo
About the role
Key responsibilities & impact- Lead Cerence's Site Reliability Engineering team and own the reliability, availability, and operational health of its cloud-native automotive AI platform
- Help select, mentor, and technically develop the team across multiple locations
- Set technical direction and priorities and contribute performance and growth input to team managers
- Design and maintain a sustainable on-call rotation and monitor page load and team health
- Own and drive the reliability roadmap across a 2–3 quarter horizon
- Define and govern SLI/SLO/SLA frameworks for customer program availability targets up to 99.95%
- Serve as Tier 2 technical escalation point for major incidents in partnership with the Global Operations Center
- Champion blameless postmortem culture and ensure actionable outcomes
- Lead and improve Production Readiness / NFR reviews with development teams
- Contribute to root cause analysis and own systemic improvements
- Approve high-risk and out-of-window production changes
- Set strategic direction for metrics, dashboards, alerting, escalation, and automation
- Drive CI/CD automation pipelines for service deployments, rollbacks, and operational tasks
- Partner with DevOps and platform teams to evolve shared infrastructure
- Embed reliability into the SDLC through collaboration with development managers and architects
- Participate in service reliability consulting and architectural reviews
- Communicate reliability posture and risk to technical and non-technical stakeholders
Requirements
What you’ll need- 8+ years of hands-on experience in site reliability, DevOps, or cloud platform roles, including time leading a team or owning a function
- A track record of setting technical direction and holding standards across a team — with or without formal authority
- Hands-on experience with container orchestration frameworks (Kubernetes, Docker, Istio)
- Experience with public cloud platforms (Azure primarily; AWS and Google Cloud)
- Familiarity with observability tooling — metrics pipelines, dashboarding, and alerting (e.g., Zabbix, Prometheus, Grafana)
- Experience with CI/CD pipelines and infrastructure-as-code practices (e.g., Terraform, Flux)
- Proficiency in at least one scripting or programming language (Python, Go, Shell, etc.)
- Strong UNIX/Linux background, including system configuration, performance debugging, and network fundamentals (Layer 4/5, DNS, HTTP/S, TLS)
- Excellent written and verbal communication skills in English
- Previous site reliability leadership experience
- Experience leading distributed or multi-site technical teams
- Background in high-availability service design (redundancy, failover, blast radius)
- Experience with log aggregation and analytics platforms (Loki, Thanos)
- Familiarity with ITSM and project tooling (Jira, Confluence)
- Experience in automotive, embedded, or latency-sensitive production environments
Benefits
Comp & perks- Annual bonus opportunity
- Insurance coverage (medical, dental, vision, life, and disability)
- Paid time off
- Paid holidays
- Company contribution to the RRSP (Registered Retirement Savings Plan)
- Equity awards for certain positions and levels
- Remote and/or hybrid work available depending on the position