FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Site Reliability Engineering Technical Leader
Cisco. Lead feasibility assessment, technical planning, and phased migration of eligible workloads from AWS-hosted Kubernetes clusters to the internal Kubernetes platform .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates extensive experience in leading the migration and operational support of Kubernetes workloads, with a strong focus on performance, reliability, and incident response. Proficient in designing and implementing production-quality Go software for distributed systems while collaborating effectively across teams.
Highest-signal resume keywords
Kubernetes MigrationGo ProgrammingIncident ResponseLinux NetworkingObservability Tools
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
KubernetesGoDistributed SystemsIncident ResponsePerformance AnalysisAutomationOperational ReadinessLoad BalancingService DiscoveryTraffic Management
Soft Skills
Technical LeadershipStakeholder AlignmentSound Judgment
Tools & Technologies
AWSDockerGitLab CI/CDOpenTelemetryPrometheusDatadog
Industry Keywords
Site Reliability EngineeringInfrastructure EngineeringProduction SystemsNetwork SecurityPerformance Optimization
Tech Stack
Tools & technologiesAWSCloudDNSDockerGRPCKubernetesLinuxNode.jsPrometheusGo
About the role
Key responsibilities & impact- Lead feasibility assessment, technical planning, and phased migration of eligible workloads from AWS-hosted Kubernetes clusters to the internal Kubernetes platform
- Establish technical direction and staged delivery plans
- Support the operation and reliability of specialized Kubernetes workloads with non-standard load-balancing, networking, and traffic requirements
- Support the application after migration, including production readiness, incident response, performance improvements, and operational improvements
- Solve complex issues across applications, Kubernetes, Linux, networking, containers, and infrastructure
- Improve scalability, reliability, security, performance, and operability of Kubernetes-hosted services
- Partner with Node Connectivity, firmware, cloud infrastructure, security, SRE, and product teams
- Prioritize support and coordinate cross-system changes
- Contribute to broader Kubernetes SRE and platform reliability initiatives
- Design, implement, and maintain production-quality Go software for distributed, concurrent, and networked systems
- Drive operational readiness through SLIs/SLOs, comprehensive testing, on-call support, root-cause analysis, and long-term corrective actions
Requirements
What you’ll need- 10+ years of professional software, site reliability, or infrastructure engineering experience, including technical leadership of substantial production systems
- Experience designing, deploying, and operating large distributed services on Kubernetes
- 5+ years of programming experience in Go or a similar systems programming language
- Experience supporting production services through incident response, performance analysis, Kubernetes reliability practices, observability, automation, and operational readiness
- Knowledge of Linux and networking concepts, including IPv4/IPv6, TCP, routing, DNS, and TLS
- Sound judgment in architecture, incident response, prioritization, and technical tradeoffs
- Ability to align stakeholders across teams without relying on formal authority
- Preferred: Experience operating Kubernetes across AWS, on-premises, or hybrid environments using AWS/EKS, Docker, Kustomize, GitLab CI/CD, or similar deployment systems
- Preferred: Experience with Kubernetes networking, ingress, load balancing, service discovery, and traffic management for high-throughput or geographically distributed services
- Preferred: Experience with gRPC, Protocol Buffers, mutual TLS, PKI, VPNs, tunneling, or network security
- Preferred: Experience with OpenTelemetry, Prometheus, Datadog, or similar observability tools, distributed routing, packet processing, performance optimization, or failure testing
Benefits
Comp & perks- Medical, dental and vision insurance
- 401(k) plan with a Cisco matching contribution
- Paid parental leave
- Short- and long-term disability coverage
- Basic life insurance
- Cisco restricted stock unit grants may be available, subject to eligibility and continued employment
- 10 paid holidays per full calendar year
- 1 floating holiday for non-exempt employees
- Paid birthday day off
- Paid year-end holiday shutdown
- 4 paid personal wellness days
- 16 days of paid vacation per full calendar year for non-exempt employees
- Flexible vacation time off program with no defined limit for eligible exempt employees
- 80 hours of sick time on hire date and each January 1st thereafter
- Up to 80 hours of unused sick time carried forward
- Additional paid time away for critical or emergency family issues
- Optional 10 paid volunteer days per full calendar year
- Annual bonuses for non-sales roles, subject to Cisco policies
- Incentive compensation for sales-plan employees, subject to applicable Cisco plan