FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Site Reliability Engineer
Climavision. Own production reliability for customer-facing platform and data services across Azure, colocation, and edge Kubernetes environments.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in Site Reliability Engineering with a strong focus on Kubernetes management, operational visibility, and incident response. Proficient in optimizing production systems for reliability and cost efficiency while ensuring high availability and performance.
Highest-signal resume keywords
Site Reliability EngineeringKubernetes ManagementOperational VisibilityIncident ResponseInfrastructure Automation
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
KubernetesProduction EngineeringInfrastructure as CodeCI/CD PipelinesMonitoring and LoggingPerformance EngineeringCapacity PlanningDisaster RecoveryHigh Availability ArchitecturesResource Management
Soft Skills
Strong Troubleshooting SkillsWritten and Verbal Communication
Tools & Technologies
TerraformAnsibleDataDogPrometheusGrafanaLokiOpenTelemetryGitHub ActionsJiraConfluence
Industry Keywords
AzureColocationEdge ComputingDistributed SystemsMicroservicesOperational MetricsSLOsSLIsIncident DocumentationPostmortem Analysis
Tech Stack
Tools & technologiesAnsibleAzureCloudDistributed SystemsGrafanaKubernetesNode.jsPrometheusTerraform
About the role
Key responsibilities & impact- Own production reliability for customer-facing platform and data services across Azure, colocation, and edge Kubernetes environments.
- Support the shared SRE function across the radar network and weather intelligence business.
- Define and improve SLIs, SLOs, alerting standards, and operational metrics.
- Build and own fleet observability through shared dashboards and proactive alerting.
- Design and build automated recovery and self-healing for production systems.
- Optimize cluster resources and costs, including right-sizing workloads and nodes and migrating workloads off Azure.
- Coordinate production incident response, including troubleshooting, mitigation, communication, and postmortem analysis.
- Diagnose and resolve complex issues across application services, Kubernetes infrastructure, storage, and distributed systems.
- Drive multi-replica and multi-cluster high availability, including workload placement, scheduling, failover, traffic routing, and data replication.
- Operate and improve self-managed Kubernetes platforms across cloud-hosted, colocation, and edge clusters.
- Execute Kubernetes upgrades, patching, cluster health, node management, and production change management.
- Improve observability, autoscaling, ingress, distributed storage, resiliency, and operational performance.
- Design and validate Kubernetes workloads for resiliency, scalability, efficiency, and graceful degradation.
- Partner with software engineering teams on production readiness, deployment safety, and operational visibility.
- Maintain deployment pipelines, Helm charts, Kubernetes manifests, and infrastructure automation.
- Support metrics, logging, distributed tracing, dashboarding, and alerting platforms.
- Conduct performance engineering and capacity planning for peak weather-event demand.
- Facilitate blameless postmortems and complete operational follow-up items.
- Improve disaster recovery, failover, and business continuity capabilities.
- Drive automation, toil reduction, game days, production readiness reviews, and reliability best practices.
- Mentor teams on reliability engineering and production operations practices.
- Participate in rotating weekday and weekend on-call schedules.
Requirements
What you’ll need- A bachelor's degree in computer science, software engineering, or a related field; equivalent professional experience considered.
- Minimum of 7 years of experience in Site Reliability Engineering, DevOps, Production Engineering, Platform Engineering, or a related infrastructure-focused role, with at least 4 years in a role formally titled Site Reliability Engineer or carrying explicit SLO / error-budget accountability.
- Deep, hands-on experience operating native Kubernetes; managed distributions such as AKS and EKS are acceptable, but native or self-managed Kubernetes is strongly preferred and is the primary technical requirement.
- Experience optimizing Kubernetes clusters, including right-sizing workloads and node pools, resource management, and reducing infrastructure cost without sacrificing reliability.
- Experience increasing operational visibility, including building dashboards, metrics pipelines, and alerting.
- Experience designing and operating workloads for safe horizontal scaling across multiple replicas, including idempotency, concurrency, and state handling.
- Experience designing or operating multi-cluster high-availability architectures, including failover behavior, traffic routing, and cross-cluster service deployment.
- Experience supporting customer-facing production systems with uptime, reliability, and incident-response responsibilities.
- Experience diagnosing and resolving production incidents across application, platform, and Kubernetes infrastructure layers.
- Experience operating Kubernetes outside strictly managed cloud environments, including bare-metal, colocation, edge, or hybrid infrastructure.
- Experience with Kubernetes operational tooling and ecosystem technologies such as Rancher, Helm, autoscaling frameworks, observability stacks, or distributed storage systems.
- Strong understanding of infrastructure automation and Infrastructure as Code using tools such as Terraform and Ansible.
- Experience supporting CI/CD and production deployment pipelines; GitHub Actions is used for CI/CD.
- Experience with monitoring, logging, and observability platforms such as DataDog, Prometheus, Grafana, Loki, OpenTelemetry, or comparable technologies.
- Experience operating distributed systems and microservice-based architectures in production.
- Working knowledge of Microsoft Azure infrastructure.
- Strong troubleshooting skills across infrastructure, application, and platform layers.
- Experience participating in a structured production on-call rotation supporting business-critical systems.
- Working familiarity with Jira, Confluence, and Microsoft Entra.
- Strong written and verbal communication skills, including incident documentation and postmortem authoring.
- Experience working in start-up, scale-up, or other fast-moving engineering environments.
- Ability to complete a background check to company standard.
- Willingness to participate in separate rotating weekday and weekend on-call schedules.
Benefits
Comp & perks- Benefits of a dynamic and growing organization
- A challenging, hands-on role that will have real impact on the business
- Competitive compensation
- Comprehensive benefits package
- 401(k) Savings Plan
- Medical/Dental/Vision Benefits
- Health Savings Account (HSA) and Flexible Spending Account (FSA)
- Unlimited Paid Time-off
- 11 Paid Holidays
- Paid Parental Leave
- Company Paid Short-term Disability (STD)
- Company Paid Long-term Disability (LTD)
- Company Paid Life Insurance