Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Climavision

Senior Site Reliability Engineer

Climavision

. Own production reliability for customer-facing platform and data services across Azure, colocation, and edge Kubernetes environments.

Posted 9/29/2026full-timeRemote • United StatesSenior💰 $130,000 - $170,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in Site Reliability Engineering with a strong focus on Kubernetes management, operational visibility, and incident response. Proficient in optimizing production systems for reliability and cost efficiency while ensuring high availability and performance.

Highest-signal resume keywords
Site Reliability EngineeringKubernetes ManagementOperational VisibilityIncident ResponseInfrastructure Automation

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
KubernetesProduction EngineeringInfrastructure as CodeCI/CD PipelinesMonitoring and LoggingPerformance EngineeringCapacity PlanningDisaster RecoveryHigh Availability ArchitecturesResource Management
Soft Skills
Strong Troubleshooting SkillsWritten and Verbal Communication
Tools & Technologies
TerraformAnsibleDataDogPrometheusGrafanaLokiOpenTelemetryGitHub ActionsJiraConfluence
Industry Keywords
AzureColocationEdge ComputingDistributed SystemsMicroservicesOperational MetricsSLOsSLIsIncident DocumentationPostmortem Analysis

Tech Stack

Tools & technologies
AnsibleAzureCloudDistributed SystemsGrafanaKubernetesNode.jsPrometheusTerraform

About the role

Key responsibilities & impact
  • Own production reliability for customer-facing platform and data services across Azure, colocation, and edge Kubernetes environments.
  • Support the shared SRE function across the radar network and weather intelligence business.
  • Define and improve SLIs, SLOs, alerting standards, and operational metrics.
  • Build and own fleet observability through shared dashboards and proactive alerting.
  • Design and build automated recovery and self-healing for production systems.
  • Optimize cluster resources and costs, including right-sizing workloads and nodes and migrating workloads off Azure.
  • Coordinate production incident response, including troubleshooting, mitigation, communication, and postmortem analysis.
  • Diagnose and resolve complex issues across application services, Kubernetes infrastructure, storage, and distributed systems.
  • Drive multi-replica and multi-cluster high availability, including workload placement, scheduling, failover, traffic routing, and data replication.
  • Operate and improve self-managed Kubernetes platforms across cloud-hosted, colocation, and edge clusters.
  • Execute Kubernetes upgrades, patching, cluster health, node management, and production change management.
  • Improve observability, autoscaling, ingress, distributed storage, resiliency, and operational performance.
  • Design and validate Kubernetes workloads for resiliency, scalability, efficiency, and graceful degradation.
  • Partner with software engineering teams on production readiness, deployment safety, and operational visibility.
  • Maintain deployment pipelines, Helm charts, Kubernetes manifests, and infrastructure automation.
  • Support metrics, logging, distributed tracing, dashboarding, and alerting platforms.
  • Conduct performance engineering and capacity planning for peak weather-event demand.
  • Facilitate blameless postmortems and complete operational follow-up items.
  • Improve disaster recovery, failover, and business continuity capabilities.
  • Drive automation, toil reduction, game days, production readiness reviews, and reliability best practices.
  • Mentor teams on reliability engineering and production operations practices.
  • Participate in rotating weekday and weekend on-call schedules.

Requirements

What you’ll need
  • A bachelor's degree in computer science, software engineering, or a related field; equivalent professional experience considered.
  • Minimum of 7 years of experience in Site Reliability Engineering, DevOps, Production Engineering, Platform Engineering, or a related infrastructure-focused role, with at least 4 years in a role formally titled Site Reliability Engineer or carrying explicit SLO / error-budget accountability.
  • Deep, hands-on experience operating native Kubernetes; managed distributions such as AKS and EKS are acceptable, but native or self-managed Kubernetes is strongly preferred and is the primary technical requirement.
  • Experience optimizing Kubernetes clusters, including right-sizing workloads and node pools, resource management, and reducing infrastructure cost without sacrificing reliability.
  • Experience increasing operational visibility, including building dashboards, metrics pipelines, and alerting.
  • Experience designing and operating workloads for safe horizontal scaling across multiple replicas, including idempotency, concurrency, and state handling.
  • Experience designing or operating multi-cluster high-availability architectures, including failover behavior, traffic routing, and cross-cluster service deployment.
  • Experience supporting customer-facing production systems with uptime, reliability, and incident-response responsibilities.
  • Experience diagnosing and resolving production incidents across application, platform, and Kubernetes infrastructure layers.
  • Experience operating Kubernetes outside strictly managed cloud environments, including bare-metal, colocation, edge, or hybrid infrastructure.
  • Experience with Kubernetes operational tooling and ecosystem technologies such as Rancher, Helm, autoscaling frameworks, observability stacks, or distributed storage systems.
  • Strong understanding of infrastructure automation and Infrastructure as Code using tools such as Terraform and Ansible.
  • Experience supporting CI/CD and production deployment pipelines; GitHub Actions is used for CI/CD.
  • Experience with monitoring, logging, and observability platforms such as DataDog, Prometheus, Grafana, Loki, OpenTelemetry, or comparable technologies.
  • Experience operating distributed systems and microservice-based architectures in production.
  • Working knowledge of Microsoft Azure infrastructure.
  • Strong troubleshooting skills across infrastructure, application, and platform layers.
  • Experience participating in a structured production on-call rotation supporting business-critical systems.
  • Working familiarity with Jira, Confluence, and Microsoft Entra.
  • Strong written and verbal communication skills, including incident documentation and postmortem authoring.
  • Experience working in start-up, scale-up, or other fast-moving engineering environments.
  • Ability to complete a background check to company standard.
  • Willingness to participate in separate rotating weekday and weekend on-call schedules.

Benefits

Comp & perks
  • Benefits of a dynamic and growing organization
  • A challenging, hands-on role that will have real impact on the business
  • Competitive compensation
  • Comprehensive benefits package
  • 401(k) Savings Plan
  • Medical/Dental/Vision Benefits
  • Health Savings Account (HSA) and Flexible Spending Account (FSA)
  • Unlimited Paid Time-off
  • 11 Paid Holidays
  • Paid Parental Leave
  • Company Paid Short-term Disability (STD)
  • Company Paid Long-term Disability (LTD)
  • Company Paid Life Insurance