Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Climavision

Senior Site Reliability Engineer, AUS

Climavision

. Own production reliability for customer-facing platform and data services across Azure, colocation, and edge Kubernetes environments.

Posted 9/29/2026full-timeRemote • AustraliaSenior💰 $130,000 - $170,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in Site Reliability Engineering with a focus on Kubernetes management, production incident response, and operational metrics. Proficient in infrastructure automation and observability, ensuring high availability and reliability of customer-facing platforms.

Highest-signal resume keywords
Site Reliability EngineeringKubernetes ManagementInfrastructure AutomationProduction Incident ResponseObservability and Monitoring

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
KubernetesInfrastructure as CodeTerraformAnsibleCI/CDMetrics PipelinesDashboardingDistributed SystemsPerformance EngineeringCapacity Planning
Soft Skills
Strong Troubleshooting SkillsWritten and Verbal Communication
Tools & Technologies
DataDogPrometheusGrafanaLokiOpenTelemetryGitHub ActionsJiraConfluenceMicrosoft AzureRancher
Industry Keywords
Production EngineeringPlatform EngineeringHigh AvailabilityIncident ResponseDisaster Recovery

Tech Stack

Tools & technologies
AnsibleAzureCloudDistributed SystemsGrafanaKubernetesNode.jsPrometheusTerraform

About the role

Key responsibilities & impact
  • Own production reliability for customer-facing platform and data services across Azure, colocation, and edge Kubernetes environments.
  • Support the whole company through a shared SRE function across radar network and weather intelligence operations.
  • Define and improve SLIs, SLOs, alerting standards, and operational metrics.
  • Build shared observability dashboards and alerting for the full fleet.
  • Design and build automated recovery and self-healing for production systems.
  • Optimize cluster resources and costs, including right-sizing workloads and nodes and migrating workloads off Azure.
  • Coordinate production incident response, including troubleshooting, mitigation, communication, and postmortem analysis.
  • Diagnose and resolve complex issues across application services, Kubernetes infrastructure, storage, and distributed systems.
  • Drive multi-replica and multi-cluster high availability, including workload placement, scheduling, failover, traffic routing, and data replication.
  • Operate and improve self-managed Kubernetes across cloud-hosted, colocation, and edge clusters.
  • Perform Kubernetes upgrades, patching, cluster health management, node management, and production change management.
  • Improve observability, autoscaling, ingress, distributed storage, resiliency, and operational maturity.
  • Design and validate Kubernetes workloads for resiliency, scalability, and operational efficiency.
  • Partner with software engineering teams on production readiness, deployment safety, resiliency, and operational visibility.
  • Maintain deployment pipelines, Helm charts, Kubernetes manifests, and infrastructure automation.
  • Support metrics, logging, distributed tracing, dashboarding, and alerting platforms.
  • Conduct performance engineering and capacity planning for peak weather-event demand.
  • Facilitate blameless postmortems and complete operational follow-up actions.
  • Improve disaster recovery, failover, and business continuity across cloud, colocation, and edge environments.
  • Drive automation, toil reduction, game days, production readiness reviews, and reliability best practices.
  • Serve as a senior technical resource and mentor.
  • Participate in separate rotating weekday and weekend on-call schedules, approximately every five weeks each.

Requirements

What you’ll need
  • A bachelor's degree in computer science, software engineering, or a related field; equivalent professional experience considered.
  • Minimum of 7 years of experience in Site Reliability Engineering, DevOps, Production Engineering, Platform Engineering, or a related infrastructure-focused role, with at least 4 years in a role formally titled Site Reliability Engineer or carrying explicit SLO / error-budget accountability.
  • Deep, hands-on experience operating native Kubernetes; managed distributions such as AKS and EKS are acceptable.
  • Experience optimizing Kubernetes clusters, including right-sizing workloads and node pools, resource management, and reducing infrastructure cost.
  • Experience building dashboards, metrics pipelines, and alerting.
  • Experience designing and operating workloads for safe horizontal scaling across multiple replicas, including idempotency, concurrency, and state handling.
  • Experience designing or operating multi-cluster high-availability architectures, including failover behavior, traffic routing, and cross-cluster service deployment.
  • Experience supporting customer-facing production systems with uptime, reliability, and incident-response responsibilities.
  • Experience diagnosing and resolving production incidents across application, platform, and Kubernetes infrastructure layers.
  • Experience operating Kubernetes outside strictly managed cloud environments, including bare-metal, colocation, edge, or hybrid infrastructure.
  • Experience with Kubernetes tooling and ecosystem technologies such as Rancher, Helm, autoscaling frameworks, observability stacks, or distributed storage systems.
  • Strong understanding of infrastructure automation and Infrastructure as Code using tools such as Terraform and Ansible.
  • Experience supporting CI/CD and production deployment pipelines; GitHub Actions is used at Climavision.
  • Experience with monitoring, logging, and observability platforms such as DataDog, Prometheus, Grafana, Loki, OpenTelemetry, or comparable technologies.
  • Experience operating distributed systems and microservice-based architectures in production.
  • Working knowledge of Microsoft Azure infrastructure.
  • Strong troubleshooting skills across infrastructure, application, and platform layers.
  • Experience participating in a structured production on-call rotation supporting business-critical systems.
  • Working familiarity with Jira, Confluence, and Microsoft Entra.
  • Strong written and verbal communication skills, including incident documentation and postmortem authoring.
  • Experience working in start-up, scale-up, or other fast-moving engineering environments.
  • Any offer of employment is contingent on completion of a background check to company standard.

Benefits

Comp & perks
  • Benefits of a dynamic and growing organization
  • A challenging, hands-on role that will have real impact on the business
  • Competitive compensation
  • Comprehensive benefits package
  • 401(k) Savings Plan
  • Medical/Dental/Vision Benefits
  • Health Savings Account (HSA) and Flexible Spending Account (FSA)
  • Unlimited Paid Time-off
  • 11 Paid Holidays
  • Paid Parental Leave
  • Company Paid Short-term Disability (STD)
  • Company Paid Long-term Disability (LTD)
  • Company Paid Life Insurance
  • Rotating weekday and weekend on-call schedule with separate rotations