FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Site Reliability Engineer, AUS
Climavision. Own production reliability for customer-facing platform and data services across Azure, colocation, and edge Kubernetes environments.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in Site Reliability Engineering with a focus on Kubernetes management, production incident response, and operational metrics. Proficient in infrastructure automation and observability, ensuring high availability and reliability of customer-facing platforms.
Highest-signal resume keywords
Site Reliability EngineeringKubernetes ManagementInfrastructure AutomationProduction Incident ResponseObservability and Monitoring
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
KubernetesInfrastructure as CodeTerraformAnsibleCI/CDMetrics PipelinesDashboardingDistributed SystemsPerformance EngineeringCapacity Planning
Soft Skills
Strong Troubleshooting SkillsWritten and Verbal Communication
Tools & Technologies
DataDogPrometheusGrafanaLokiOpenTelemetryGitHub ActionsJiraConfluenceMicrosoft AzureRancher
Industry Keywords
Production EngineeringPlatform EngineeringHigh AvailabilityIncident ResponseDisaster Recovery
Tech Stack
Tools & technologiesAnsibleAzureCloudDistributed SystemsGrafanaKubernetesNode.jsPrometheusTerraform
About the role
Key responsibilities & impact- Own production reliability for customer-facing platform and data services across Azure, colocation, and edge Kubernetes environments.
- Support the whole company through a shared SRE function across radar network and weather intelligence operations.
- Define and improve SLIs, SLOs, alerting standards, and operational metrics.
- Build shared observability dashboards and alerting for the full fleet.
- Design and build automated recovery and self-healing for production systems.
- Optimize cluster resources and costs, including right-sizing workloads and nodes and migrating workloads off Azure.
- Coordinate production incident response, including troubleshooting, mitigation, communication, and postmortem analysis.
- Diagnose and resolve complex issues across application services, Kubernetes infrastructure, storage, and distributed systems.
- Drive multi-replica and multi-cluster high availability, including workload placement, scheduling, failover, traffic routing, and data replication.
- Operate and improve self-managed Kubernetes across cloud-hosted, colocation, and edge clusters.
- Perform Kubernetes upgrades, patching, cluster health management, node management, and production change management.
- Improve observability, autoscaling, ingress, distributed storage, resiliency, and operational maturity.
- Design and validate Kubernetes workloads for resiliency, scalability, and operational efficiency.
- Partner with software engineering teams on production readiness, deployment safety, resiliency, and operational visibility.
- Maintain deployment pipelines, Helm charts, Kubernetes manifests, and infrastructure automation.
- Support metrics, logging, distributed tracing, dashboarding, and alerting platforms.
- Conduct performance engineering and capacity planning for peak weather-event demand.
- Facilitate blameless postmortems and complete operational follow-up actions.
- Improve disaster recovery, failover, and business continuity across cloud, colocation, and edge environments.
- Drive automation, toil reduction, game days, production readiness reviews, and reliability best practices.
- Serve as a senior technical resource and mentor.
- Participate in separate rotating weekday and weekend on-call schedules, approximately every five weeks each.
Requirements
What you’ll need- A bachelor's degree in computer science, software engineering, or a related field; equivalent professional experience considered.
- Minimum of 7 years of experience in Site Reliability Engineering, DevOps, Production Engineering, Platform Engineering, or a related infrastructure-focused role, with at least 4 years in a role formally titled Site Reliability Engineer or carrying explicit SLO / error-budget accountability.
- Deep, hands-on experience operating native Kubernetes; managed distributions such as AKS and EKS are acceptable.
- Experience optimizing Kubernetes clusters, including right-sizing workloads and node pools, resource management, and reducing infrastructure cost.
- Experience building dashboards, metrics pipelines, and alerting.
- Experience designing and operating workloads for safe horizontal scaling across multiple replicas, including idempotency, concurrency, and state handling.
- Experience designing or operating multi-cluster high-availability architectures, including failover behavior, traffic routing, and cross-cluster service deployment.
- Experience supporting customer-facing production systems with uptime, reliability, and incident-response responsibilities.
- Experience diagnosing and resolving production incidents across application, platform, and Kubernetes infrastructure layers.
- Experience operating Kubernetes outside strictly managed cloud environments, including bare-metal, colocation, edge, or hybrid infrastructure.
- Experience with Kubernetes tooling and ecosystem technologies such as Rancher, Helm, autoscaling frameworks, observability stacks, or distributed storage systems.
- Strong understanding of infrastructure automation and Infrastructure as Code using tools such as Terraform and Ansible.
- Experience supporting CI/CD and production deployment pipelines; GitHub Actions is used at Climavision.
- Experience with monitoring, logging, and observability platforms such as DataDog, Prometheus, Grafana, Loki, OpenTelemetry, or comparable technologies.
- Experience operating distributed systems and microservice-based architectures in production.
- Working knowledge of Microsoft Azure infrastructure.
- Strong troubleshooting skills across infrastructure, application, and platform layers.
- Experience participating in a structured production on-call rotation supporting business-critical systems.
- Working familiarity with Jira, Confluence, and Microsoft Entra.
- Strong written and verbal communication skills, including incident documentation and postmortem authoring.
- Experience working in start-up, scale-up, or other fast-moving engineering environments.
- Any offer of employment is contingent on completion of a background check to company standard.
Benefits
Comp & perks- Benefits of a dynamic and growing organization
- A challenging, hands-on role that will have real impact on the business
- Competitive compensation
- Comprehensive benefits package
- 401(k) Savings Plan
- Medical/Dental/Vision Benefits
- Health Savings Account (HSA) and Flexible Spending Account (FSA)
- Unlimited Paid Time-off
- 11 Paid Holidays
- Paid Parental Leave
- Company Paid Short-term Disability (STD)
- Company Paid Long-term Disability (LTD)
- Company Paid Life Insurance
- Rotating weekday and weekend on-call schedule with separate rotations