Senior Site Reliability Engineer

New
C
ClimavisionWeather technology
Fully Remote - Australia, Standard local business hours on the east coast of Australia, providing overlap with the United States team in the early morning Eastern time.Full-TimeSenior
Salary135,000 - 170,000 USD per year
Apply NowOpens the employer's application page

Job Details

Experience
Minimum of 7 years of experience in Site Reliability Engineering, DevOps, Production Engineering, Platform Engineering, or a related infrastructure-focused role, with at least 4 years in a role formally titled Site Reliability Engineer or carrying explicit SLO / error-budget accountability.
Required Skills
KubernetesMicrosoft AzureGrafanaPrometheusTerraformAnsibleGitHub ActionsHelm

Requirements

  • Bachelor's degree in computer science, software engineering, or a related field; equivalent professional experience is considered.
  • At least 7 years of experience in SRE, DevOps, Production Engineering, Platform Engineering, or a related infrastructure-focused role.
  • At least 4 years in a role formally titled Site Reliability Engineer or carrying explicit SLO / error-budget accountability.
  • Deep hands-on experience operating native Kubernetes; native or self-managed Kubernetes experience is strongly preferred.
  • Experience optimizing Kubernetes clusters, including right-sizing workloads and node pools, resource management, and infrastructure cost reduction.
  • Experience building dashboards, metrics pipelines, and alerting, including in environments with little or no existing observability.
  • Experience designing workloads for horizontal scaling across multiple replicas and multi-cluster high-availability architectures.
  • Experience supporting customer-facing production systems, incident response, and structured on-call rotations.
  • Experience operating Kubernetes in bare-metal, colocation, edge, or hybrid infrastructure environments.
  • Strong understanding of infrastructure automation and Infrastructure as Code concepts using tools such as Terraform and Ansible.
  • Experience supporting CI/CD and production deployment pipelines; GitHub Actions is used for CI/CD.
  • Experience with monitoring, logging, and observability platforms such as DataDog, Prometheus, Grafana, Loki, or OpenTelemetry.
  • Experience operating distributed systems and microservice-based architectures; working knowledge of Microsoft Azure.
  • Working familiarity with Jira, Confluence, and Microsoft Entra.

Responsibilities

  • Own production reliability for customer-facing platform and data services across Azure, colocation, and edge Kubernetes environments.
  • Build and maintain fleet observability dashboards, metrics, logging, distributed tracing, and alerting.
  • Design automated recovery and self-healing for production systems.
  • Define and improve SLIs, SLOs, alerting standards, and operational metrics.
  • Optimize cluster resources and costs, including right-sizing workloads and supporting workload migration off Azure.
  • Coordinate incident response, including troubleshooting, mitigation, communication, and postmortem analysis.
  • Drive multi-replica and multi-cluster high availability, including failover behavior, traffic routing, and graceful degradation.
  • Operate and improve self-managed Kubernetes platforms, including upgrades, patching, node management, and production change management.
  • Maintain deployment pipelines, Helm charts, Kubernetes manifests, and infrastructure automation.
  • Partner with software engineering teams on production readiness, deployment safety, reliability practices, and operational visibility.
View Full Description & ApplyYou'll be redirected to the employer's site
135,000 - 170,000 USD per year
Apply Now