Senior Site Reliability Engineer
New
C
ClimavisionWeather technology
Fully Remote - Australia, Standard local business hours on the east coast of Australia, providing overlap with the United States team in the early morning Eastern time.Full-TimeSenior
Salary135,000 - 170,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- Minimum of 7 years of experience in Site Reliability Engineering, DevOps, Production Engineering, Platform Engineering, or a related infrastructure-focused role, with at least 4 years in a role formally titled Site Reliability Engineer or carrying explicit SLO / error-budget accountability.
- Required Skills
- KubernetesMicrosoft AzureGrafanaPrometheusTerraformAnsibleGitHub ActionsHelm
Requirements
- Bachelor's degree in computer science, software engineering, or a related field; equivalent professional experience is considered.
- At least 7 years of experience in SRE, DevOps, Production Engineering, Platform Engineering, or a related infrastructure-focused role.
- At least 4 years in a role formally titled Site Reliability Engineer or carrying explicit SLO / error-budget accountability.
- Deep hands-on experience operating native Kubernetes; native or self-managed Kubernetes experience is strongly preferred.
- Experience optimizing Kubernetes clusters, including right-sizing workloads and node pools, resource management, and infrastructure cost reduction.
- Experience building dashboards, metrics pipelines, and alerting, including in environments with little or no existing observability.
- Experience designing workloads for horizontal scaling across multiple replicas and multi-cluster high-availability architectures.
- Experience supporting customer-facing production systems, incident response, and structured on-call rotations.
- Experience operating Kubernetes in bare-metal, colocation, edge, or hybrid infrastructure environments.
- Strong understanding of infrastructure automation and Infrastructure as Code concepts using tools such as Terraform and Ansible.
- Experience supporting CI/CD and production deployment pipelines; GitHub Actions is used for CI/CD.
- Experience with monitoring, logging, and observability platforms such as DataDog, Prometheus, Grafana, Loki, or OpenTelemetry.
- Experience operating distributed systems and microservice-based architectures; working knowledge of Microsoft Azure.
- Working familiarity with Jira, Confluence, and Microsoft Entra.
Responsibilities
- Own production reliability for customer-facing platform and data services across Azure, colocation, and edge Kubernetes environments.
- Build and maintain fleet observability dashboards, metrics, logging, distributed tracing, and alerting.
- Design automated recovery and self-healing for production systems.
- Define and improve SLIs, SLOs, alerting standards, and operational metrics.
- Optimize cluster resources and costs, including right-sizing workloads and supporting workload migration off Azure.
- Coordinate incident response, including troubleshooting, mitigation, communication, and postmortem analysis.
- Drive multi-replica and multi-cluster high availability, including failover behavior, traffic routing, and graceful degradation.
- Operate and improve self-managed Kubernetes platforms, including upgrades, patching, node management, and production change management.
- Maintain deployment pipelines, Helm charts, Kubernetes manifests, and infrastructure automation.
- Partner with software engineering teams on production readiness, deployment safety, reliability practices, and operational visibility.
View Full Description & ApplyYou'll be redirected to the employer's site