Senior Site Reliability Engineer
New
C
ClimavisionWeather technology
Fully Remote - United StatesFull-TimeSenior
Salary130,000 - 170,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- Minimum of 7 years of experience in Site Reliability Engineering, DevOps, Production Engineering, Platform Engineering, or a related infrastructure-focused role, with at least 4 years in a role formally titled Site Reliability Engineer or carrying explicit SLO / error-budget accountability.
- Required Skills
- KubernetesMicrosoft AzureGrafanaPrometheusTerraformAnsibleGitHub ActionsDatadog
Requirements
- Have a bachelor’s degree in computer science, software engineering, or a related field, or equivalent professional experience.
- Bring at least 7 years of experience in SRE, DevOps, Production Engineering, Platform Engineering, or related infrastructure work.
- Have at least 4 years in a formally titled SRE role or a role with explicit SLO or error-budget accountability.
- Have deep hands-on experience operating native Kubernetes; self-managed Kubernetes experience is strongly preferred.
- Demonstrate Kubernetes cluster optimization, including workload and node-pool right-sizing, resource management, and infrastructure cost reduction.
- Have experience building dashboards, metrics pipelines, and alerting where little or no observability existed.
- Have experience designing horizontally scalable workloads across multiple replicas, including idempotency, concurrency, and state handling.
- Have experience with multi-cluster high-availability architectures, including failover, traffic routing, and cross-cluster deployment.
- Have production incident-response experience supporting customer-facing systems.
- Have Kubernetes experience in bare-metal, colocation, edge, or hybrid infrastructure environments.
- Understand infrastructure automation and Infrastructure as Code, using tools such as Terraform and Ansible.
- Have experience with CI/CD and production deployment pipelines; GitHub Actions is used by the team.
Responsibilities
- Own production reliability for customer-facing platform and data services across Azure, colocation, and edge Kubernetes environments.
- Build and maintain fleet-wide observability dashboards, metrics, logging, tracing, and alerting.
- Design automated recovery and self-healing for production systems.
- Optimize cluster resources and costs, including right-sizing workloads and nodes and supporting workload migration off Azure.
- Coordinate incident response, troubleshooting, mitigation, communication, and postmortem analysis.
- Drive multi-replica and multi-cluster high availability, including failover behavior, traffic routing, and data replication considerations.
- Operate and improve Kubernetes platforms, including upgrades, patching, cluster health, node management, and production changes.
- Maintain deployment pipelines, Helm charts, Kubernetes manifests, and infrastructure automation.
- Conduct performance engineering and capacity planning for customer-facing services during peak weather-event demand.
- Improve disaster recovery, failover, business continuity, and operational maturity across cloud, colocation, and edge environments.
View Full Description & ApplyYou'll be redirected to the employer's site