Senior Site Reliability Engineer

New
C
ClimavisionWeather technology
Fully Remote - United StatesFull-TimeSenior
Salary130,000 - 170,000 USD per year
Apply NowOpens the employer's application page

Job Details

Experience
Minimum of 7 years of experience in Site Reliability Engineering, DevOps, Production Engineering, Platform Engineering, or a related infrastructure-focused role, with at least 4 years in a role formally titled Site Reliability Engineer or carrying explicit SLO / error-budget accountability.
Required Skills
KubernetesMicrosoft AzureGrafanaPrometheusTerraformAnsibleGitHub ActionsDatadog

Requirements

  • Have a bachelor’s degree in computer science, software engineering, or a related field, or equivalent professional experience.
  • Bring at least 7 years of experience in SRE, DevOps, Production Engineering, Platform Engineering, or related infrastructure work.
  • Have at least 4 years in a formally titled SRE role or a role with explicit SLO or error-budget accountability.
  • Have deep hands-on experience operating native Kubernetes; self-managed Kubernetes experience is strongly preferred.
  • Demonstrate Kubernetes cluster optimization, including workload and node-pool right-sizing, resource management, and infrastructure cost reduction.
  • Have experience building dashboards, metrics pipelines, and alerting where little or no observability existed.
  • Have experience designing horizontally scalable workloads across multiple replicas, including idempotency, concurrency, and state handling.
  • Have experience with multi-cluster high-availability architectures, including failover, traffic routing, and cross-cluster deployment.
  • Have production incident-response experience supporting customer-facing systems.
  • Have Kubernetes experience in bare-metal, colocation, edge, or hybrid infrastructure environments.
  • Understand infrastructure automation and Infrastructure as Code, using tools such as Terraform and Ansible.
  • Have experience with CI/CD and production deployment pipelines; GitHub Actions is used by the team.

Responsibilities

  • Own production reliability for customer-facing platform and data services across Azure, colocation, and edge Kubernetes environments.
  • Build and maintain fleet-wide observability dashboards, metrics, logging, tracing, and alerting.
  • Design automated recovery and self-healing for production systems.
  • Optimize cluster resources and costs, including right-sizing workloads and nodes and supporting workload migration off Azure.
  • Coordinate incident response, troubleshooting, mitigation, communication, and postmortem analysis.
  • Drive multi-replica and multi-cluster high availability, including failover behavior, traffic routing, and data replication considerations.
  • Operate and improve Kubernetes platforms, including upgrades, patching, cluster health, node management, and production changes.
  • Maintain deployment pipelines, Helm charts, Kubernetes manifests, and infrastructure automation.
  • Conduct performance engineering and capacity planning for customer-facing services during peak weather-event demand.
  • Improve disaster recovery, failover, business continuity, and operational maturity across cloud, colocation, and edge environments.
View Full Description & ApplyYou'll be redirected to the employer's site
130,000 - 170,000 USD per year
Apply Now