Senior Site Reliability Engineer

New
G
Garner HealthHealthtech
RemoteFull-TimeSenior
Salary$191,000 - $226,000
Apply NowOpens the employer's application page

Job Details

Experience
4+ years
Required Skills
AWSPythonKubernetesGoTerraformHIPAA

Requirements

  • 4+ years of hands-on experience operating production cloud infrastructure at scale in an SRE, DevOps, or platform engineering role.
  • Deep expertise with Kubernetes and Terraform in a cloud-first environment (AWS preferred).
  • Strong track record with production observability, including defining SLOs, building monitoring, and leading incident response.
  • Strong software engineering fundamentals in Python or Go, applied to infrastructure automation.
  • Experience driving cloud cost-efficiency and performance optimization.
  • Fluency with AI tools applied to engineering and operations workflows.
  • Experience supporting AI/ML or data-intensive workloads in production is a plus.
  • Experience operating in a security-conscious or regulated environment (HIPAA, SOC 2) is a plus.
  • Experience with Kubernetes APIs is a plus.

Responsibilities

  • Own the end-to-end reliability, performance, and resilience of Garner’s cloud environments (AWS, Kubernetes).
  • Define, measure, and uphold SLOs across critical services.
  • Serve in the on-call rotation, lead incident response, and drive deep-dive root cause analysis.
  • Build and maintain monitoring, alerting, and observability systems.
  • Translate scaling requirements into automated infrastructure-as-code deliverables using Terraform.
  • Automate repetitive operational work using AI tools to create hands-free, monitored processes.
  • Build and maintain deployment and observability standards to enable the engineering team.
  • Ensure infrastructure and operations meet security and HIPAA compliance obligations.
View Full Description & ApplyYou'll be redirected to the employer's site
$191,000 - $226,000
Apply Now