Staff Site Reliability Engineer

New
G
Garner HealthHealthcare Technology
Garner is headquartered in NYC, but this position is available for individuals who are comfortable with remote work and occasional travel to HQ.Full-TimeStaff
Salary$241,000 - $270,000
Apply NowOpens the employer's application page

Job Details

Experience
7+ years
Required Skills
AWSPythonKubernetesTypeScriptGoPostgresTerraformDatadog

Requirements

  • 7+ years of hands-on experience operating production cloud infrastructure at scale in an SRE, DevOps, or platform engineering role.
  • Deep expertise with Kubernetes and Terraform in a cloud-first environment (AWS preferred).
  • Proven track record of architecting reliability for systems at scale.
  • Experience designing organizational reliability practices including SLO frameworks, observability, and incident response programs.
  • Strong Python or Go skills applied to infrastructure automation.
  • Experience driving cloud cost-efficiency and performance optimization.
  • Mentorship experience with the ability to set technical direction.
  • Excellent communication skills for both technical and non-technical stakeholders.
  • Fluency with AI tools applied to engineering and operations workflows.
  • Experience supporting AI/ML or data-intensive workloads is a plus.
  • Experience operating in a security-conscious or regulated environment (HIPAA, SOC 2) is a plus.

Responsibilities

  • Architect and own the end-to-end reliability, performance, and resilience of cloud environments (AWS, Kubernetes) powering products and AI/ML workloads.
  • Design and implement the SLO framework for critical services and lead technical decision-making at scale.
  • Set the standard for incident response, serve in the on-call rotation, and drive root cause analysis for complex escalations.
  • Architect monitoring, alerting, and observability platforms to ensure proactive issue detection.
  • Transform high-level scaling requirements into automated, composable infrastructure-as-code deliverables using Terraform.
  • Optimize cloud cost-efficiency and performance across compute, storage, and networking.
  • Mentor engineers across the organization to raise the bar for operational rigor and discipline.
  • Ensure infrastructure and operations meet security and HIPAA compliance requirements.
View Full Description & ApplyYou'll be redirected to the employer's site
$241,000 - $270,000
Apply Now