Staff Site Reliability Engineer
New
G
Garner HealthHealthcare Technology
Garner is headquartered in NYC, but this position is available for individuals who are comfortable with remote work and occasional travel to HQ.Full-TimeStaff
Salary$241,000 - $270,000
Apply NowOpens the employer's application page
Job Details
- Experience
- 7+ years
- Required Skills
- AWSPythonKubernetesTypeScriptGoPostgresTerraformDatadog
Requirements
- 7+ years of hands-on experience operating production cloud infrastructure at scale in an SRE, DevOps, or platform engineering role.
- Deep expertise with Kubernetes and Terraform in a cloud-first environment (AWS preferred).
- Proven track record of architecting reliability for systems at scale.
- Experience designing organizational reliability practices including SLO frameworks, observability, and incident response programs.
- Strong Python or Go skills applied to infrastructure automation.
- Experience driving cloud cost-efficiency and performance optimization.
- Mentorship experience with the ability to set technical direction.
- Excellent communication skills for both technical and non-technical stakeholders.
- Fluency with AI tools applied to engineering and operations workflows.
- Experience supporting AI/ML or data-intensive workloads is a plus.
- Experience operating in a security-conscious or regulated environment (HIPAA, SOC 2) is a plus.
Responsibilities
- Architect and own the end-to-end reliability, performance, and resilience of cloud environments (AWS, Kubernetes) powering products and AI/ML workloads.
- Design and implement the SLO framework for critical services and lead technical decision-making at scale.
- Set the standard for incident response, serve in the on-call rotation, and drive root cause analysis for complex escalations.
- Architect monitoring, alerting, and observability platforms to ensure proactive issue detection.
- Transform high-level scaling requirements into automated, composable infrastructure-as-code deliverables using Terraform.
- Optimize cloud cost-efficiency and performance across compute, storage, and networking.
- Mentor engineers across the organization to raise the bar for operational rigor and discipline.
- Ensure infrastructure and operations meet security and HIPAA compliance requirements.
View Full Description & ApplyYou'll be redirected to the employer's site