Senior Site Reliability Engineer
New
G
Garner HealthHealthtech
RemoteFull-TimeSenior
Salary$191,000 - $226,000
Apply NowOpens the employer's application page
Job Details
- Experience
- 4+ years
- Required Skills
- AWSPythonKubernetesGoTerraformHIPAA
Requirements
- 4+ years of hands-on experience operating production cloud infrastructure at scale in an SRE, DevOps, or platform engineering role.
- Deep expertise with Kubernetes and Terraform in a cloud-first environment (AWS preferred).
- Strong track record with production observability, including defining SLOs, building monitoring, and leading incident response.
- Strong software engineering fundamentals in Python or Go, applied to infrastructure automation.
- Experience driving cloud cost-efficiency and performance optimization.
- Fluency with AI tools applied to engineering and operations workflows.
- Experience supporting AI/ML or data-intensive workloads in production is a plus.
- Experience operating in a security-conscious or regulated environment (HIPAA, SOC 2) is a plus.
- Experience with Kubernetes APIs is a plus.
Responsibilities
- Own the end-to-end reliability, performance, and resilience of Garner’s cloud environments (AWS, Kubernetes).
- Define, measure, and uphold SLOs across critical services.
- Serve in the on-call rotation, lead incident response, and drive deep-dive root cause analysis.
- Build and maintain monitoring, alerting, and observability systems.
- Translate scaling requirements into automated infrastructure-as-code deliverables using Terraform.
- Automate repetitive operational work using AI tools to create hands-free, monitored processes.
- Build and maintain deployment and observability standards to enable the engineering team.
- Ensure infrastructure and operations meet security and HIPAA compliance obligations.
View Full Description & ApplyYou'll be redirected to the employer's site