Senior Site Reliability Engineer (Systems Engineer III)

New
O
OpheliaHealthcare telehealth
Location: United StatesFull-TimeSenior
Salary140,000 - 150,000 USD per year
Apply NowOpens the employer's application page

Job Details

Experience
5+ years experience in engineering, including 3+ in a cloud infrastructure, DevOps, or SRE-focused role
Required Skills
PythonFirebaseTerraform

Requirements

  • Have 5+ years of engineering experience, including 3+ years in a cloud infrastructure, DevOps, or SRE-focused role.
  • Bring expert, hands-on GCP fluency, including Cloud Run, Cloud Operations, IAM, and Firebase/Firestore.
  • Have Terraform fluency; it is strongly preferred.
  • Be expert in at least one language used for infrastructure automation, such as Python, and strong in shell scripting.
  • Be able to read and debug TypeScript/Node; SQL for BigQuery and Log Analytics is a plus.
  • Have production operations experience, including on-call participation, incident response, postmortems, and defining SLOs.
  • Apply quantitative fluency to production systems, including percentiles, tail latency, error-budget math, burn-rate alerting, time series, and histograms.
  • Own complex, ambiguous projects end-to-end and communicate proactively with technical and non-technical stakeholders.
  • Use written documentation such as runbooks, postmortems, and design docs, and pair to share knowledge.
  • Use AI tools hands-on in daily engineering work and have experience building automation with them.

Responsibilities

  • Own reliability, availability, and performance of the GCP platform, and define and operationalize SLOs and SLIs.
  • Drive improvements to endpoint latency, Cloud Run cold starts, time to recovery, and disaster recovery posture.
  • Extend GCP Cloud Operations logging, metrics, alerting policies, and dashboards.
  • Move alerting from Slack into PagerDuty and mature incident categorization, escalation paths, and on-call practices.
  • Expand Terraform coverage across GCP configuration and improve the GitHub Actions CI/CD pipeline.
  • Partner with the head of IT on IAM hardening, secrets management, audit logging, and HIPAA/SOC 2 infrastructure controls.
  • Develop engineers' GCP and DevOps skills through documentation, runbooks, presentations, pairing, and code review.
  • Bring reliability recommendations to new features and share reusable AI skills and automation across the team.
View Full Description & ApplyYou'll be redirected to the employer's site
140,000 - 150,000 USD per year
Apply Now