Senior Site Reliability Engineer (Systems Engineer III)
New
O
OpheliaHealthcare telehealth
Location: United StatesFull-TimeSenior
Salary140,000 - 150,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years experience in engineering, including 3+ in a cloud infrastructure, DevOps, or SRE-focused role
- Required Skills
- PythonFirebaseTerraform
Requirements
- Have 5+ years of engineering experience, including 3+ years in a cloud infrastructure, DevOps, or SRE-focused role.
- Bring expert, hands-on GCP fluency, including Cloud Run, Cloud Operations, IAM, and Firebase/Firestore.
- Have Terraform fluency; it is strongly preferred.
- Be expert in at least one language used for infrastructure automation, such as Python, and strong in shell scripting.
- Be able to read and debug TypeScript/Node; SQL for BigQuery and Log Analytics is a plus.
- Have production operations experience, including on-call participation, incident response, postmortems, and defining SLOs.
- Apply quantitative fluency to production systems, including percentiles, tail latency, error-budget math, burn-rate alerting, time series, and histograms.
- Own complex, ambiguous projects end-to-end and communicate proactively with technical and non-technical stakeholders.
- Use written documentation such as runbooks, postmortems, and design docs, and pair to share knowledge.
- Use AI tools hands-on in daily engineering work and have experience building automation with them.
Responsibilities
- Own reliability, availability, and performance of the GCP platform, and define and operationalize SLOs and SLIs.
- Drive improvements to endpoint latency, Cloud Run cold starts, time to recovery, and disaster recovery posture.
- Extend GCP Cloud Operations logging, metrics, alerting policies, and dashboards.
- Move alerting from Slack into PagerDuty and mature incident categorization, escalation paths, and on-call practices.
- Expand Terraform coverage across GCP configuration and improve the GitHub Actions CI/CD pipeline.
- Partner with the head of IT on IAM hardening, secrets management, audit logging, and HIPAA/SOC 2 infrastructure controls.
- Develop engineers' GCP and DevOps skills through documentation, runbooks, presentations, pairing, and code review.
- Bring reliability recommendations to new features and share reusable AI skills and automation across the team.
View Full Description & ApplyYou'll be redirected to the employer's site