Site Reliability Engineer III

New
V
VidaVirtual healthcare
All Vida Employees must reside in/be able to work from the U.S.- international work is prohibited., no time zone restrictionsFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
5+ years of experience in SRE, DevOps, or infrastructure engineering, with real ownership of production systems.
Required Skills
PostgreSQLPythonKubernetesMySQLTerraformDatadog

Requirements

  • Bachelor's degree at a minimum.
  • 5+ years of experience in SRE, DevOps, or infrastructure engineering, with ownership of production systems.
  • Deep hands-on Terraform experience, including structuring modules and managing state across environments.
  • Strong working knowledge of GCP, including GKE, Cloud SQL (MySQL and PostgreSQL), IAM, networking and load balancing, and cost management.
  • Production Kubernetes experience, including autoscaling, resource management, and deciding what belongs in the cluster versus outside it.
  • Hands-on experience building monitoring, alerting, and dashboards with Datadog or Cloud Monitoring.
  • Proficiency in Python for tooling and automation.
  • Ability to work across teams and disciplines and explain infrastructure decisions to non-specialists.
  • Preferred: experience as an early or first SRE hire and consolidating a large Terraform codebase.
  • Preferred: experience improving observability, working in a HIPAA-regulated or compliance-driven environment, or using GitHub Actions for CI/CD.
  • Preferred: experience running Django applications or Airflow in production on Kubernetes, or designing or migrating multi-cluster Kubernetes architectures.

Responsibilities

  • Consolidate Terraform repositories into a clean, documented structure and establish conventions for state management, modules, code review, and CI checks.
  • Normalize environments and improve build and deploy automation in GitHub Actions.
  • Add infrastructure drift detection and alerting.
  • Apply patches and upgrades across Cloud SQL databases and application runtimes.
  • Right-size compute and database workloads, including connection pooling and scaling improvements.
  • Evaluate Kubernetes architecture, including whether and when to move to a multi-cluster setup.
  • Improve monitoring and observability in Datadog and Cloud Monitoring.
  • Design observability access for contractors and external partners while keeping protected health information out of view.
  • Retire legacy infrastructure and tooling that has been replaced.
  • Build operational processes, including runbooks, an on-call rotation, and escalation documentation.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now