Site Reliability Engineer III
New
V
VidaVirtual healthcare
All Vida Employees must reside in/be able to work from the U.S.- international work is prohibited., no time zone restrictionsFull-TimeMiddle
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years of experience in SRE, DevOps, or infrastructure engineering, with real ownership of production systems.
- Required Skills
- PostgreSQLPythonKubernetesMySQLTerraformDatadog
Requirements
- Bachelor's degree at a minimum.
- 5+ years of experience in SRE, DevOps, or infrastructure engineering, with real ownership of production systems.
- Deep hands-on Terraform experience, including structuring modules and managing state across environments.
- Strong working knowledge of GCP, including GKE, Cloud SQL (MySQL and PostgreSQL), IAM, networking and load balancing, and cost management.
- Production Kubernetes experience, including autoscaling, resource management, and judgment about what belongs in the cluster versus outside it.
- Hands-on experience building monitoring, alerting, and dashboards with tools like Datadog or Cloud Monitoring.
- Proficiency in Python for tooling and automation.
- Experience working across multiple teams and disciplines and explaining infrastructure decisions to non-specialists.
- Preferred: experience as an early or first SRE hire, or refactoring or consolidating a large Terraform codebase.
- Preferred: experience improving observability from a less mature baseline or in a HIPAA-regulated or other compliance-driven environment.
- Preferred: GitHub Actions CI/CD experience, or experience running Django applications or Airflow in production on Kubernetes.
- Preferred: experience designing or migrating to multi-cluster Kubernetes architectures.
Responsibilities
- Consolidate Terraform repositories into a well-documented structure and establish conventions for state management, modules, code review, and CI checks.
- Normalize environments and improve build and deploy automation in GitHub Actions.
- Add infrastructure drift detection and alerting.
- Apply patches and upgrades across Cloud SQL databases and application runtimes.
- Right-size compute and database workloads, including connection pooling and scaling improvements.
- Evaluate Kubernetes architecture, including whether and when to move to a multi-cluster setup.
- Improve monitoring and observability in Datadog and Cloud Monitoring.
- Design observability access for contractors and external partners while keeping protected health information out of view.
- Retire legacy infrastructure and tooling that has been replaced.
- Build operational processes, including runbooks, an on-call rotation, and escalation documentation, and support infrastructure readiness for enterprise launches.
View Full Description & ApplyYou'll be redirected to the employer's site