Site Reliability Engineer III

New
V
VidaVirtual healthcare
Listing location: United States; Workplace type: Remote; This is a fully remote role with no time zone restrictions., no time zone restrictionsFull-TimeSenior
Salary175,000 - 185,000 USD per year
Apply NowOpens the employer's application page

Job Details

Experience
5+ years of experience in SRE, DevOps, or infrastructure engineering, with real ownership of production systems.
Required Skills
PythonKubernetesCI/CDTerraformGitHub ActionsDatadog

Requirements

  • Bachelor's degree at a minimum.
  • 5+ years of experience in SRE, DevOps, or infrastructure engineering, with ownership of production systems.
  • Deep hands-on Terraform experience, including structuring modules and managing state across environments.
  • Strong working knowledge of GCP, including GKE, Cloud SQL (MySQL and PostgreSQL), IAM, networking, load balancing, and cost management.
  • Production Kubernetes experience, including autoscaling, resource management, and deciding what belongs in the cluster versus outside it.
  • Hands-on experience building monitoring, alerting, and dashboards with tools such as Datadog or Cloud Monitoring.
  • Proficiency in Python for tooling and automation.
  • Ability to work across teams and disciplines and explain infrastructure decisions to non-specialists.
  • Preferred: experience as an early or first SRE hire and refactoring or consolidating a large Terraform codebase.
  • Preferred: experience improving observability from a less mature baseline or working in a HIPAA-regulated or other compliance-driven environment.
  • Preferred: GitHub Actions CI/CD experience and production experience running Django applications or Airflow on Kubernetes.
  • Preferred: experience designing or migrating to multi-cluster Kubernetes architectures.

Responsibilities

  • Consolidate Terraform repositories into a clean, documented structure and establish conventions for state management, modules, code review, and CI checks.
  • Normalize environments and improve build and deploy automation in GitHub Actions, including drift detection and alerting.
  • Apply patches and upgrades across Cloud SQL databases and application runtimes.
  • Right-size compute and database workloads, including connection pooling and scaling improvements for high-traffic services.
  • Evaluate Kubernetes architecture, including whether and when to move to a multi-cluster setup.
  • Improve monitoring and observability in Datadog and Cloud Monitoring to detect issues before they become incidents.
  • Design observability access for contractors and external partners while keeping protected health information out of view.
  • Retire legacy infrastructure and tooling that has been replaced but not decommissioned.
  • Build operational processes, including runbooks, an on-call rotation, and escalation documentation.
  • Support infrastructure readiness for enterprise launches starting January 1.
View Full Description & ApplyYou'll be redirected to the employer's site
175,000 - 185,000 USD per year
Apply Now