Site Reliability Engineer III
New
V
VidaVirtual healthcare
Listing location: United States; Workplace type: Remote; This is a fully remote role with no time zone restrictions., no time zone restrictionsFull-TimeSenior
Salary175,000 - 185,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years of experience in SRE, DevOps, or infrastructure engineering, with real ownership of production systems.
- Required Skills
- PythonKubernetesCI/CDTerraformGitHub ActionsDatadog
Requirements
- Bachelor's degree at a minimum.
- 5+ years of experience in SRE, DevOps, or infrastructure engineering, with ownership of production systems.
- Deep hands-on Terraform experience, including structuring modules and managing state across environments.
- Strong working knowledge of GCP, including GKE, Cloud SQL (MySQL and PostgreSQL), IAM, networking, load balancing, and cost management.
- Production Kubernetes experience, including autoscaling, resource management, and deciding what belongs in the cluster versus outside it.
- Hands-on experience building monitoring, alerting, and dashboards with tools such as Datadog or Cloud Monitoring.
- Proficiency in Python for tooling and automation.
- Ability to work across teams and disciplines and explain infrastructure decisions to non-specialists.
- Preferred: experience as an early or first SRE hire and refactoring or consolidating a large Terraform codebase.
- Preferred: experience improving observability from a less mature baseline or working in a HIPAA-regulated or other compliance-driven environment.
- Preferred: GitHub Actions CI/CD experience and production experience running Django applications or Airflow on Kubernetes.
- Preferred: experience designing or migrating to multi-cluster Kubernetes architectures.
Responsibilities
- Consolidate Terraform repositories into a clean, documented structure and establish conventions for state management, modules, code review, and CI checks.
- Normalize environments and improve build and deploy automation in GitHub Actions, including drift detection and alerting.
- Apply patches and upgrades across Cloud SQL databases and application runtimes.
- Right-size compute and database workloads, including connection pooling and scaling improvements for high-traffic services.
- Evaluate Kubernetes architecture, including whether and when to move to a multi-cluster setup.
- Improve monitoring and observability in Datadog and Cloud Monitoring to detect issues before they become incidents.
- Design observability access for contractors and external partners while keeping protected health information out of view.
- Retire legacy infrastructure and tooling that has been replaced but not decommissioned.
- Build operational processes, including runbooks, an on-call rotation, and escalation documentation.
- Support infrastructure readiness for enterprise launches starting January 1.
View Full Description & ApplyYou'll be redirected to the employer's site