Site Reliability Engineer III (DBA)

New
B
BackblazeCloud storage
Remote - USFull-TimeSenior
Salary125,000 - 150,000 USD per year
Apply NowOpens the employer's application page

Job Details

Experience
6–8 years of experience in site reliability engineering, systems engineering, infrastructure operations, database engineering, or similar roles, with meaningful experience supporting production database systems.
Required Skills
DockerPythonSQLKubernetesMySQLCassandraLinuxTerraformAnsible

Requirements

  • Have 6–8 years of experience in site reliability engineering, systems engineering, infrastructure operations, database engineering, or similar roles, including meaningful production database experience.
  • Bring deep hands-on experience with MySQL and distributed or sharded database systems.
  • Have experience administering and supporting NoSQL databases such as Cassandra.
  • Have experience designing high-availability database architecture, replication topology, backup strategies, and disaster recovery processes.
  • Demonstrate strong SQL skills, including query performance analysis, indexing, schema design, and troubleshooting.
  • Have solid Linux systems administration and troubleshooting skills.
  • Have experience with security-focused operations, including patching, hardening, access controls, and vulnerability remediation.
  • Understand monitoring, alerting, incident response, root cause analysis, SLIs, SLOs, and error budgets.
  • Have experience with containers and orchestration platforms, including Kubernetes and Docker.
  • Have experience with Terraform, Ansible, Jenkins, and HashiCorp products such as Vault and Nomad.
  • Be proficient in at least one scripting language such as Python, Bash, or Go.
  • Have experience establishing operational procedures, runbooks, documentation, and escalation processes, and mentoring or onboarding engineers.
  • Have a bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent professional experience.

Responsibilities

  • Design, deploy, and own highly available database architecture for Vitess and Cassandra.
  • Optimize database performance through query tuning, indexing, schema design, and capacity planning.
  • Own database backup, recovery, replication, and disaster recovery strategies, and validate recovery procedures.
  • Establish operational procedures, runbooks, escalation guidance, and training materials for Level 1 and Level 2 SRE Database Engineers.
  • Monitor production service health using SLIs, SLOs, error budgets, monitoring, logging, and alerting platforms.
  • Participate in on-call rotations, respond to production incidents, and lead or contribute to root cause analysis and post-incident reviews.
  • Develop automation and operational tooling using scripting languages and infrastructure/configuration management tools.
  • Operate and troubleshoot production environments using Kubernetes, Docker, and Vitess tools.
  • Lead Production Readiness Reviews and support the operational readiness of new database-backed services.
View Full Description & ApplyYou'll be redirected to the employer's site
125,000 - 150,000 USD per year
Apply Now