Site Reliability Engineer III (DBA)
New
B
BackblazeCloud storage
Remote - USFull-TimeSenior
Salary125,000 - 150,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- 6–8 years of experience in site reliability engineering, systems engineering, infrastructure operations, database engineering, or similar roles, with meaningful experience supporting production database systems.
- Required Skills
- DockerPythonSQLKubernetesMySQLCassandraLinuxTerraformAnsible
Requirements
- Have 6–8 years of experience in site reliability engineering, systems engineering, infrastructure operations, database engineering, or similar roles, including meaningful production database experience.
- Bring deep hands-on experience with MySQL and distributed or sharded database systems.
- Have experience administering and supporting NoSQL databases such as Cassandra.
- Have experience designing high-availability database architecture, replication topology, backup strategies, and disaster recovery processes.
- Demonstrate strong SQL skills, including query performance analysis, indexing, schema design, and troubleshooting.
- Have solid Linux systems administration and troubleshooting skills.
- Have experience with security-focused operations, including patching, hardening, access controls, and vulnerability remediation.
- Understand monitoring, alerting, incident response, root cause analysis, SLIs, SLOs, and error budgets.
- Have experience with containers and orchestration platforms, including Kubernetes and Docker.
- Have experience with Terraform, Ansible, Jenkins, and HashiCorp products such as Vault and Nomad.
- Be proficient in at least one scripting language such as Python, Bash, or Go.
- Have experience establishing operational procedures, runbooks, documentation, and escalation processes, and mentoring or onboarding engineers.
- Have a bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent professional experience.
Responsibilities
- Design, deploy, and own highly available database architecture for Vitess and Cassandra.
- Optimize database performance through query tuning, indexing, schema design, and capacity planning.
- Own database backup, recovery, replication, and disaster recovery strategies, and validate recovery procedures.
- Establish operational procedures, runbooks, escalation guidance, and training materials for Level 1 and Level 2 SRE Database Engineers.
- Monitor production service health using SLIs, SLOs, error budgets, monitoring, logging, and alerting platforms.
- Participate in on-call rotations, respond to production incidents, and lead or contribute to root cause analysis and post-incident reviews.
- Develop automation and operational tooling using scripting languages and infrastructure/configuration management tools.
- Operate and troubleshoot production environments using Kubernetes, Docker, and Vitess tools.
- Lead Production Readiness Reviews and support the operational readiness of new database-backed services.
View Full Description & ApplyYou'll be redirected to the employer's site