Manager, Production Site Reliability Engineering
New
V
VultrCloud infrastructure
Remote - United StatesFull-TimeManager
Salary140,000 - 160,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- 10+ years of professional experience in site reliability engineering or infrastructure operations, with at least 2 years in a team lead or management role
- Required Skills
- RedisLinuxAnsibleNetworking
Requirements
- Bring 10+ years of professional experience in site reliability engineering or infrastructure operations.
- Have at least 2 years of experience in a team lead or management role.
- Demonstrate experience operating production web stacks at scale and debugging performance across the request path from load balancer to database.
- Bring strong Linux systems knowledge, including networking, systemd, package management, firewall configuration, and performance tuning.
- Have experience building and operating monitoring and alerting systems and driving observability and proactive incident response.
- Have experience with configuration management at scale, such as Puppet, Ansible, or Chef.
- Have experience hiring and building engineering teams.
- Understand database replication, backup strategies, failure modes, query plans, and database failover.
- Have experience leading production migrations or major infrastructure transitions, including planning, parallel runs, validation, and safe cutover.
- Communicate effectively through runbooks, postmortems, incident communications, and cross-functional coordination.
- Bonus: Experience with Harvester or similar HCI platforms, Redis clusters, HAProxy, keepalived, PCI environments, PHP, or Cloudflare and DNS cutovers.
Responsibilities
- Build, hire, mentor, and lead a team of SREs and database engineers responsible for the production control plane.
- Own the availability, performance, and operability of the production web stack.
- Lead incident response, drive postmortems, and turn lessons into system improvements, runbooks, and alerting.
- Own configuration management for the production environment across control plane hosts.
- Partner with engineering teams to ensure new services are observable, deployable, and documented before launch.
- Lead control-plane re-architecture, including infrastructure setup, parallel runs, and cutover.
- Set the production roadmap for capacity planning, scaling, disaster recovery, and architecture evolution.
- Establish monitoring, alerting, observability, SLO tracking, and operational reviews.
- Manage production database and caching operations, including replication, performance tuning, backups, and failover testing.
View Full Description & ApplyYou'll be redirected to the employer's site