Manager, Production Site Reliability Engineering

New
V
VultrCloud infrastructure
Remote - United StatesFull-TimeManager
Salary140,000 - 160,000 USD per year
Apply NowOpens the employer's application page

Job Details

Experience
10+ years of professional experience in site reliability engineering or infrastructure operations, with at least 2 years in a team lead or management role
Required Skills
RedisLinuxAnsibleNetworking

Requirements

  • Bring 10+ years of professional experience in site reliability engineering or infrastructure operations.
  • Have at least 2 years of experience in a team lead or management role.
  • Demonstrate experience operating production web stacks at scale and debugging performance across the request path from load balancer to database.
  • Bring strong Linux systems knowledge, including networking, systemd, package management, firewall configuration, and performance tuning.
  • Have experience building and operating monitoring and alerting systems and driving observability and proactive incident response.
  • Have experience with configuration management at scale, such as Puppet, Ansible, or Chef.
  • Have experience hiring and building engineering teams.
  • Understand database replication, backup strategies, failure modes, query plans, and database failover.
  • Have experience leading production migrations or major infrastructure transitions, including planning, parallel runs, validation, and safe cutover.
  • Communicate effectively through runbooks, postmortems, incident communications, and cross-functional coordination.
  • Bonus: Experience with Harvester or similar HCI platforms, Redis clusters, HAProxy, keepalived, PCI environments, PHP, or Cloudflare and DNS cutovers.

Responsibilities

  • Build, hire, mentor, and lead a team of SREs and database engineers responsible for the production control plane.
  • Own the availability, performance, and operability of the production web stack.
  • Lead incident response, drive postmortems, and turn lessons into system improvements, runbooks, and alerting.
  • Own configuration management for the production environment across control plane hosts.
  • Partner with engineering teams to ensure new services are observable, deployable, and documented before launch.
  • Lead control-plane re-architecture, including infrastructure setup, parallel runs, and cutover.
  • Set the production roadmap for capacity planning, scaling, disaster recovery, and architecture evolution.
  • Establish monitoring, alerting, observability, SLO tracking, and operational reviews.
  • Manage production database and caching operations, including replication, performance tuning, backups, and failover testing.
View Full Description & ApplyYou'll be redirected to the employer's site
140,000 - 160,000 USD per year
Apply Now