Vice President, Global Production Operations & Reliability

New
E
EverbridgeCritical Event Management
United StatesFull-TimeVp
Salary195,300 - 270,000 USD per year
Apply NowOpens the employer's application page

Job Details

Experience
15+ years
Required Skills
AWSKubernetesCI/CDSaaS

Requirements

  • 15+ years of experience in production operations, cloud infrastructure, platform engineering, SRE, or related technology leadership roles.
  • Proven experience leading global production operations for large-scale, mission-critical SaaS or cloud platforms.
  • Deep expertise in AWS, Kubernetes, cloud-native architectures, distributed systems, and high-availability environments.
  • Strong understanding of SRE principles, incident management, observability, disaster recovery, change management, CI/CD, and Infrastructure as Code.
  • Demonstrated success building and leading high-performing global engineering and operations teams.
  • Excellent executive communication, stakeholder management, and cross-functional leadership skills.

Responsibilities

  • Define and execute Everbridge's global production operations and reliability strategy.
  • Lead and develop high-performing global SRE and DRE teams responsible for platform reliability and operational engineering.
  • Own the operational excellence of Everbridge's AWS and Kubernetes-based cloud platform, ensuring scalability, resilience, security, and performance.
  • Establish best practices for service reliability, including Service Level Objectives (SLOs), error budgets, production readiness, capacity planning, observability, and operational automation.
  • Drive disciplined incident management, change governance, release management, and post-incident reviews to continuously improve platform stability and reduce operational risk.
  • Lead disaster recovery planning, resilience testing, business continuity initiatives, and operational readiness across global production environments.
  • Partner closely with Engineering, Product, Security, Customer Support, and Customer Success to embed reliability into the software development lifecycle and improve customer outcomes.
  • Champion automation, cloud-native engineering practices, and continuous improvement to enhance operational efficiency and platform performance.
  • Provide executive leadership and reporting on operational health, reliability metrics, customer-impacting incidents, and strategic initiatives.
View Full Description & ApplyYou'll be redirected to the employer's site
195,300 - 270,000 USD per year
Apply Now