Vice President, Global Production Operations & Reliability
New
E
EverbridgeCritical Event Management
United StatesFull-TimeVp
Salary195,300 - 270,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- 15+ years
- Required Skills
- AWSKubernetesCI/CDSaaS
Requirements
- 15+ years of experience in production operations, cloud infrastructure, platform engineering, SRE, or related technology leadership roles.
- Proven experience leading global production operations for large-scale, mission-critical SaaS or cloud platforms.
- Deep expertise in AWS, Kubernetes, cloud-native architectures, distributed systems, and high-availability environments.
- Strong understanding of SRE principles, incident management, observability, disaster recovery, change management, CI/CD, and Infrastructure as Code.
- Demonstrated success building and leading high-performing global engineering and operations teams.
- Excellent executive communication, stakeholder management, and cross-functional leadership skills.
Responsibilities
- Define and execute Everbridge's global production operations and reliability strategy.
- Lead and develop high-performing global SRE and DRE teams responsible for platform reliability and operational engineering.
- Own the operational excellence of Everbridge's AWS and Kubernetes-based cloud platform, ensuring scalability, resilience, security, and performance.
- Establish best practices for service reliability, including Service Level Objectives (SLOs), error budgets, production readiness, capacity planning, observability, and operational automation.
- Drive disciplined incident management, change governance, release management, and post-incident reviews to continuously improve platform stability and reduce operational risk.
- Lead disaster recovery planning, resilience testing, business continuity initiatives, and operational readiness across global production environments.
- Partner closely with Engineering, Product, Security, Customer Support, and Customer Success to embed reliability into the software development lifecycle and improve customer outcomes.
- Champion automation, cloud-native engineering practices, and continuous improvement to enhance operational efficiency and platform performance.
- Provide executive leadership and reporting on operational health, reliability metrics, customer-impacting incidents, and strategic initiatives.
View Full Description & ApplyYou'll be redirected to the employer's site