- Define and execute Everbridge's global production operations and reliability strategy.
- Lead and develop high-performing global SRE and DRE teams responsible for platform reliability and operational engineering.
- Own the operational excellence of Everbridge's AWS and Kubernetes-based cloud platform, ensuring scalability, resilience, security, and performance.
- Establish best practices for service reliability, including Service Level Objectives (SLOs), error budgets, production readiness, capacity planning, observability, and operational automation.
- Drive disciplined incident management, change governance, release management, and post-incident reviews to continuously improve platform stability and reduce operational risk.
- Lead disaster recovery planning, resilience testing, business continuity initiatives, and operational readiness across global production environments.
- Partner closely with Engineering, Product, Security, Customer Support, and Customer Success to embed reliability into the software development lifecycle and improve customer outcomes.
- Champion automation, cloud-native engineering practices, and continuous improvement to enhance operational efficiency and platform performance.
- Provide executive leadership and reporting on operational health, reliability metrics, customer-impacting incidents, and strategic initiatives.
AWSKubernetesCI/CD+1 more