- Build, lead, and develop the UK SRE team, establishing operational standards and reliability goals.
- Ensure high availability, stability, security, and performance of business-critical SaaS platforms.
- Define and drive the operational strategy for the global platform.
- Lead major incident management as the senior escalation point for critical production events.
- Establish and monitor reliability metrics including SLIs, SLOs, and operational KPIs.
- Drive automation across infrastructure, deployments, and monitoring workflows.
- Champion the adoption of AI-powered operations to enhance engineering productivity.
- Partner with Engineering, Product, Security, and Infrastructure teams to improve platform scalability.
- Lead disaster recovery planning and operational resilience initiatives.
AWSDockerArtificial Intelligence+4 more