- Lead and execute on complex technical troubleshooting and incident resolution from investigation to delivery of permanent solutions.
- Design and implement comprehensive monitoring processes, including creating detailed playbooks and runbooks for common scenarios.
- Draft detailed documentation including post-incident reviews and knowledge base articles.
- Leverage AI-powered tools and workflows to automate issue detection, diagnosis, and resolution processes.
- Collaborate with cross-functional teams to deliver solutions, optimize support processes, and implement scalable monitoring systems.
- Take ownership of complex production issues and provide technical expertise across distributed systems, infrastructure, and application layers.
- Develop and maintain automation scripts and tools to reduce manual intervention and improve system reliability.
- Bridge communication between engineering teams, customer experience, and clients during critical incidents and implementation challenges.