- Act as incident commander on major and critical incidents, owning coordination, decision-making cadence, and escalation.
- Drive down time to detect, time to engage, and time to recover (MTTR).
- Manage on-call rotations, escalation policies, and paging hygiene to ensure alert quality.
- Own internal stakeholder alignment and merchant-facing communications during incidents.
- Conduct blameless postmortems and ensure action items are tracked to closure.
- Partner with engineering teams to translate incident patterns into reliability roadmap items.
- Define and maintain incident severity levels, response runbooks, and the incident operating model.
- Report on incident trends, reliability posture, and SLA/SLO performance to engineering leadership.