- Define and operate Service Level Objectives (SLOs) aligned with customer SLAs.
- Build and maintain the observability stack including metrics, logs, traces, and alerting.
- Lead incident response and chair post-incident reviews.
- Drive automation to reduce toil and improve mean-time-to-recover (MTTR).
- Author and maintain operational runbooks alongside the NOC.
- Manage on-call rotation, escalation paths, and incident-management tooling.
- Coordinate cross-functionally with NOC, Platform Engineering, and Network Engineering.
- Drive chaos engineering, game days, and reliability testing programs.
- Produce SLA performance reports in coordination with the SLA Manager.
- Mentor junior engineers and contribute to engineering culture.
PythonKubernetesGo+3 more