- Configure and maintain cloud infrastructure automation using Terraform, focusing on CDN optimization and content delivery performance
- Develop capacity planning strategies and performance optimization initiatives for high-volume spatial content delivery
- Instrument services to understand system health and create/optimize monitoring dashboards and alerting systems
- Design and implement comprehensive observability strategies including SLI/SLO definition and error budget management
- Help define escalation policies and participate in on-call rotations to ensure the 24/7 health of Miris pipelines
- Lead incident response efforts and conduct thorough post-mortems to drive systemic improvements
- Establish reliability engineering practices including code review processes and deployment safety measures
- Mentor DevOps engineers on operational best practices and production readiness standards
KubernetesGrafanaPrometheus+2 more