- Own reliability and observability across the organization including SLAs/SLOs, instrumentation, and on-call health.
- Design, deploy, and maintain Kubernetes infrastructure and core AWS services.
- Build and maintain CI/CD pipelines using GitHub Actions, Argo, and Helm.
- Drive adoption of Infrastructure as Code standards across multiple engineering teams.
- Partner with engineering teams to roll out new standards and processes.
- Participate in on-call rotation and lead incident response and root-cause analysis.
- Use and promote AI tools to improve development, debugging, and operational speed.
AWSPostgreSQLKubernetes+4 more