- Design, provision, and manage AWS infrastructure using Terraform.
- Operate, maintain, and scale production workloads running on Kubernetes.
- Package, deploy, and manage applications using Helm and automation tools.
- Define and maintain SLIs, SLOs, and error budgets.
- Develop automation for deployment, scaling, monitoring, and incident response.
- Own platform observability via metrics, logging, tracing, and alerting.
- Lead incident response and facilitate blameless postmortems.
- Implement security best practices for HIPAA and SOC 2 compliance.
- Participate in an on-call rotation for production systems.