- Architect system monitoring, alerting, and logging with Prometheus, Grafana, and ELK Stack, and establish actionable SLIs/SLOs.
- Design and implement automated infrastructure and workflows using Terraform and Ansible across AWS, GCP, and Azure.
- Optimize CI/CD pipeline architectures for fast, secure, zero-downtime releases.
- Guide high-priority incident triage, lead blameless post-mortems, and implement preventative measures.
- Partner with product and software engineering teams to incorporate SRE practices into product roadmaps.
- Identify performance bottlenecks and evaluate technologies such as eBPF and container orchestration.
- Own reliability engineering workstreams across multi-cloud environments.