- Own platform availability: monitor, triage, and resolve incidents within defined SLA windows
- Manage cloud infrastructure on AWS and/or GCP - provisioning, scaling, and day-to-day operations
- Maintain and improve CI/CD pipelines and GitOps workflows
- Operate observability systems: monitoring, logging, and alerting at production scale
- Participate in on-call rotation as part of the global follow-the-sun coverage model
- Configure, deploy, and manage AI tooling and MCP servers in production environments
- Contribute to infrastructure automation, scripting, and internal tooling
- Write clear post-incident reviews and contribute to the monthly operational report
- Collaborate closely with engineering teams across multiple time zones
AWSPythonBash+5 more