- Own platform availability by monitoring, triaging, and resolving incidents within defined SLA windows.
- Manage cloud infrastructure on AWS and/or GCP, including provisioning, scaling, and daily operations.
- Maintain and improve CI/CD pipelines and GitOps workflows.
- Operate observability systems, including monitoring, logging, and alerting at production scale.
- Participate in an on-call rotation as part of a global follow-the-sun coverage model.
- Configure, deploy, and manage AI tooling and MCP servers in production environments.
- Contribute to infrastructure automation, scripting, and internal tooling development.
- Write clear post-incident reviews and contribute to monthly operational reports.
AWSPythonBash+5 more