- Enhance high availability and elasticity to handle growth automatically.
- Boost observability capabilities, including telemetry, dashboards, and alerting.
- Improve disaster recovery tools and incident discovery processes.
- Handle Kubernetes lifecycle tasks, cluster management, and autoscaling.
- Extract peak performance from ClickHouse and identify system bottlenecks.
- Transform manual processes into repeatable, well-managed systems.
- Strengthen CI/CD foundations to support confident deployments.
- Participate in on-call rotations to resolve incidents and support service health.
AWSPythonSQL+6 more