- Design systems for cluster management, deployment automation, and production monitoring.
- Build the operational backbone that keeps vLLM running reliably at massive scale.
- Ensure vLLM deployments are observable, debuggable, and recoverable.
- Enable teams to serve AI models without friction.
- Manage GPU clusters and debug hardware-related issues.
- Work across AWS, GCP, Azure, and on-premise infrastructure.
AWSPythonGCP+5 more