- Set the platform's technical direction, define roadmaps, and decide on scheduling/storage architectures.
- Manage, hire, and coach a team of senior infrastructure engineers.
- Operate and maintain a fleet of GPU clusters, ensuring performance and fault tolerance for large-scale AI research.
- Design security protocols for shared clusters, covering identity, access, and workload isolation.
- Define operational practices including on-call, incident response, and observability standards.
- Collaborate with research teams to provide infrastructure support and resolve performance issues.
- Serve as the escalation point for research teams and infrastructure providers.
PythonKubernetesGo+4 more