- Take end-to-end ownership of production GPU clusters for training and inference across AWS and OCI, keeping them highly available and performant.
- Join critical re-architecture sessions to redesign systems for higher efficiency and scale.
- Tune Linux performance deeply, at the OS and kernel level.
- Build automation in Python, Go, or Bash to manage, monitor, and self-heal infrastructure.
- Serve as the final escalation for the hardest GPU, networking (InfiniBand/RDMA), and system failures.
- Help achieve and maintain security certifications like SOC 2 and ISO.
AWSPythonBash+4 more