Tech Lead Manager, GPU Cluster Infrastructure
F
FAR.AIAI Research
Remote (International). We can hire remotely in most countries., Working hours should overlap with Berkeley, CA (PST/PDT).Full-TimeManager
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years in systems or infrastructure engineering
- Required Skills
- PythonKubernetesGoPrometheusLinuxTerraformAnsible
Requirements
- 5+ years in systems or infrastructure engineering on production Linux, GPU, HPC, or large-scale batch platforms.
- Experience managing engineers as a tech lead or project lead.
- Deep expertise in running production Kubernetes for GPU workloads with batch layers (Slurm, Kueue, Volcano, or similar).
- Proficiency in infrastructure as code (Terraform, Ansible) and deployment tools (Helm, ArgoCD).
- Strong programming skills in Python, Go, Rust, or C++.
- Experience with observability and monitoring tools like Prometheus.
- Ability to write technical documentation including design docs, roadmaps, and incident summaries.
- Knowledge of GPU systems, networking, and distributed storage (e.g., VAST, Weka, Lustre, Ceph).
- Understanding of cluster security, scheduling internals, and multi-provider platform management.
Responsibilities
- Set the platform's technical direction, define roadmaps, and decide on scheduling/storage architectures.
- Manage, hire, and coach a team of senior infrastructure engineers.
- Operate and maintain a fleet of GPU clusters, ensuring performance and fault tolerance for large-scale AI research.
- Design security protocols for shared clusters, covering identity, access, and workload isolation.
- Define operational practices including on-call, incident response, and observability standards.
- Collaborate with research teams to provide infrastructure support and resolve performance issues.
- Serve as the escalation point for research teams and infrastructure providers.
View Full Description & ApplyYou'll be redirected to the employer's site