Senior Software Engineer, GPU Cluster Infrastructure
New
F
FAR.AIAI Research
Remote (International). We can hire remotely in most countries., Berkeley-adjacent timezone overlap preferred.Full-TimeSenior
Salary150,000 - 275,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- 3+ years
- Required Skills
- PythonKubernetesC++GoPrometheusRustTerraformAnsibleHelm
Requirements
- 3+ years of experience in systems or infrastructure engineering.
- Experience operating production Linux and large-scale batch platforms.
- Proven experience running production Kubernetes for GPU workloads.
- Hands-on experience with batch layers such as Slurm, Kueue, or Volcano.
- Proficiency in infrastructure as code (Terraform or Ansible).
- Experience with deployment tools such as Helm or ArgoCD.
- Familiarity with monitoring tools such as Prometheus.
- Strong programming skills in Python, Go, Rust, or C++.
- Ability to clearly document designs, incidents, and escalations.
- Experience with distributed training (PyTorch, NCCL) or cluster security is a plus.
Responsibilities
- Operate the Kubernetes GPU fleet, including node lifecycle management, driver/image rollouts, and capacity planning.
- Manage batch scheduling and multi-tenancy with queues, priorities, and preemption.
- Design and maintain storage solutions, including high-performance shared filesystems and object storage.
- Ensure fault-tolerant multi-node training by owning node health, NCCL debugging, and checkpoint patterns.
- Harden the platform through identity, network policy, secrets management, and container sandboxing.
- Provision new capacity and integrate clusters using infrastructure as code.
- Collaborate with research teams to solve infrastructure bottlenecks and contribute to shared on-call/runbook efforts.
View Full Description & ApplyYou'll be redirected to the employer's site