Senior Software Engineer, GPU Cluster Infrastructure

New
F
FAR.AIAI Research
Remote (International). We can hire remotely in most countries., Berkeley-adjacent timezone overlap preferred.Full-TimeSenior
Salary150,000 - 275,000 USD per year
Apply NowOpens the employer's application page

Job Details

Experience
3+ years
Required Skills
PythonKubernetesC++GoPrometheusRustTerraformAnsibleHelm

Requirements

  • 3+ years of experience in systems or infrastructure engineering.
  • Experience operating production Linux and large-scale batch platforms.
  • Proven experience running production Kubernetes for GPU workloads.
  • Hands-on experience with batch layers such as Slurm, Kueue, or Volcano.
  • Proficiency in infrastructure as code (Terraform or Ansible).
  • Experience with deployment tools such as Helm or ArgoCD.
  • Familiarity with monitoring tools such as Prometheus.
  • Strong programming skills in Python, Go, Rust, or C++.
  • Ability to clearly document designs, incidents, and escalations.
  • Experience with distributed training (PyTorch, NCCL) or cluster security is a plus.

Responsibilities

  • Operate the Kubernetes GPU fleet, including node lifecycle management, driver/image rollouts, and capacity planning.
  • Manage batch scheduling and multi-tenancy with queues, priorities, and preemption.
  • Design and maintain storage solutions, including high-performance shared filesystems and object storage.
  • Ensure fault-tolerant multi-node training by owning node health, NCCL debugging, and checkpoint patterns.
  • Harden the platform through identity, network policy, secrets management, and container sandboxing.
  • Provision new capacity and integrate clusters using infrastructure as code.
  • Collaborate with research teams to solve infrastructure bottlenecks and contribute to shared on-call/runbook efforts.
View Full Description & ApplyYou'll be redirected to the employer's site
150,000 - 275,000 USD per year
Apply Now