Technical Lead - GPU Infrastructure

New
J
JobgetherGPU Infrastructure
Based in United States, UTC to UTC+5:30Full-TimeLead
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Languages
English
Experience
8+ years of hands-on engineering experience, including at least 3 years leading teams
Required Skills
Node.jsJavascriptKubernetesGrafanaPrometheusLinux

Requirements

  • 8+ years of hands-on engineering experience, including 3+ years in a technical leadership role.
  • Bachelor's or Master's degree in CS, engineering, or equivalent experience.
  • Extensive production experience operating Slurm.
  • Proven experience operating HPC or GPU training clusters.
  • Deep expertise operating NVIDIA GPU fleets on bare metal (CUDA, DCGM, NVSwitch).
  • Strong knowledge of InfiniBand, RDMA, SR-IOV, and NCCL performance tuning.
  • Linux systems expertise (kernel modules, PCIe passthrough, cgroups).
  • Production Kubernetes experience (control planes, operators, multi-tenancy).
  • Experience with HPC storage systems like VAST, Lustre, or NFS.
  • Working proficiency in JavaScript and Node.js for code review and architecture.
  • Excellent written and spoken English.
  • Ability to work in a time zone between UTC and UTC+5:30.

Responsibilities

  • Own end-to-end platform architecture, designs, and documentation.
  • Lead and line-manage a distributed engineering team.
  • Design, build, and operate a managed Slurm service for research and training workloads.
  • Lead GPU infrastructure operations on bare-metal, including NVIDIA drivers and CUDA lifecycles.
  • Manage Kubernetes cluster bootstrap, lifecycle, and GPU isolation.
  • Define managed inference architecture, including scaling and request routing.
  • Establish observability using metrics, logging, and SLOs.
  • Lead incident response and on-call operations.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now