Technical Lead - GPU Infrastructure
New
J
JobgetherGPU Infrastructure
Based in United States, UTC to UTC+5:30Full-TimeLead
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Languages
- English
- Experience
- 8+ years of hands-on engineering experience, including at least 3 years leading teams
- Required Skills
- Node.jsJavascriptKubernetesGrafanaPrometheusLinux
Requirements
- 8+ years of hands-on engineering experience, including 3+ years in a technical leadership role.
- Bachelor's or Master's degree in CS, engineering, or equivalent experience.
- Extensive production experience operating Slurm.
- Proven experience operating HPC or GPU training clusters.
- Deep expertise operating NVIDIA GPU fleets on bare metal (CUDA, DCGM, NVSwitch).
- Strong knowledge of InfiniBand, RDMA, SR-IOV, and NCCL performance tuning.
- Linux systems expertise (kernel modules, PCIe passthrough, cgroups).
- Production Kubernetes experience (control planes, operators, multi-tenancy).
- Experience with HPC storage systems like VAST, Lustre, or NFS.
- Working proficiency in JavaScript and Node.js for code review and architecture.
- Excellent written and spoken English.
- Ability to work in a time zone between UTC and UTC+5:30.
Responsibilities
- Own end-to-end platform architecture, designs, and documentation.
- Lead and line-manage a distributed engineering team.
- Design, build, and operate a managed Slurm service for research and training workloads.
- Lead GPU infrastructure operations on bare-metal, including NVIDIA drivers and CUDA lifecycles.
- Manage Kubernetes cluster bootstrap, lifecycle, and GPU isolation.
- Define managed inference architecture, including scaling and request routing.
- Establish observability using metrics, logging, and SLOs.
- Lead incident response and on-call operations.
View Full Description & ApplyYou'll be redirected to the employer's site