Senior HPC & AMD GPU Infrastructure Engineer
New
E
EvergridAI Infrastructure
New York City or U.SFull-TimeSenior
SalaryCompetitive base salary, Performance-based commission, Equity
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years
- Required Skills
- PythonBashPyTorchLinuxNetworking
Requirements
- 5+ years of experience in HPC, GPU cluster operations, or Linux systems engineering.
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
- Deep hands-on experience with AMD MI-series GPUs, including driver and kernel-level debugging.
- Strong understanding of Linux internals, hardware bring-up, and performance tuning.
- Proficiency in Bash and Python for automation and operational tooling.
- Deep familiarity with ML software stacks in ROCm environments (ROCm/HIP, MIOpen, RCCL, PyTorch, JAX).
- Experience securing production infrastructure, including VPNs, firewalls, and identity systems.
- Experience debugging high-performance networking and RDMA technologies such as InfiniBand or RoCE.
Responsibilities
- Serve as primary on-call responder for GPU failures, system outages, and cluster-wide incidents.
- Act as the key point person for customer GPU cluster reliability and SLA management.
- Design and maintain system monitoring for thermal, memory, and cluster-load health metrics.
- Coordinate with data center operators and hardware vendors for physical maintenance and RMAs.
- Install, patch, and maintain large fleets of Linux systems including kernel-level tuning.
- Configure secure networking, identity systems, and distributed storage such as Lustre or GPFS.
- Lead the deployment, BIOS configuration, and NUMA tuning of new GPU nodes.
- Debug and maintain the AMD ROCm-based software stack including PyTorch, JAX, and RCCL.
- Partner with ML and platform teams to optimize infrastructure for research and production workflows.
View Full Description & ApplyYou'll be redirected to the employer's site