Senior Network Engineer - InfiniBand / UFM

L
Lightning AIAI Infrastructure
This role may be fully remote within the U.S., or hybrid out of one of our office hubs (NYC, SF, Seattle, or London)Full-TimeSenior
Salary$170,000 — $210,000 USD
Apply NowOpens the employer's application page

Job Details

Experience
7+ years
Required Skills
PythonLinuxTerraformAnsible

Requirements

  • 7+ years of experience in large-scale data center networking.
  • Experience deploying, operating, and troubleshooting InfiniBand environments.
  • Strong understanding of Layer 2 and Layer 3 networking, including BGP, EVPN, VXLAN, and spine-leaf architectures.
  • Strong Linux administration experience.
  • Experience with network automation using Python, Ansible, Terraform, or similar tools.
  • Experience with network troubleshooting, observability, and telemetry tooling.
  • Excellent troubleshooting, documentation, and communication skills.

Responsibilities

  • Design, deploy, and maintain large-scale NVIDIA InfiniBand fabrics supporting AI/ML GPU clusters.
  • Deploy and administer NVIDIA Unified Fabric Manager (UFM) Enterprise for monitoring, provisioning, telemetry, and fabric health.
  • Configure and optimize NVIDIA Quantum and Quantum-2 InfiniBand switches.
  • Troubleshoot fabric performance issues impacting NCCL, MPI, GPUDirect RDMA, and AI training jobs.
  • Implement and validate fat-tree, Dragonfly+, Clos, and spine-leaf network architectures.
  • Perform firmware lifecycle management for InfiniBand switches, adapters (HCAs), and UFM infrastructure.
  • Optimize congestion control, adaptive routing, QoS, and traffic engineering for high-performance GPU communication.
  • Automate network provisioning using Python, Ansible, Git, REST APIs, and Infrastructure-as-Code methodologies.
View Full Description & ApplyYou'll be redirected to the employer's site
$170,000 — $210,000 USD
Apply Now