Senior Network Engineer - InfiniBand / UFM
L
Lightning AIAI Infrastructure
This role may be fully remote within the U.S., or hybrid out of one of our office hubs (NYC, SF, Seattle, or London)Full-TimeSenior
Salary$170,000 — $210,000 USD
Apply NowOpens the employer's application page
Job Details
- Experience
- 7+ years
- Required Skills
- PythonLinuxTerraformAnsible
Requirements
- 7+ years of experience in large-scale data center networking.
- Experience deploying, operating, and troubleshooting InfiniBand environments.
- Strong understanding of Layer 2 and Layer 3 networking, including BGP, EVPN, VXLAN, and spine-leaf architectures.
- Strong Linux administration experience.
- Experience with network automation using Python, Ansible, Terraform, or similar tools.
- Experience with network troubleshooting, observability, and telemetry tooling.
- Excellent troubleshooting, documentation, and communication skills.
Responsibilities
- Design, deploy, and maintain large-scale NVIDIA InfiniBand fabrics supporting AI/ML GPU clusters.
- Deploy and administer NVIDIA Unified Fabric Manager (UFM) Enterprise for monitoring, provisioning, telemetry, and fabric health.
- Configure and optimize NVIDIA Quantum and Quantum-2 InfiniBand switches.
- Troubleshoot fabric performance issues impacting NCCL, MPI, GPUDirect RDMA, and AI training jobs.
- Implement and validate fat-tree, Dragonfly+, Clos, and spine-leaf network architectures.
- Perform firmware lifecycle management for InfiniBand switches, adapters (HCAs), and UFM infrastructure.
- Optimize congestion control, adaptive routing, QoS, and traffic engineering for high-performance GPU communication.
- Automate network provisioning using Python, Ansible, Git, REST APIs, and Infrastructure-as-Code methodologies.
View Full Description & ApplyYou'll be redirected to the employer's site