AI Infrastructure Engineer
New
B
Bright Vision TechnologiesAI infrastructure
100% Remote (U.S.)Full-TimeSenior
Salary100,000 - 160,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- 10+ years
- Required Skills
- PythonKubernetesC++GoLinux
Requirements
- Bachelor’s or Master’s degree in Computer Science or a related field.
- Ten or more years of experience in infrastructure, platform, or HPC engineering.
- Hands-on experience operating GPU clusters or large-scale ML training infrastructure.
- Strong proficiency in Python and at least one systems language, such as Go or C++.
- Deep understanding of distributed training, accelerator architectures, and collective communication.
- Experience with Kubernetes, Slurm, Ray, or similar scheduling systems for ML workloads.
- Strong understanding of Linux internals, networking, and high-performance storage.
- Experience with at least one major cloud provider’s ML infrastructure offerings.
- Strong software engineering practices, including testing, CI/CD, and code review.
- Experience operating InfiniBand or RDMA networking at scale is preferred.
- Contributions to open-source ML infrastructure projects are preferred.
- Familiarity with custom orchestrators or research-grade training stacks is preferred.
Responsibilities
- Design and operate GPU and accelerator infrastructure for training and inference across on-prem clusters, cloud-managed services, and hybrid configurations.
- Build scheduling, queueing, and resource-sharing systems to maximize accelerator utilization across teams.
- Integrate PyTorch, JAX, DeepSpeed, FSDP, Megatron-LM, and Ray Train into a unified platform.
- Operate high-performance storage systems and data pipelines for training workloads.
- Design networking architectures that support RDMA, InfiniBand, NCCL, and high-bandwidth collective communication.
- Build observability for AI workloads, including utilization, throughput, training stability, and failure-mode analytics.
- Implement checkpointing, restart, and fault-tolerance patterns for large-scale training jobs.
- Optimize compute, storage, and networking costs through scheduling, spot capacity, and right-sizing.
- Develop developer tooling and workflows for researchers to launch experiments.
- Partner with research and applied ML teams to plan training-run capacity.
View Full Description & ApplyYou'll be redirected to the employer's site