AI Infrastructure Engineer

New
B
Bright Vision TechnologiesAI infrastructure
100% Remote (U.S.)Full-TimeSenior
Salary100,000 - 160,000 USD per year
Apply NowOpens the employer's application page

Job Details

Experience
10+ years
Required Skills
PythonKubernetesC++GoLinux

Requirements

  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • Ten or more years of experience in infrastructure, platform, or HPC engineering.
  • Hands-on experience operating GPU clusters or large-scale ML training infrastructure.
  • Strong proficiency in Python and at least one systems language, such as Go or C++.
  • Deep understanding of distributed training, accelerator architectures, and collective communication.
  • Experience with Kubernetes, Slurm, Ray, or similar scheduling systems for ML workloads.
  • Strong understanding of Linux internals, networking, and high-performance storage.
  • Experience with at least one major cloud provider’s ML infrastructure offerings.
  • Strong software engineering practices, including testing, CI/CD, and code review.
  • Experience operating InfiniBand or RDMA networking at scale is preferred.
  • Contributions to open-source ML infrastructure projects are preferred.
  • Familiarity with custom orchestrators or research-grade training stacks is preferred.

Responsibilities

  • Design and operate GPU and accelerator infrastructure for training and inference across on-prem clusters, cloud-managed services, and hybrid configurations.
  • Build scheduling, queueing, and resource-sharing systems to maximize accelerator utilization across teams.
  • Integrate PyTorch, JAX, DeepSpeed, FSDP, Megatron-LM, and Ray Train into a unified platform.
  • Operate high-performance storage systems and data pipelines for training workloads.
  • Design networking architectures that support RDMA, InfiniBand, NCCL, and high-bandwidth collective communication.
  • Build observability for AI workloads, including utilization, throughput, training stability, and failure-mode analytics.
  • Implement checkpointing, restart, and fault-tolerance patterns for large-scale training jobs.
  • Optimize compute, storage, and networking costs through scheduling, spot capacity, and right-sizing.
  • Develop developer tooling and workflows for researchers to launch experiments.
  • Partner with research and applied ML teams to plan training-run capacity.
View Full Description & ApplyYou'll be redirected to the employer's site
100,000 - 160,000 USD per year
Apply Now