HPC Architect

New
A
Andromeda ClusterAI Infrastructure
North America Remote / San Francisco, CAFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Required Skills
KubernetesLinux

Requirements

  • Deep HPC experience designing, building, or operating GPU clusters at scale.
  • Strong fabric knowledge, specifically InfiniBand and RoCE.
  • Experience with distributed orchestration and the HPC software stack, including Slurm, Kubernetes, and OpenMPI.
  • Solid understanding of Linux systems engineering.
  • Data-center literacy, including power, cooling, cabling, and physical-layer realities.
  • Strong benchmarking judgment for large-scale training workloads.
  • Ability to write precise, testable technical standards for engineering teams.
  • Credibility to deliver technical assessments and maintain strong external engineering relationships.
  • Experience inside a neocloud, hyperscaler, colocation, or data-center provider is a plus.
  • NVIDIA data-center GPU depth (DGX/HGX, NVLink/NVSwitch) is highly desirable.
  • Familiarity with AI research lab workload requirements.

Responsibilities

  • Vet prospective compute providers by assessing GPU hardware, network fabric, storage, and orchestration against quality metrics.
  • Define the qualification bar by building acceptance test suites, benchmark methodologies, and quality thresholds from scratch.
  • Run hands-on validation including burn-in testing, fabric validation, and application-level benchmarks.
  • Guide providers through technical onboarding to remediate gaps and tune configurations.
  • Partner with compute procurement to provide technical due diligence during sourcing.
  • Build and maintain technical relationships with existing providers to monitor for architectural or quality drift.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now