HPC Architect
New
A
Andromeda ClusterAI Infrastructure
North America Remote / San Francisco, CAFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Required Skills
- KubernetesLinux
Requirements
- Deep HPC experience designing, building, or operating GPU clusters at scale.
- Strong fabric knowledge, specifically InfiniBand and RoCE.
- Experience with distributed orchestration and the HPC software stack, including Slurm, Kubernetes, and OpenMPI.
- Solid understanding of Linux systems engineering.
- Data-center literacy, including power, cooling, cabling, and physical-layer realities.
- Strong benchmarking judgment for large-scale training workloads.
- Ability to write precise, testable technical standards for engineering teams.
- Credibility to deliver technical assessments and maintain strong external engineering relationships.
- Experience inside a neocloud, hyperscaler, colocation, or data-center provider is a plus.
- NVIDIA data-center GPU depth (DGX/HGX, NVLink/NVSwitch) is highly desirable.
- Familiarity with AI research lab workload requirements.
Responsibilities
- Vet prospective compute providers by assessing GPU hardware, network fabric, storage, and orchestration against quality metrics.
- Define the qualification bar by building acceptance test suites, benchmark methodologies, and quality thresholds from scratch.
- Run hands-on validation including burn-in testing, fabric validation, and application-level benchmarks.
- Guide providers through technical onboarding to remediate gaps and tune configurations.
- Partner with compute procurement to provide technical due diligence during sourcing.
- Build and maintain technical relationships with existing providers to monitor for architectural or quality drift.
View Full Description & ApplyYou'll be redirected to the employer's site