Senior Manager, Cluster Engineering & Deployment

New
T
TensorWaveData Center / AI
Location: RemoteFull-TimeManager
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
10+ years
Required Skills
PythonAnsible

Requirements

  • 10+ years of experience in network deployment, cluster/HPC bring-up, or large-scale infrastructure delivery.
  • Proven experience managing engineers in a field or deployment setting.
  • Hands-on fabric bring-up experience at scale with hundreds of switches and thousands of links.
  • Strong operational rigor with the ability to build and enforce playbooks, gates, and metrics.
  • Team leadership experience managing schedule accountability across multiple concurrent builds or sites.
  • GPU cluster validation experience (NCCL/RCCL benchmarking) preferred.
  • Automation skills (Python, Ansible) applied to deployment preferred.
  • Optics/link-layer debugging depth preferred.
  • Experience with acceptance testing as a commercial gate preferred.

Responsibilities

  • Own the cluster deployment playbook and drive its evolution including staged bring-up, automated config push, and fault triage.
  • Lead deployment engineering across concurrent cluster builds and coordinate daily with Data Center Integration teams and vendors.
  • Drive deployment velocity by reducing bring-up time per cluster through tooling and pre-staging.
  • Manage defect feedback loops to Network Engineering and hardware vendors to ensure resolution.
  • Define and standardize requirements for spares, test equipment, and deployment tooling across sites.
  • Oversee end-to-end cluster validation including performance benchmarking and burn-in criteria to maintain acceptance gates.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now