Senior Manager, Cluster Engineering & Deployment
New
T
TensorWaveData Center / AI
Location: RemoteFull-TimeManager
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 10+ years
- Required Skills
- PythonAnsible
Requirements
- 10+ years of experience in network deployment, cluster/HPC bring-up, or large-scale infrastructure delivery.
- Proven experience managing engineers in a field or deployment setting.
- Hands-on fabric bring-up experience at scale with hundreds of switches and thousands of links.
- Strong operational rigor with the ability to build and enforce playbooks, gates, and metrics.
- Team leadership experience managing schedule accountability across multiple concurrent builds or sites.
- GPU cluster validation experience (NCCL/RCCL benchmarking) preferred.
- Automation skills (Python, Ansible) applied to deployment preferred.
- Optics/link-layer debugging depth preferred.
- Experience with acceptance testing as a commercial gate preferred.
Responsibilities
- Own the cluster deployment playbook and drive its evolution including staged bring-up, automated config push, and fault triage.
- Lead deployment engineering across concurrent cluster builds and coordinate daily with Data Center Integration teams and vendors.
- Drive deployment velocity by reducing bring-up time per cluster through tooling and pre-staging.
- Manage defect feedback loops to Network Engineering and hardware vendors to ensure resolution.
- Define and standardize requirements for spares, test equipment, and deployment tooling across sites.
- Oversee end-to-end cluster validation including performance benchmarking and burn-in criteria to maintain acceptance gates.
View Full Description & ApplyYou'll be redirected to the employer's site