Senior HPC Cluster Engineer - AI, ML
New
J
JobgetherAI, HPC Infrastructure
Based in IndiaFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years
- Required Skills
- DockerPythonBashKubernetesLinuxTerraformAnsible
Requirements
- Bachelor's degree in Computer Science, Electrical Engineering, or a related discipline, or equivalent practical experience.
- 5+ years of experience designing, deploying, and operating large-scale compute or infrastructure environments.
- Strong experience with AI/HPC job schedulers such as Slurm, Kubernetes, PBS, RTDA, BCM, or LSF.
- Proficiency administering Linux distributions such as CentOS/RHEL and/or Ubuntu.
- Hands-on experience with cluster configuration and infrastructure management tools including BCM, Terraform, Ansible, Puppet, Salt, or similar technologies.
- Strong knowledge of container technologies such as Docker, Singularity, Podman, Shifter, or Charliecloud.
- Proficiency with Python programming and Bash scripting for infrastructure automation and operational tooling.
- Practical experience supporting AI/HPC workflows that use MPI.
- Experience analyzing and tuning the performance of diverse AI and HPC workloads.
- Strong troubleshooting, analytical, and root-cause analysis skills.
- Excellent collaboration and communication skills, particularly when working with researchers, infrastructure engineers, and globally distributed teams.
Responsibilities
- Provide technical leadership for systems administration and service delivery across large-scale AI/HPC infrastructure, coordinating upgrades, incident response, and reliability improvements.
- Own the day-to-day operation of production AI/HPC clusters, monitoring system health, user experience, resource utilization, and adherence to internal service-level targets.
- Build, maintain, and scale heterogeneous AI/ML clusters across on-premises and cloud environments, covering compute, networking, storage, and GPU-accelerated infrastructure.
- Develop scalable automation and tooling to improve the deployment, configuration, management, and operational efficiency of AI/HPC environments.
- Collaborate with global engineering teams and internal users to understand evolving research and computing requirements and deliver reliable infrastructure solutions.
- Support researchers in running complex AI/HPC workloads, including performance analysis, troubleshooting, optimization, and workload tuning.
- Analyze cluster efficiency, job fragmentation, and GPU utilization to identify opportunities to reduce resource waste and improve overall capacity.
- Conduct root cause analysis for infrastructure and service issues, proactively identifying risks and implementing corrective actions before they affect users.
- Lead SEV triage, incident response, and postmortems for reliability events affecting production infrastructure or users.
- Participate in an on-call rotation and provide timely support for critical production GPU clusters.
View Full Description & ApplyYou'll be redirected to the employer's site