Senior HPC Cluster Engineer - AI, ML

New
J
JobgetherAI, HPC Infrastructure
Based in IndiaFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
5+ years
Required Skills
DockerPythonBashKubernetesLinuxTerraformAnsible

Requirements

  • Bachelor's degree in Computer Science, Electrical Engineering, or a related discipline, or equivalent practical experience.
  • 5+ years of experience designing, deploying, and operating large-scale compute or infrastructure environments.
  • Strong experience with AI/HPC job schedulers such as Slurm, Kubernetes, PBS, RTDA, BCM, or LSF.
  • Proficiency administering Linux distributions such as CentOS/RHEL and/or Ubuntu.
  • Hands-on experience with cluster configuration and infrastructure management tools including BCM, Terraform, Ansible, Puppet, Salt, or similar technologies.
  • Strong knowledge of container technologies such as Docker, Singularity, Podman, Shifter, or Charliecloud.
  • Proficiency with Python programming and Bash scripting for infrastructure automation and operational tooling.
  • Practical experience supporting AI/HPC workflows that use MPI.
  • Experience analyzing and tuning the performance of diverse AI and HPC workloads.
  • Strong troubleshooting, analytical, and root-cause analysis skills.
  • Excellent collaboration and communication skills, particularly when working with researchers, infrastructure engineers, and globally distributed teams.

Responsibilities

  • Provide technical leadership for systems administration and service delivery across large-scale AI/HPC infrastructure, coordinating upgrades, incident response, and reliability improvements.
  • Own the day-to-day operation of production AI/HPC clusters, monitoring system health, user experience, resource utilization, and adherence to internal service-level targets.
  • Build, maintain, and scale heterogeneous AI/ML clusters across on-premises and cloud environments, covering compute, networking, storage, and GPU-accelerated infrastructure.
  • Develop scalable automation and tooling to improve the deployment, configuration, management, and operational efficiency of AI/HPC environments.
  • Collaborate with global engineering teams and internal users to understand evolving research and computing requirements and deliver reliable infrastructure solutions.
  • Support researchers in running complex AI/HPC workloads, including performance analysis, troubleshooting, optimization, and workload tuning.
  • Analyze cluster efficiency, job fragmentation, and GPU utilization to identify opportunities to reduce resource waste and improve overall capacity.
  • Conduct root cause analysis for infrastructure and service issues, proactively identifying risks and implementing corrective actions before they affect users.
  • Lead SEV triage, incident response, and postmortems for reliability events affecting production infrastructure or users.
  • Participate in an on-call rotation and provide timely support for critical production GPU clusters.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now