Senior HPC Systems Engineer

New
P
Parallel WorksHigh Performance Computing
Chicago, Illinois, United StatesFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
10 or more years operating production Linux systems
Required Skills
PythonBashKubernetesLinuxTerraformAnsible

Requirements

  • 10 or more years of experience operating production Linux systems across multiple distribution families (RHEL, Debian, Ubuntu).
  • Expertise in kernel and network tuning, systemd, cgroups, and NUMA.
  • Production Slurm administration experience including configuration, debugging, and upgrades.
  • At least one parallel or high throughput filesystem in production.
  • Proficiency with InfiniBand or RoCE fabric operations.
  • Experience with NVIDIA GPU node operations at multi-node scale, including driver stack management and fault triage.
  • Experience with customer-owned or on-premises clusters and public cloud, including bare metal provisioning.
  • DevOps experience with infrastructure as code frameworks such as Ansible or Terraform.
  • Proficiency with Bash and Python scripting.

Responsibilities

  • Build and operate production Slurm clusters including partitions, QOS, account management, and cgroup enforcement.
  • Connect customer-owned clusters to the control plane, reconciling scheduler, storage, and identity sources.
  • Perform bare metal provisioning, manage out-of-band hardware, and coordinate fault resolution with site staff.
  • Validate GPU nodes including driver/CUDA stack management, NVLink checks, and NCCL tuning.
  • Tune parallel filesystems and write infrastructure pipelines using Ansible, Terraform, and image build tools.
  • Execute security hardening (STIG), remediate scans, manage FIPS cryptography, and handle Tier 3 escalations.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now