Senior HPC Systems Engineer
New
P
Parallel WorksHigh Performance Computing
Chicago, Illinois, United StatesFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 10 or more years operating production Linux systems
- Required Skills
- PythonBashKubernetesLinuxTerraformAnsible
Requirements
- 10 or more years of experience operating production Linux systems across multiple distribution families (RHEL, Debian, Ubuntu).
- Expertise in kernel and network tuning, systemd, cgroups, and NUMA.
- Production Slurm administration experience including configuration, debugging, and upgrades.
- At least one parallel or high throughput filesystem in production.
- Proficiency with InfiniBand or RoCE fabric operations.
- Experience with NVIDIA GPU node operations at multi-node scale, including driver stack management and fault triage.
- Experience with customer-owned or on-premises clusters and public cloud, including bare metal provisioning.
- DevOps experience with infrastructure as code frameworks such as Ansible or Terraform.
- Proficiency with Bash and Python scripting.
Responsibilities
- Build and operate production Slurm clusters including partitions, QOS, account management, and cgroup enforcement.
- Connect customer-owned clusters to the control plane, reconciling scheduler, storage, and identity sources.
- Perform bare metal provisioning, manage out-of-band hardware, and coordinate fault resolution with site staff.
- Validate GPU nodes including driver/CUDA stack management, NVLink checks, and NCCL tuning.
- Tune parallel filesystems and write infrastructure pipelines using Ansible, Terraform, and image build tools.
- Execute security hardening (STIG), remediate scans, manage FIPS cryptography, and handle Tier 3 escalations.
View Full Description & ApplyYou'll be redirected to the employer's site