HPC Support Engineer

New
L
LambdaAI Cloud Infrastructure
Remote, USAFull-TimeSenior
Salary122,000 - 162,000 USD per year
Apply NowOpens the employer's application page

Job Details

Experience
3+ years
Required Skills
KubernetesGrafanaPrometheusLinuxDatadog

Requirements

  • 3+ years of hands-on HPC experience in administration, support, or engineering.
  • Strong experience supporting Linux in a system administration role.
  • Expertise in HPC environments, specifically Linux cluster administration.
  • Strong preference for experience with Kubernetes and/or Slurm for cluster orchestration.
  • Strong coding ability and CI/CD experience.
  • Proficiency with monitoring and logging tools such as Prometheus, Grafana, and Datadog.
  • Skills in log analysis, kernel-level debugging, and performance profiling.
  • Experience with CUDA, NCCL, NVLink, and GPUDirect RDMA.
  • Experience with high-throughput networking technologies (IB/RoCE).
  • Knowledge of distributed AI/ML or HPC workloads, TCP/IP, VPN, and firewalls.

Responsibilities

  • Serve as a senior technical escalation point, troubleshooting infrastructure and platform issues at hardware, driver, or kernel levels.
  • Distinguish between hardware, driver, kernel, and configuration failures to ensure correct resolution.
  • Identify and remediate process, tooling, and documentation gaps.
  • Develop scripts and automations using AI tools to address operational gaps.
  • Perform root-cause analysis across distributed systems and GPU infrastructure.
  • Collaborate with engineering teams to permanently resolve customer issues.
  • Mentor junior support engineers and lead major incident responses.
  • Participate in a rotating on-call schedule.
View Full Description & ApplyYou'll be redirected to the employer's site
122,000 - 162,000 USD per year
Apply Now