HPC Support Engineer
New
L
LambdaAI Cloud Infrastructure
Remote, USAFull-TimeSenior
Salary122,000 - 162,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- 3+ years
- Required Skills
- KubernetesGrafanaPrometheusLinuxDatadog
Requirements
- 3+ years of hands-on HPC experience in administration, support, or engineering.
- Strong experience supporting Linux in a system administration role.
- Expertise in HPC environments, specifically Linux cluster administration.
- Strong preference for experience with Kubernetes and/or Slurm for cluster orchestration.
- Strong coding ability and CI/CD experience.
- Proficiency with monitoring and logging tools such as Prometheus, Grafana, and Datadog.
- Skills in log analysis, kernel-level debugging, and performance profiling.
- Experience with CUDA, NCCL, NVLink, and GPUDirect RDMA.
- Experience with high-throughput networking technologies (IB/RoCE).
- Knowledge of distributed AI/ML or HPC workloads, TCP/IP, VPN, and firewalls.
Responsibilities
- Serve as a senior technical escalation point, troubleshooting infrastructure and platform issues at hardware, driver, or kernel levels.
- Distinguish between hardware, driver, kernel, and configuration failures to ensure correct resolution.
- Identify and remediate process, tooling, and documentation gaps.
- Develop scripts and automations using AI tools to address operational gaps.
- Perform root-cause analysis across distributed systems and GPU infrastructure.
- Collaborate with engineering teams to permanently resolve customer issues.
- Mentor junior support engineers and lead major incident responses.
- Participate in a rotating on-call schedule.
View Full Description & ApplyYou'll be redirected to the employer's site