Forward Deployed Engineer - SRE

New
A
Andromeda ClusterAI Infrastructure
North America Remote / San Francisco, CAFull-TimeMiddle
SalaryCompetitive compensation + meaningful equity
Apply NowOpens the employer's application page

Job Details

Required Skills
PythonBashKubernetesGoLinuxTerraform

Requirements

  • Hands-on experience operating GPU clusters (NVIDIA A100/H100/H200/B200) in production environments.
  • Production experience with high-speed interconnects such as InfiniBand, RoCE, or NVLink.
  • Expert-level Linux experience including kernel tuning, driver management, and performance profiling.
  • Strong proficiency running Kubernetes in production with GPU workloads, or equivalent experience with Slurm.
  • Advanced programming skills in Python, Go, or Bash to build production-grade tools.
  • Infrastructure-as-Code proficiency using tools like Terraform, Helm, or Ansible.
  • Experience building monitoring and alerting systems for GPU telemetry.
  • Proven track record of leading incident response for complex distributed systems.
  • Ability to articulate technical trade-offs to engineering leadership.

Responsibilities

  • Serve as the primary technical point of contact for teams running large-scale training and inference workloads.
  • Own end-to-end onboarding, including environment setup, orchestration, and storage layout.
  • Diagnose and resolve complex real-time failures such as NCCL timeouts, checkpoint I/O stalls, and driver mismatches.
  • Profile and improve distributed training performance, focusing on MFU and time-to-first-successful-run.
  • Ensure the health of high-speed interconnects (InfiniBand, RoCE, NVLink) and build GPU-specific monitoring dashboards.
  • Automate cluster provisioning, health checks, and preflight validation to eliminate repeated issues.
  • Lead incident response for multi-layer failures and manage customer-facing communication.
View Full Description & ApplyYou'll be redirected to the employer's site
Competitive compensation + meaningful equity
Apply Now