Forward Deployed Engineer - SRE
New
A
Andromeda ClusterAI Infrastructure
North America Remote / San Francisco, CAFull-TimeMiddle
SalaryCompetitive compensation + meaningful equity
Apply NowOpens the employer's application page
Job Details
- Required Skills
- PythonBashKubernetesGoLinuxTerraform
Requirements
- Hands-on experience operating GPU clusters (NVIDIA A100/H100/H200/B200) in production environments.
- Production experience with high-speed interconnects such as InfiniBand, RoCE, or NVLink.
- Expert-level Linux experience including kernel tuning, driver management, and performance profiling.
- Strong proficiency running Kubernetes in production with GPU workloads, or equivalent experience with Slurm.
- Advanced programming skills in Python, Go, or Bash to build production-grade tools.
- Infrastructure-as-Code proficiency using tools like Terraform, Helm, or Ansible.
- Experience building monitoring and alerting systems for GPU telemetry.
- Proven track record of leading incident response for complex distributed systems.
- Ability to articulate technical trade-offs to engineering leadership.
Responsibilities
- Serve as the primary technical point of contact for teams running large-scale training and inference workloads.
- Own end-to-end onboarding, including environment setup, orchestration, and storage layout.
- Diagnose and resolve complex real-time failures such as NCCL timeouts, checkpoint I/O stalls, and driver mismatches.
- Profile and improve distributed training performance, focusing on MFU and time-to-first-successful-run.
- Ensure the health of high-speed interconnects (InfiniBand, RoCE, NVLink) and build GPU-specific monitoring dashboards.
- Automate cluster provisioning, health checks, and preflight validation to eliminate repeated issues.
- Lead incident response for multi-layer failures and manage customer-facing communication.
View Full Description & ApplyYou'll be redirected to the employer's site