Technical Support Engineer (GPU Clusters)
New
T
Together AIArtificial Intelligence
US, US daytime hoursFull-TimeMiddle
Salary$160K - $230K + equity + benefits
Apply NowOpens the employer's application page
Job Details
- Experience
- 3+ years of experience in a customer-facing technical role with at least 1 year in a support function
- Required Skills
- KubernetesDevOpsAnsibleNetworking
Requirements
- 3+ years of experience in a customer-facing technical role.
- At least 1 year in a support function for an AI service or supporting a mission-critical API in SaaS.
- Experience as an SRE or DevOps engineer working with Kubernetes.
- Strong technical background in AI, ML, GPU technologies, and high-performance computing (HPC) environments.
- Advanced knowledge with infrastructure services (e.g., Kubernetes, SLURM) and infrastructure as code solutions (e.g., Ansible).
- Experience with HPC/Slurm cluster environments including node draining, job scheduling, and maintenance.
- Familiarity with high-speed networking concepts including InfiniBand, RDMA, and interface diagnostics.
- Experience with distributed storage systems such as Weka or NFS.
- Foundational understanding in the installation, configuration, administration, and securing of compute clusters.
- Ability to work a 4-day, 10-hour shift including weekend days.
- Excellent communication skills for explaining complex technical concepts to non-technical stakeholders.
Responsibilities
- Engage directly with customers to tackle and resolve complex technical challenges involving our cutting-edge Kubernetes GPU clusters.
- Act as a customer facing SRE to ensure our customer’s Kubernetes clusters remain healthy and stable.
- Become a product expert in our GPU Cluster service, serving as the last line of technical defense.
- Monitor GPU cluster health and proactively communicate hardware issues with clear remediation steps.
- Operate and maintain production infrastructure for enterprise GPU customers, including fleet rebalancing, Slurm cluster maintenance, node repair/migration, and Kubernetes-based workload management.
- Investigate and resolve storage and networking issues such as Weka filesystem degradation, InfiniBand link failures, and bandwidth anomalies.
- Collaborate across Engineering, Research, and Product teams to address customer concerns and drive roadmap improvements.
- Maintain detailed documentation of system configurations, procedures, troubleshooting guides, and FAQs.
View Full Description & ApplyYou'll be redirected to the employer's site