Technical Support Engineer (Inference) - US Weekends

New
T
Together AIArtificial Intelligence
US daytime hours, US timezonesFull-TimeSenior
Salary$160K - $230K + equity + benefits
Apply NowOpens the employer's application page

Job Details

Experience
6+ years
Required Skills
PythonJavascriptKubernetesTypeScriptGrafanaPrometheusRESTful APIsDevOps

Requirements

  • 6+ years of experience in a customer-facing technical role, SRE, DevOps, or infrastructure engineering.
  • At least 1 year in a support role for an AI service.
  • Experience as an SRE or DevOps engineer working with Kubernetes.
  • Advanced, production-level experience with infrastructure services (Kubernetes, SLURM) and infrastructure as code solutions like Ansible.
  • Expertise in high-performance network fabrics and NFS-based storage management.
  • Strong technical background in AI, ML, GPU technologies, and high-performance computing (HPC) environments.
  • Proficiency in Python, TypeScript, and/or JavaScript with debugging experience using curl and Postman.
  • Demonstrated expertise with observability tools like Prometheus and Grafana.
  • Deep familiarity with REST API debugging and HTTP semantics.
  • Experience with LLM inference frameworks, LoRA fine-tuning, and GPU cluster management.
  • Familiarity with operating HPC storage systems such as Vast and Weka.
  • Must be available for a 4-day, 10-hour shift including Saturdays and Sundays.

Responsibilities

  • Engage directly with customers to resolve complex technical challenges involving GPU clusters, inference, and fine-tuning services.
  • Act as a customer-facing SRE to ensure Inference endpoints on Kubernetes remain healthy, stable, and performant.
  • Serve as the final technical expert before issues escalate to Engineering or Product teams.
  • Monitor system health, validate traffic routing during hardware migrations, and perform data-backed anomaly detection.
  • Manage customer-facing incident communications while translating technical findings into clear updates.
  • Execute infrastructure changes via pull requests (IaC) for model deployments, cluster configuration, and capacity scaling.
  • Identify patterns in support cases and collaborate with Engineering and Go-To-Market teams to drive the product roadmap.
View Full Description & ApplyYou'll be redirected to the employer's site
$160K - $230K + equity + benefits
Apply Now