Technical Support Engineer (Inference) - US Weekends
New
T
Together AIArtificial Intelligence
US daytime hours, US timezonesFull-TimeSenior
Salary$160K - $230K + equity + benefits
Apply NowOpens the employer's application page
Job Details
- Experience
- 6+ years
- Required Skills
- PythonJavascriptKubernetesTypeScriptGrafanaPrometheusRESTful APIsDevOps
Requirements
- 6+ years of experience in a customer-facing technical role, SRE, DevOps, or infrastructure engineering.
- At least 1 year in a support role for an AI service.
- Experience as an SRE or DevOps engineer working with Kubernetes.
- Advanced, production-level experience with infrastructure services (Kubernetes, SLURM) and infrastructure as code solutions like Ansible.
- Expertise in high-performance network fabrics and NFS-based storage management.
- Strong technical background in AI, ML, GPU technologies, and high-performance computing (HPC) environments.
- Proficiency in Python, TypeScript, and/or JavaScript with debugging experience using curl and Postman.
- Demonstrated expertise with observability tools like Prometheus and Grafana.
- Deep familiarity with REST API debugging and HTTP semantics.
- Experience with LLM inference frameworks, LoRA fine-tuning, and GPU cluster management.
- Familiarity with operating HPC storage systems such as Vast and Weka.
- Must be available for a 4-day, 10-hour shift including Saturdays and Sundays.
Responsibilities
- Engage directly with customers to resolve complex technical challenges involving GPU clusters, inference, and fine-tuning services.
- Act as a customer-facing SRE to ensure Inference endpoints on Kubernetes remain healthy, stable, and performant.
- Serve as the final technical expert before issues escalate to Engineering or Product teams.
- Monitor system health, validate traffic routing during hardware migrations, and perform data-backed anomaly detection.
- Manage customer-facing incident communications while translating technical findings into clear updates.
- Execute infrastructure changes via pull requests (IaC) for model deployments, cluster configuration, and capacity scaling.
- Identify patterns in support cases and collaborate with Engineering and Go-To-Market teams to drive the product roadmap.
View Full Description & ApplyYou'll be redirected to the employer's site