Engage directly with customers to resolve complex technical challenges involving GPU clusters, inference, and fine-tuning services.
Act as a customer-facing SRE to ensure Inference endpoints on Kubernetes remain healthy, stable, and performant.
Serve as the final technical expert before issues escalate to Engineering or Product teams.
Monitor system health, validate traffic routing during hardware migrations, and perform data-backed anomaly detection.
Manage customer-facing incident communications while translating technical findings into clear updates.
Execute infrastructure changes via pull requests (IaC) for model deployments, cluster configuration, and capacity scaling.
Identify patterns in support cases and collaborate with Engineering and Go-To-Market teams to drive the product roadmap.
PythonJavascriptKubernetes+5 more
Showing 1 of 2 positions
About Together AI
Together AI is building the open-source future of generative AI. You will join a company that has raised $1.33 billion, reaching an $8.3 billion valuation in just four years. Together AI offers a cloud-based platform for developers and researchers to train, fine-tune, and deploy generative AI models. Their "AI Acceleration Cloud" provides GPU infrastructure, model inference services, and custom model building. Customers like Salesforce and Zoom trust Together AI to power their AI innovations, leveraging over 8,000 GPUs for up to 20 exaflops of compute power. They deliver 2x latency improvements and 33% cost reductions for enterprise clients.
How We Work
Together AI embraces a remote-first culture. You can work from anywhere, collaborating with colleagues across different time zones. The company avoids mandatory return-to-office policies, fostering a flexible environment. Team members often choose to meet in person for voluntary collaboration. This approach promotes work-life balance and empowers you to manage your own schedule.
Engineering at Together AI
You will tackle complex challenges at the forefront of AI infrastructure. Engineers at Together AI build and maintain high-performance computing (HPC) environments. This includes managing cutting-edge Kubernetes GPU clusters and optimizing inference services. You will work with technologies like SLURM, Ansible, InfiniBand, RDMA, Weka, and NFS. The team also contributes to leading open-source research and models. Projects like FlashAttention, Hyena, FlexGen, and RedPajama demonstrate their commitment to advancing AI.
Why Join Us
Pioneer the future of open-source generative AI with significant impact.
Solve complex technical problems at an $8.3 billion valued company.
Work in a truly remote-first environment with flexible schedules.
Contribute to groundbreaking open-source research and cutting-edge models.
Join a team that has secured $1.33 billion in funding from top investors.