- Help build and maintain the core infrastructure that powers Quora's ML platform, ensuring high availability, scalability, and performance
- Build and improve the distributed systems that serve our ML models in production, from Large Recommendation Models (LRMs) to Large Language Models (LLMs)
- Work on GPU model serving, optimizing latency, throughput, and cost to support larger and more capable models
- Contribute to platform initiatives such as PyTorch-first standardization and ML ecosystem modernization
- Improve ML developer velocity by building tooling that helps ML engineers develop, test, and deploy models more efficiently
- Modernize our feature store so ML engineers can get new features into production faster
- Participate in the team's on-call rotation, helping resolve production issues as you grow your knowledge and ownership of the platform
AWSDockerPython+6 more