- Architect and build core AI infrastructure for a platform serving millions of users.
- Implement LLM-backed features including retrieval-augmented generation (RAG), streaming responses, and tool usage.
- Develop systems for model routing, failover, token streaming, and cost/usage metering.
- Build and maintain retrieval systems including vector search, embedding pipelines, and context assembly.
- Design and implement instrumentation for monitoring failure modes like hallucinations and provider outages.
- Define evaluation processes to measure AI answer quality, relevance, and safety using repeatable tests.
- Own services end-to-end, from API design and schema to deployment and on-call support.
PythonMachine LearningGo+1 more