- Build and own evaluation infrastructure including CI/CD pipelines, scorers, and datasets for AI agents.
- Diagnose failures across agent pipelines including plan creation, tool selection, and execution trajectories.
- Grow datasets via annotation workflows and AI-generated/simulated conversations.
- Run structured experiments across prompts, models, and the agent harness to optimize cost, latency, and performance.
- Lead and grow the AI Quality engineering team while remaining hands-on.
- Partner with AI Core engineering to validate product changes and quality improvements.
PythonMachine LearningRuby on Rails+2 more