- Own end-to-end development of agents and agentic features, defining requirements, success criteria, evaluations, guardrails, and tools with engineering, design, and product teams.
- Lead model selection, development, and validation by assessing whether models or configurations perform better.
- Design and run A/B tests and other experiments to measure AI feature impact on customers.
- Set and evolve quality standards across agents in partnership with engineers building agents and evaluation infrastructure.
- Improve agent platform capabilities through prompt iteration, reusable judges, context-window management, tool use, and adoption of new models.
- Communicate AI capabilities, limitations, and quality to product, design, engineering, and leadership.