- Improve accuracy on core document extractions by identifying and fixing failure modes with subject matter experts.
- Design and maintain evaluation suites to monitor system performance and prevent regressions.
- Improve the shared agent harness, including tooling, parallel tool calls, and context strategy.
- Design interfaces for human review of AI outputs to ensure explainability and facilitate feedback loops.
- Stay informed on the state of AI to integrate latest model developments and usage pattern best practices.