- Validate task quality by checking instructions, source materials, reference solutions, and evaluation criteria for consistency and completeness.
- Review AI agent execution traces, tool calls, and deliverables to determine whether outcomes are justified.
- Audit grading logic to identify brittle checks, incorrect expected answers, unsupported rubric criteria, and unfair penalties for valid alternative solutions.
- Investigate discrepancies between model performance, grader results, and expected outcomes.
- Distinguish genuine model limitations from task defects, grader errors, and environment or tool failures.
- Independently assess automated QC findings rather than accepting them without verification.
- Document concise, evidence-backed findings and provide actionable, reproducible feedback.
- Flag uncertainty and verify that implemented revisions resolve previously identified issues.
PythonSQLAnalytical Skills+1 more