Agents are harder to evaluate than single model calls: they take many steps, different paths can succeed, and outcomes depend on environments.
Build a Task Suite
Collect realistic tasks with clear success criteria: a bug to fix with tests that must pass, a question with a known answer, a form to complete correctly.
Grade Outcomes
- Automated checks: tests pass, the file contains the right values, the database is in the expected state.
- Model graders: judge qualities like helpfulness against a rubric, validated against human judgement.
- Human review: for nuanced or high-stakes tasks.
Look at Trajectories
Read the steps agents took, not just results. Trajectories reveal wasted steps, lucky successes, unsafe actions and tool misuse.
Measure More Than Success
Track cost, time, number of steps, errors, and whether the agent asked permission appropriately.
Run Multiple Times
Agents are non-deterministic; run each task several times and report success rates.
Reproducible Environments
Reset environments to a known state before each run so results are comparable.
Regression Testing
Run the suite whenever you change the model, prompts or tools. Improvements in one area often cause regressions in another.