Not every model makes a good agent. Agentic work demands specific capabilities.
What Matters
- Tool-use reliability: correct tool choice and well-formed arguments.
- Instruction following: respecting constraints and permissions across many steps.
- Long-context handling: staying coherent as history grows.
- Recovery: noticing errors and changing approach.
- Honesty: reporting failures rather than claiming false success.
- Reasoning: planning multi-step work.
Evaluate on Your Tasks
General benchmarks give a rough guide, but run candidate models on your own agent tasks and tools to compare success rate, cost and time.
Mixing Models
- A capable model for planning and hard steps.
- Smaller, faster models for simple sub-tasks, classification or summarisation.
- Specialised models for embeddings or vision.
Cost per Task
Per-token prices mislead. A stronger model that succeeds in fewer steps can be cheaper overall than a weaker one that flounders.
Re-evaluate
Models improve quickly. Keep your evaluation suite ready so you can test new models and switch when it pays off.