Skip to content

Choosing Models for Agent Harnesses

What makes a model good at agentic work, and how to pick and mix models for different parts of an agent.

Editorial team 1 min read

Not every model makes a good agent. Agentic work demands specific capabilities.

What Matters

  • Tool-use reliability: correct tool choice and well-formed arguments.
  • Instruction following: respecting constraints and permissions across many steps.
  • Long-context handling: staying coherent as history grows.
  • Recovery: noticing errors and changing approach.
  • Honesty: reporting failures rather than claiming false success.
  • Reasoning: planning multi-step work.

Evaluate on Your Tasks

General benchmarks give a rough guide, but run candidate models on your own agent tasks and tools to compare success rate, cost and time.

Mixing Models

  • A capable model for planning and hard steps.
  • Smaller, faster models for simple sub-tasks, classification or summarisation.
  • Specialised models for embeddings or vision.

Cost per Task

Per-token prices mislead. A stronger model that succeeds in fewer steps can be cheaper overall than a weaker one that flounders.

Re-evaluate

Models improve quickly. Keep your evaluation suite ready so you can test new models and switch when it pays off.

More in Agent harnesses

All Agent harnesses guides →
Agent harnesses Guide · 2 min

What Is an Agent Harness?

The software around a language model that turns it into an agent: the loop, tools, context, permissions and memory.

Agent harnesses 2 min read 27 Sep 2025

Agent harnesses Guide · 1 min

The Agent Loop Explained

The core cycle every agent runs: think, call a tool, observe the result, repeat — and how the loop knows when to stop.

Agent harnesses 1 min read 26 Sep 2025

Agent harnesses Guide · 1 min

Designing Tools for AI Agents

How to write tools agents use well: clear names, precise descriptions, sensible inputs and informative outputs.

Agent harnesses 1 min read 25 Sep 2025

Agent harnesses Guide · 1 min

Context Management in Agent Harnesses

How agents stay effective over long tasks: what to keep in context, what to summarise, and what to store outside.

Agent harnesses 1 min read 24 Sep 2025