When an agent does something unexpected, you need to see exactly what happened. Observability makes agents debuggable.
What to Record
- The prompts and context sent at each step.
- Model responses, including tool calls.
- Tool inputs, outputs, durations and errors.
- Permission requests and decisions.
- Token usage and cost per step and per task.
- Final outcomes and user feedback.
Traces
Structure logs as traces: one task containing nested steps, tool calls and sub-agents. Trace viewers make long runs understandable.
Dashboards
Track success rates, average steps, costs, error rates and latency over time and by task type.
Debugging
Replay failed runs to see where they went wrong: a misleading tool result, a missing instruction, a context problem.
Privacy and Security
Traces contain user data, file contents and possibly secrets. Redact sensitive values, restrict access and set retention periods.
Close the Loop
Turn interesting failures into evaluation cases, and use trace analysis to guide improvements to prompts and tools.