LLM applications need monitoring beyond standard application metrics.
Operational Metrics
- Latency, including time to first token.
- Error rates and timeouts.
- Token usage and cost per request, user and feature.
- Rate limit hits.
Quality Signals
- User feedback: ratings, edits, regenerations.
- Automated evaluation on sampled traffic using rubrics or model graders.
- Retrieval quality for RAG: empty results, low relevance.
- Task completion rates.
Safety Signals
- Safety classifier flags.
- Refusal rates, both too high and too low.
- Prompt injection attempts.
- Personal data in outputs.
Tracing
Record prompts, retrieved context, tool calls and outputs for debugging, with redaction and retention limits.
Drift
User behaviour and topics change; provider models get updated. Watch for shifts in metrics.
Acting on Data
Review samples regularly, turn failures into test cases, and alert on sudden changes.