Text classification assigns categories to text: routing support tickets, tagging documents, detecting spam, flagging urgent emails.
Define the Categories Well
Write clear definitions with examples, decide how to handle texts that fit several or none, and check that two people would agree on labels. Ambiguous categories cap achievable accuracy.
Approaches, From Simple to Powerful
- Keyword rules: fast to build and a good baseline.
- TF-IDF + logistic regression: strong, fast and explainable with a few thousand labelled examples.
- Embeddings + a simple classifier: good accuracy with fewer labels.
- Fine-tuned transformer: high accuracy with enough labelled data.
- LLM with instructions and examples: works with little or no labelled data and handles new categories easily, at higher per-request cost.
Labelled Data
Label a representative sample, including hard cases. For LLM approaches, you still need labelled examples to evaluate, even if not to train.
Evaluation
Use per-class precision and recall and the confusion matrix to see which categories are confused. Check performance on recent data, since topics shift over time.
Handling Uncertainty
Route low-confidence predictions to a person, and use their decisions as new training or evaluation data.
Monitoring
New products, issues and phrasing appear over time. Track the share of low-confidence predictions and periodically re-label a sample to check accuracy.