Before deep learning, text was converted into numerical features for classical models. These methods remain fast, effective baselines.
Preprocessing
- Lowercasing and removing punctuation.
- Tokenisation into words.
- Optionally removing common stop words.
- Stemming or lemmatisation to group word forms.
Bag of Words
Count how often each word appears in each document. Simple but surprisingly effective.
TF-IDF
Weight words by how frequent they are in a document relative to how common they are across all documents. Distinctive words gain importance.
N-Grams
Include word pairs or triples ("not good", "credit card") to capture some word order.
Models
Logistic regression, linear SVMs and naive Bayes work well on these sparse, high-dimensional features.
Embeddings
Embeddings from pretrained language models capture meaning and synonyms better, and usually perform better with limited labelled data.
When to Use Which
- Large labelled data, need for speed and interpretability → TF-IDF with a linear model.
- Nuanced meaning, less data → embeddings or fine-tuned models.
Always compare against the simple baseline.