Image classification assigns a label to a whole image — "cat", "invoice", "defective part". Pretrained models make it accessible without huge datasets.
Start With a Pretrained Model
Models such as ResNet, EfficientNet and vision transformers were trained on large datasets like ImageNet (1,000 everyday categories). Use one directly if its categories fit, or adapt it to yours.
Prepare Images Exactly as Expected
Each model expects a specific input: size (often 224×224), colour order (RGB), value range and normalisation. Mismatched preprocessing doesn't cause errors — just quietly worse predictions. Resize the shorter side and centre-crop to avoid distorting images.
Reading the Output
The model outputs a score per class; softmax turns scores into probabilities. Look at the top five predictions, because many categories are similar.
Adapting to Your Categories
- Feature extraction: use the model's internal representation as input to a simple classifier trained on your images. Works with a few hundred examples.
- Fine-tuning: retrain the final layers (or all layers, gently) on your labelled images.
Collecting Good Training Images
Use images like those the system will see in production — same cameras, lighting and backgrounds — and cover the variation that occurs.
Watch for Shortcuts
Models learn whatever separates classes in training data, which may be a watermark, background or ruler rather than the object. Inspect errors and test on images from different sources.