Support vector machines (SVMs) were among the most popular classifiers before deep learning, and remain useful for small and medium-sized datasets.
The Maximum-Margin Idea
For two classes, many lines (or planes) might separate them. An SVM chooses the one with the widest margin — the greatest distance to the nearest points of each class. Those nearest points are the support vectors; they alone define the boundary.
Soft Margins
Real data overlaps, so SVMs allow some points to fall inside the margin or on the wrong side. The hyperparameter C controls the trade-off: large C penalises errors heavily (risking overfitting), small C allows a wider, more tolerant margin.
Kernels
A kernel lets an SVM draw curved boundaries by implicitly mapping data into a higher-dimensional space where a straight separator exists. Common kernels are linear, polynomial and the radial basis function (RBF), which has its own width parameter, gamma.
Strengths
- Effective in high-dimensional spaces, such as text features.
- Robust with clear margins between classes.
Weaknesses
- Training scales poorly to very large datasets.
- Requires feature scaling and careful tuning of C and gamma.
- Doesn't output probabilities directly (they can be estimated with extra calibration).
When to Use It
Consider an SVM for small-to-medium datasets with many features, such as text classification with TF-IDF features. For large tabular datasets, gradient boosting is usually a better default.