Skip to content

Support Vector Machines

How SVMs find the widest margin between classes, and how kernels let them draw curved boundaries.

Editorial team 2 min read

Support vector machines (SVMs) were among the most popular classifiers before deep learning, and remain useful for small and medium-sized datasets.

The Maximum-Margin Idea

For two classes, many lines (or planes) might separate them. An SVM chooses the one with the widest margin — the greatest distance to the nearest points of each class. Those nearest points are the support vectors; they alone define the boundary.

Soft Margins

Real data overlaps, so SVMs allow some points to fall inside the margin or on the wrong side. The hyperparameter C controls the trade-off: large C penalises errors heavily (risking overfitting), small C allows a wider, more tolerant margin.

Kernels

A kernel lets an SVM draw curved boundaries by implicitly mapping data into a higher-dimensional space where a straight separator exists. Common kernels are linear, polynomial and the radial basis function (RBF), which has its own width parameter, gamma.

Strengths

  • Effective in high-dimensional spaces, such as text features.
  • Robust with clear margins between classes.

Weaknesses

  • Training scales poorly to very large datasets.
  • Requires feature scaling and careful tuning of C and gamma.
  • Doesn't output probabilities directly (they can be estimated with extra calibration).

When to Use It

Consider an SVM for small-to-medium datasets with many features, such as text classification with TF-IDF features. For large tabular datasets, gradient boosting is usually a better default.

More in Machine learning

All Machine learning guides →
Machine learning Guide · 2 min

Linear Regression Explained

The simplest predictive model: how linear regression fits a line through data, how to read its coefficients, and when it breaks down.

Machine learning 2 min read 17 Sep 2026

Machine learning Guide · 2 min

Logistic Regression for Classification

Despite its name, logistic regression is a classification method. How it produces probabilities and why it remains a strong baseline.

Machine learning 2 min read 16 Sep 2026

Machine learning Guide · 2 min

Decision Trees

How decision trees split data with simple questions, why they are easy to explain, and why single trees overfit.

Machine learning 2 min read 15 Sep 2026

Machine learning Guide · 2 min

Random Forests

Why averaging many randomised decision trees produces a robust, accurate model with little tuning.

Machine learning 2 min read 14 Sep 2026