Skip to content

Gradient Descent and Learning Rates

The optimisation method behind most machine learning: how gradient descent works, and how to choose a learning rate.

Editorial team 2 min read

Gradient descent is how most models — from logistic regression to large neural networks — find good parameter values.

The Intuition

Imagine standing on a hilly landscape in fog, trying to reach the lowest valley. You feel which way the ground slopes and take a step downhill. Repeat, and you gradually descend. The landscape is the loss function; your position is the model's parameters; the slope is the gradient.

The Algorithm

  1. Start with initial parameters.
  2. Compute the gradient of the loss with respect to each parameter.
  3. Move each parameter a small step in the opposite direction of its gradient.
  4. Repeat until the loss stops improving.

The Learning Rate

The step size is the learning rate, the single most important training setting.

  • Too large: training oscillates or diverges; the loss jumps around or explodes.
  • Too small: training is painfully slow and may stall.

Variants

  • Batch gradient descent uses all data per step — accurate but slow.
  • Stochastic gradient descent (SGD) uses one example per step — noisy but fast.
  • Mini-batch uses small batches — the standard compromise.
  • Momentum, RMSProp and Adam adapt steps to speed up and stabilise training; Adam is a common default for neural networks.

Learning-Rate Schedules

Reducing the learning rate during training — step decay, cosine decay, or warm-up followed by decay — often improves final results.

Practical Tip

Plot the training and validation loss over time. A smooth decline is healthy; spikes or a flat line suggest adjusting the learning rate.

More in Machine learning

All Machine learning guides →
Machine learning Guide · 2 min

Linear Regression Explained

The simplest predictive model: how linear regression fits a line through data, how to read its coefficients, and when it breaks down.

Machine learning 2 min read 17 Sep 2026

Machine learning Guide · 2 min

Logistic Regression for Classification

Despite its name, logistic regression is a classification method. How it produces probabilities and why it remains a strong baseline.

Machine learning 2 min read 16 Sep 2026

Machine learning Guide · 2 min

Decision Trees

How decision trees split data with simple questions, why they are easy to explain, and why single trees overfit.

Machine learning 2 min read 15 Sep 2026

Machine learning Guide · 2 min

Random Forests

Why averaging many randomised decision trees produces a robust, accurate model with little tuning.

Machine learning 2 min read 14 Sep 2026