Gradient descent is how most models — from logistic regression to large neural networks — find good parameter values.
The Intuition
Imagine standing on a hilly landscape in fog, trying to reach the lowest valley. You feel which way the ground slopes and take a step downhill. Repeat, and you gradually descend. The landscape is the loss function; your position is the model's parameters; the slope is the gradient.
The Algorithm
- Start with initial parameters.
- Compute the gradient of the loss with respect to each parameter.
- Move each parameter a small step in the opposite direction of its gradient.
- Repeat until the loss stops improving.
The Learning Rate
The step size is the learning rate, the single most important training setting.
- Too large: training oscillates or diverges; the loss jumps around or explodes.
- Too small: training is painfully slow and may stall.
Variants
- Batch gradient descent uses all data per step — accurate but slow.
- Stochastic gradient descent (SGD) uses one example per step — noisy but fast.
- Mini-batch uses small batches — the standard compromise.
- Momentum, RMSProp and Adam adapt steps to speed up and stabilise training; Adam is a common default for neural networks.
Learning-Rate Schedules
Reducing the learning rate during training — step decay, cosine decay, or warm-up followed by decay — often improves final results.
Practical Tip
Plot the training and validation loss over time. A smooth decline is healthy; spikes or a flat line suggest adjusting the learning rate.