Skip to content

Recurrent Neural Networks and LSTMs

How RNNs and LSTMs process sequences step by step, why they struggled with long contexts, and what replaced them.

Editorial team 2 min read

Recurrent neural networks (RNNs) were the standard deep learning approach to sequences — text, speech, time series — before transformers.

How RNNs Work

An RNN reads a sequence one element at a time, maintaining a hidden state that summarises what it has seen so far. At each step it combines the new input with the previous state.

The Vanishing Gradient Problem

When training on long sequences, the signal from early steps fades as it is propagated back through many time steps. Plain RNNs therefore struggle to learn long-range dependencies.

LSTMs and GRUs

Long short-term memory networks add gates that control what information to keep, forget and output, allowing them to remember across longer spans. Gated recurrent units (GRUs) are a simpler variant with similar performance. For years they powered machine translation, speech recognition and text generation.

Limitations

  • Sequential processing is hard to parallelise, so training is slow.
  • Very long dependencies remain difficult.

Why Transformers Took Over

Transformers process all positions in parallel using attention, letting each element directly relate to every other. They train faster on modern hardware and handle long-range context better, so they replaced RNNs for most language tasks.

Where RNNs Still Appear

Small, efficient RNNs remain useful for some streaming and on-device applications, and for certain time-series problems. Recent state-space models revisit the recurrent idea with better long-range performance.

More in Machine learning

All Machine learning guides →
Machine learning Guide · 2 min

Linear Regression Explained

The simplest predictive model: how linear regression fits a line through data, how to read its coefficients, and when it breaks down.

Machine learning 2 min read 17 Sep 2026

Machine learning Guide · 2 min

Logistic Regression for Classification

Despite its name, logistic regression is a classification method. How it produces probabilities and why it remains a strong baseline.

Machine learning 2 min read 16 Sep 2026

Machine learning Guide · 2 min

Decision Trees

How decision trees split data with simple questions, why they are easy to explain, and why single trees overfit.

Machine learning 2 min read 15 Sep 2026

Machine learning Guide · 2 min

Random Forests

Why averaging many randomised decision trees produces a robust, accurate model with little tuning.

Machine learning 2 min read 14 Sep 2026