Skip to content

Attention and the Transformer Architecture

The mechanism behind modern language models: self-attention, multi-head attention and positional information, explained intuitively.

Editorial team 2 min read

The transformer, introduced in the 2017 paper Attention Is All You Need, is the architecture behind today's language models and increasingly behind vision and audio models too.

Self-Attention

For each token in a sequence, self-attention decides how much to draw on every other token when building that token's representation. In "The animal didn't cross the road because it was tired", attention helps the model link "it" to "animal".

Mechanically, each token is projected into a query, a key and a value. The similarity between one token's query and every token's key gives attention weights, which are used to combine the values.

Multi-Head Attention

Several attention "heads" run in parallel, each free to focus on different relationships — syntax, reference, position — and their outputs are combined.

Positional Information

Attention itself ignores order, so transformers add positional encodings that tell the model where each token sits in the sequence.

Building Blocks

A transformer layer combines attention with a feed-forward network, residual connections and normalisation. Models stack dozens of such layers.

Encoders and Decoders

  • Encoder-only models (such as BERT) build rich representations for understanding tasks.
  • Decoder-only models (most chat LLMs) generate text one token at a time.
  • Encoder–decoder models (such as T5 and Whisper) map an input sequence to an output sequence.

Why It Won

Transformers process sequences in parallel, scale well with data and compute, and capture long-range relationships — which made large pretrained models possible.

More in Machine learning

All Machine learning guides →
Machine learning Guide · 2 min

Linear Regression Explained

The simplest predictive model: how linear regression fits a line through data, how to read its coefficients, and when it breaks down.

Machine learning 2 min read 17 Sep 2026

Machine learning Guide · 2 min

Logistic Regression for Classification

Despite its name, logistic regression is a classification method. How it produces probabilities and why it remains a strong baseline.

Machine learning 2 min read 16 Sep 2026

Machine learning Guide · 2 min

Decision Trees

How decision trees split data with simple questions, why they are easy to explain, and why single trees overfit.

Machine learning 2 min read 15 Sep 2026

Machine learning Guide · 2 min

Random Forests

Why averaging many randomised decision trees produces a robust, accurate model with little tuning.

Machine learning 2 min read 14 Sep 2026