๐ง Large Language Models ยท Transformers
How the Attention Mechanism Works:
The Idea Behind Every Modern AI Model
The most consequential sentence in modern AI fits inside a paper title. In June 2017, eight Google researchers published "Attention Is All You Need" โ and for a while it read like a pun. Nearly a decade later, every model you'd recognise โ GPT-4o, Gemini, Claude, Llama โ is that paper's architecture scaled up. This guide is the plain-English version: what problem attention solved, the three little vectors that do all the work, why "multi-head" matters, and why the same idea now writes your emails, folds proteins, and paints images.
No prerequisites needed. The actual maths fits on a napkin; the rest is bookkeeping.
๐ก Key insight: Attention lets a network decide, for each word it produces, which words in the input matter most โ the way you re-read the subject of a long sentence to understand a pronoun. Technically: every token computes a weighted sum of every other token's representation, with weights learned from data.
What Was Wrong With the Old Way
Before transformers, sequence models were Recurrent Neural Networks โ machines that read text one token at a time, left to right, carrying a "hidden state" that was supposed to remember everything so far. Two problems. First, memory decayed: by the time an RNN reached the end of a hundred-word sentence, the beginning was a rumour. Second, the left-to-right walk meant training couldn't be parallelized โ a GPU with ten thousand cores was forced to watch one word at a time.
That combination capped how much context models could use and how much data they could train on. Attention removed both caps at once.
Self-Attention: Every Token Looks at Every Token
Self-attention deletes the queue. Every token gets a direct line to every other token, all at once, in parallel. Consider the sentence the original paper's descendants still use: "The animal didn't cross the street because it was too tired." When the model processes the word "it", self-attention scores every other word and correctly loads most of the weight onto "animal" โ not "street" โ because learned patterns say tiredness applies to animals, not streets.
๐ง Watch it happen: visualise the attention matrix for that sentence and you'll see a bright cell at ("it" โ "animal"). The model isn't looking the word up in a grammar table; it's computing that relationship from raw dot products, every time, for every sentence it has never seen.
Queries, Keys, and Values โ the Only Maths You Need
Each token's embedding is multiplied by three learned matrices to produce three vectors with honest names:
Attention scores are then dot products โ a Query dotted against every Key, scaled so gradients behave โ passed through softmax to become weights, and the token's new representation is the weighted sum of all Values:
That's the whole mechanism. A library-search analogy works well: your Query is your question, every book's spine is a Key, and the pages inside are Values. You scan spines (scores), pick what matches (weights), and read the relevant pages (weighted sum of Values).
Multi-Head Attention: 96 Ways to Read a Sentence
One attention head can track one kind of relationship at a time โ coreference, say, or word order. So transformers run many heads in parallel, each with its own Q/K/V matrices, and stitch the results together. One head learns syntax, another tracks pronouns, another watches position, another smells sentiment.
๐ Scale: GPT-3 stacked 96 layers with 96 heads each โ over 9,000 attention heads learning different patterns simultaneously. Modern frontier models go far beyond that, and the 2026 generation adds sparse mixtures of experts on top so only part of the network fires per token.
Why Transformers Beat Everything Else
| Property | RNN / LSTM | Transformer |
|---|---|---|
| Context reach | Decays with distance | Direct, any distance |
| Training parallelism | Sequential, slow | Fully parallel on GPU/TPU |
| Scaling behaviour | Plateaus early | Predictable gains with data & params |
| Modalities | Mostly text/audio | Text, vision, audio, protein, video |
| Long documents | Painful | Million-token contexts in 2026 |
From One Paper to Every Model
The follow-up wave was fast: BERT and GPT in 2018, both transformers, both state-of-the-art on arrival. Vision transformers arrived in 2020 and beat convolutional networks at their own game; the same architecture then went on to generate images from pure noise, predict protein structures, and score speech. Today's frontier models are the 2017 design with three orders of magnitude more compute โ the core attention operation is essentially unchanged.
And the frontier keeps stretching. The million-token contexts you can now paste into a model exist because of attention variants โ KV-caching, flash attention, sliding windows โ that made the quadratic cost of "every token looks at every token" survivable. Our explainer on context windows covers how those tricks turned a memory problem into a product feature.
The same architecture is also walking into the physical world: attention over sensor streams and video is how modern robot policies reason about kitchens and warehouses, which we trace in LLMs in physical space. And the question of keeping all this capability pointed at what we actually want is the subject of our alignment problem guide โ the necessary companion to everything on this page.