25 September 2026 · By Hadi Ataei
What Is a Transformer? The Idea Behind Modern AI, Explained with Pictures

Nearly every AI system that's made headlines in the past few years, including ChatGPT, Gemini, Claude, Llama and many image and speech models, is built on the same design: the transformer. It was introduced in 2017 by a team at Google in a paper titled "Attention Is All You Need". This article explains what a transformer does, one picture at a time.
The big picture
A language model built from a transformer does one thing over and over: given some text, it predicts what comes next. Here are the five stages that turn a sentence into a prediction:
- Text. The input, for example "The cat sat on the".
- Tokens. The text is cut into pieces called tokens. Often a token is a whole word, but long or rare words are split into parts.
- Vectors. Each token is turned into a list of numbers (a vector, also called an embedding) that represents its meaning. Words with similar meanings get similar vectors. Information about each token's position in the sentence is added too, because otherwise the model couldn't tell "dog bites man" from "man bites dog".
- Transformer layers. The vectors pass through a stack of identical layers. In each one, the tokens share information with each other (attention) and then each token is processed on its own (feed-forward). Large models stack dozens of these layers.
- Prediction. At the end, the model produces a probability for every token in its vocabulary. "mat" is likely; "roof" is not. This last step is a softmax, the multi-class version of logistic regression.
To write a whole paragraph, the model picks a token, adds it to the input, and runs the whole thing again for the next one. That's all text generation is: predicting the next token, one after another.
The key idea: attention
Stages 1–3 and 5 existed before transformers. What made the transformer different is stage 4, and specifically a mechanism called self-attention.
Look at this sentence: "The animal didn't cross the street because it was too tired." What does "it" mean? You know immediately that it's the animal, because streets don't get tired. But nothing in the word "it" tells you that. You have to look at the other words.
That's what attention does. When the model processes "it", it looks at every other word in the sentence and decides how much each one matters for understanding "it":
After this step, the vector for "it" is no longer just the generic meaning of the word "it". It has absorbed information from "animal" (and a bit from "tired"), so it now carries the meaning "it, meaning the tired animal". Every word in the sentence gets updated this way at the same time.
Older language models (recurrent neural networks) read text one word at a time, left to right, and had to carry everything they'd read in a single running memory. Details from early in a long text tended to fade. Attention lets every word look directly at every other word, however far apart they are. And because all words are processed at once instead of in sequence, transformers train very efficiently on modern GPUs. Those two properties are the main reason they took over.
How attention decides: queries, keys and values
How does the model know "animal" is the word to pay attention to? Each token's vector is used to make three new vectors, each by multiplying with a table of numbers (a weight matrix) that the model learned during training:
- a query (q): what this word is looking for. For "it": "I'm a pronoun; which noun do I refer to?"
- a key (k): what this word offers to others. For "animal": "I'm a noun, a living thing."
- a value (v): the information this word passes on if someone pays attention to it.
The steps, as in Figure 3:
- Score. Compare the query of "it" with the key of every word using a dot product (multiply matching numbers and add them up). A high score means a good match.
- Softmax. Turn the scores into weights between 0 and 1 that add up to 1. The best match gets most of the weight.
- Blend. Multiply each word's value by its weight and add them together. The result is the new vector for "it".
For readers who like formulas, the whole thing is one line: Attention(Q, K, V) = softmax(QKᵀ / √d) · V. The √d just keeps the scores in a reasonable range.
The important point: nobody programmed the rule "pronouns look for nouns". The model learned the query, key and value tables from examples, by adjusting them over billions of training sentences until its next-word predictions got better. The attention patterns are a result of that training.
Inside one transformer block
A full transformer layer, or block, wraps attention with a few more parts:
- Multi-head attention. Instead of one attention step, the block runs several in parallel, called heads, each with its own query/key/value tables. One head might learn to link pronouns to nouns, another to connect verbs with their subjects, another to focus on nearby words. Their results are combined.
- Feed-forward network. After the words have exchanged information, each word's vector goes through a small neural network on its own. Researchers think this is where much of what the model "knows" about the world is stored.
- Shortcuts (residual connections). The input to each part is added back to its output (the + circles in the diagram). This means each part only needs to learn an adjustment, and it keeps information flowing through very deep stacks. Each part also includes a normalisation step (not shown) that keeps the numbers stable.
Stack this block many times and each layer builds on the one below. Early layers tend to capture simple things like grammar and word pairs; later layers capture meaning, facts and reasoning across the whole text.
Encoders, decoders and the models you know
A recent serving paper routes each generated token between a small and a large model. See TokenRouter.
The original 2017 transformer was built for translation and had two halves: an encoder that reads the input sentence, and a decoder that writes the translation. Later models usually keep just one half:
| Type | What it's good at | Examples |
|---|---|---|
| Encoder-only | Understanding text: classification, search, tagging | BERT |
| Decoder-only | Generating text one token at a time | GPT, Claude, Gemini, Llama |
| Encoder–decoder | Turning one text into another: translation, summarisation | The original Transformer, T5 |
We look at the best-known encoder-only model, and how it learned by filling in hidden words, in What Is BERT?
The same idea also works beyond text. Split an image into small patches and treat each patch as a token, and you get a Vision Transformer. Do the same with audio, protein sequences or DNA, and you get models for speech, biology and genomics. Transformers now even control robots, turning camera images and instructions into movements (see Robotics: From Factory Arms to Humanoids).
Why this is worth understanding
Once you understand tokens, embeddings, attention and stacked blocks, most AI news gets easier to follow. "Context window" is how many tokens attention can look at. "Parameters" are mostly the numbers in the attention and feed-forward tables. "Fine-tuning" is continuing to adjust those numbers on new data. It's also the first step from using these models to building them, which is why we think learning how models work is worth the effort.
If you want to go further, our Introduction to Natural Language Processing course covers how language models are built, and the AI & Machine Learning Specialization works up to building neural networks yourself.



