Qatar AI Institute
← Back to articles

25 September 2026 · By

What Is a Transformer? The Idea Behind Modern AI, Explained with Pictures

What Is a Transformer? The Idea Behind Modern AI, Explained with Pictures

Nearly every AI system that's made headlines in the past few years, including ChatGPT, Gemini, Claude, Llama and many image and speech models, is built on the same design: the transformer. It was introduced in 2017 by a team at Google in a paper titled "Attention Is All You Need". This article explains what a transformer does, one picture at a time.

The big picture

A language model built from a transformer does one thing over and over: given some text, it predicts what comes next. Here are the five stages that turn a sentence into a prediction:

Diagram: text is split into tokens, tokens become vectors, vectors pass through a stack of attention and feed-forward layers, and the output is a probability for each possible next word, with 'mat' the most likely.
Figure 1. From "The cat sat on the" to a prediction of the next word.
  1. Text. The input, for example "The cat sat on the".
  2. Tokens. The text is cut into pieces called tokens. Often a token is a whole word, but long or rare words are split into parts.
  3. Vectors. Each token is turned into a list of numbers (a vector, also called an embedding) that represents its meaning. Words with similar meanings get similar vectors. Information about each token's position in the sentence is added too, because otherwise the model couldn't tell "dog bites man" from "man bites dog".
  4. Transformer layers. The vectors pass through a stack of identical layers. In each one, the tokens share information with each other (attention) and then each token is processed on its own (feed-forward). Large models stack dozens of these layers.
  5. Prediction. At the end, the model produces a probability for every token in its vocabulary. "mat" is likely; "roof" is not. This last step is a softmax, the multi-class version of logistic regression.

To write a whole paragraph, the model picks a token, adds it to the input, and runs the whole thing again for the next one. That's all text generation is: predicting the next token, one after another.

The key idea: attention

Stages 1–3 and 5 existed before transformers. What made the transformer different is stage 4, and specifically a mechanism called self-attention.

Look at this sentence: "The animal didn't cross the street because it was too tired." What does "it" mean? You know immediately that it's the animal, because streets don't get tired. But nothing in the word "it" tells you that. You have to look at the other words.

That's what attention does. When the model processes "it", it looks at every other word in the sentence and decides how much each one matters for understanding "it":

Diagram: the word 'it' is connected by lines to every other word in the sentence. The line to 'animal' is by far the thickest, the line to 'tired' is medium, and the rest are thin.
Figure 2. When processing "it", attention puts most of its weight on "animal".

After this step, the vector for "it" is no longer just the generic meaning of the word "it". It has absorbed information from "animal" (and a bit from "tired"), so it now carries the meaning "it, meaning the tired animal". Every word in the sentence gets updated this way at the same time.

Older language models (recurrent neural networks) read text one word at a time, left to right, and had to carry everything they'd read in a single running memory. Details from early in a long text tended to fade. Attention lets every word look directly at every other word, however far apart they are. And because all words are processed at once instead of in sequence, transformers train very efficiently on modern GPUs. Those two properties are the main reason they took over.

How attention decides: queries, keys and values

How does the model know "animal" is the word to pay attention to? Each token's vector is used to make three new vectors, each by multiplying with a table of numbers (a weight matrix) that the model learned during training:

  • a query (q): what this word is looking for. For "it": "I'm a pronoun; which noun do I refer to?"
  • a key (k): what this word offers to others. For "animal": "I'm a noun, a living thing."
  • a value (v): the information this word passes on if someone pays attention to it.
Diagram: the query from 'it' is compared with the keys of 'The', 'animal', 'street' and 'tired', giving scores 0.4, 3.1, 1.2 and 2.0. Softmax turns these into weights 0.04, 0.65, 0.10 and 0.21. Each word's value is multiplied by its weight and the results are added up to form the new vector for 'it'.
Figure 3. One attention step for the word "it". Numbers are illustrative.

The steps, as in Figure 3:

  1. Score. Compare the query of "it" with the key of every word using a dot product (multiply matching numbers and add them up). A high score means a good match.
  2. Softmax. Turn the scores into weights between 0 and 1 that add up to 1. The best match gets most of the weight.
  3. Blend. Multiply each word's value by its weight and add them together. The result is the new vector for "it".

For readers who like formulas, the whole thing is one line: Attention(Q, K, V) = softmax(QKᵀ / √d) · V. The √d just keeps the scores in a reasonable range.

The important point: nobody programmed the rule "pronouns look for nouns". The model learned the query, key and value tables from examples, by adjusting them over billions of training sentences until its next-word predictions got better. The attention patterns are a result of that training.

Inside one transformer block

A full transformer layer, or block, wraps attention with a few more parts:

Diagram of one transformer block: word vectors enter at the bottom, pass through multi-head attention, are added to a shortcut copy of the input, then pass through a feed-forward network and are added to another shortcut, producing richer vectors for the next block. A side panel shows three attention heads each tracking a different relationship.
Figure 4. One transformer block. A model stacks many of these.
  • Multi-head attention. Instead of one attention step, the block runs several in parallel, called heads, each with its own query/key/value tables. One head might learn to link pronouns to nouns, another to connect verbs with their subjects, another to focus on nearby words. Their results are combined.
  • Feed-forward network. After the words have exchanged information, each word's vector goes through a small neural network on its own. Researchers think this is where much of what the model "knows" about the world is stored.
  • Shortcuts (residual connections). The input to each part is added back to its output (the + circles in the diagram). This means each part only needs to learn an adjustment, and it keeps information flowing through very deep stacks. Each part also includes a normalisation step (not shown) that keeps the numbers stable.

Stack this block many times and each layer builds on the one below. Early layers tend to capture simple things like grammar and word pairs; later layers capture meaning, facts and reasoning across the whole text.

Encoders, decoders and the models you know

A recent serving paper routes each generated token between a small and a large model. See TokenRouter.

The original 2017 transformer was built for translation and had two halves: an encoder that reads the input sentence, and a decoder that writes the translation. Later models usually keep just one half:

TypeWhat it's good atExamples
Encoder-onlyUnderstanding text: classification, search, taggingBERT
Decoder-onlyGenerating text one token at a timeGPT, Claude, Gemini, Llama
Encoder–decoderTurning one text into another: translation, summarisationThe original Transformer, T5

We look at the best-known encoder-only model, and how it learned by filling in hidden words, in What Is BERT?

The same idea also works beyond text. Split an image into small patches and treat each patch as a token, and you get a Vision Transformer. Do the same with audio, protein sequences or DNA, and you get models for speech, biology and genomics. Transformers now even control robots, turning camera images and instructions into movements (see Robotics: From Factory Arms to Humanoids).

Why this is worth understanding

Once you understand tokens, embeddings, attention and stacked blocks, most AI news gets easier to follow. "Context window" is how many tokens attention can look at. "Parameters" are mostly the numbers in the attention and feed-forward tables. "Fine-tuning" is continuing to adjust those numbers on new data. It's also the first step from using these models to building them, which is why we think learning how models work is worth the effort.

If you want to go further, our Introduction to Natural Language Processing course covers how language models are built, and the AI & Machine Learning Specialization works up to building neural networks yourself.

Related articles