25 September 2026 · By Hadi Ataei
What Is BERT? How Google Taught Computers to Read in Context

In October 2018, researchers at Google published a model called BERT, short for Bidirectional Encoder Representations from Transformers. Within weeks it had set new records on eleven standard language-understanding benchmarks, from sentiment analysis to answering questions about a passage. A year later Google announced it was using BERT in Search, where it was helping with about one in ten English searches in the US.
ChatGPT-style models get the headlines today, but BERT and its descendants still quietly run a huge share of real-world language technology: search engines, spam filters, document classifiers, and the retrieval step behind many AI assistants. This article explains how it works.
The problem BERT solved
Words change meaning with context. "Bank" means something different in "river bank" and "bank loan". Earlier word representations such as word2vec gave each word a single fixed vector, the same in every sentence, so the model had to untangle the meaning some other way.
BERT gives every word a vector that depends on the whole sentence around it. The "bank" in "the bank of the river" and the "bank" in "the bank approved the loan" come out as different vectors. That is what "contextual" means.
Built from half a transformer
BERT is built from the transformer, the architecture behind almost every modern language model. (If you haven't read it yet, our illustrated guide to transformers explains attention, the key idea.) The original transformer has two halves: an encoder that reads text and a decoder that writes it. BERT keeps only the encoder.
That choice matters. A decoder, as in GPT models, reads left to right: each word can only look at the words before it, because when generating text the future hasn't been written yet. An encoder has no such restriction. Every word can look at every other word, before and after it.
That's the "bidirectional" in BERT's name. To understand "bank", BERT can use "river", even though it comes later. This makes BERT very good at understanding text, and not suited to writing it.
How BERT learned: fill in the blank
The clever part is how BERT was trained. Labelled data (text with human-written answers) is expensive and scarce. Unlabelled text is almost unlimited. BERT's creators found a way to learn from plain text alone: hide some words and ask the model to guess them.
This task is called masked language modelling. In each training sentence, 15% of the tokens were chosen for prediction. Most of those were replaced with a special [MASK] token, and a few were swapped for a random word or left unchanged, so the model couldn't simply rely on spotting the mask. To guess the hidden word, the model has to use clues from both sides. Doing this billions of times forces it to learn grammar, word meanings, and a surprising amount about the world.
BERT was also trained on a second task, next sentence prediction: given two sentences, decide whether the second one really followed the first in the original text. Later research (notably RoBERTa, below) found this second task wasn't very useful and dropped it. Masked language modelling was the part that mattered.
The training data was English Wikipedia (about 2.5 billion words) and a collection of books (about 800 million words). No human labelling was needed.
You can try masked prediction yourself in a few lines of Python with the Hugging Face transformers library:
from transformers import pipeline
unmasker = pipeline("fill-mask", model="bert-base-uncased")
unmasker("The [MASK] sat on the mat and purred.")
What BERT sees: tokens, segments and positions
Different models tokenize text differently, which makes distilling one into another harder. See Distillation Across Tokenizers.
Before any of this, text has to be turned into numbers. BERT splits text into WordPiece tokens from a vocabulary of about 30,000. Common words stay whole, and rarer words are broken into familiar pieces, so the model never meets a word it can't represent.
Two special tokens appear in every input. [CLS] sits at the start, and its final vector is used as a summary of the whole input for classification tasks. [SEP] marks the end of each sentence, which lets BERT handle pairs of sentences, such as a question and a passage. Each token's vector is the sum of three learned embeddings: what the token is, which sentence it belongs to, and where it sits. BERT can read up to 512 tokens at a time.
The original release came in two sizes:
| Layers | Vector size | Attention heads | Parameters | |
|---|---|---|---|---|
| BERT-Base | 12 | 768 | 12 | 110 million |
| BERT-Large | 24 | 1,024 | 16 | 340 million |
Those numbers look tiny next to today's large language models, which have hundreds of billions of parameters. That's part of BERT's lasting appeal: BERT-Base runs comfortably on an ordinary server, or even a laptop.
Pre-train once, fine-tune for anything
BERT's biggest influence wasn't the model itself but the way of working it made standard: pre-train a general model once, at great expense, then fine-tune it cheaply for each specific task.
Fine-tuning adds a small output layer on top of BERT and trains the whole thing briefly on a labelled dataset, often a few thousand examples, sometimes fewer:
- Classifying a text (spam, sentiment, topic, intent): feed the
[CLS]vector into the output layer. That layer is a softmax, which is multinomial logistic regression, applied to BERT's features. - Tagging each word (finding names of people, places, organisations and dates, called named-entity recognition): classify every token's vector.
- Answering questions from a passage: predict which token starts the answer and which one ends it.
Before BERT, each of these tasks usually needed its own specially designed model. Afterwards, one pre-trained model and a small head could beat them all. Nearly every language model since, including GPT, has followed this pre-train-then-adapt pattern.
BERT in the real world
When Google added BERT to Search in 2019, it gave an example query: "2019 brazil traveler to usa need a visa". Older systems tended to ignore small words like "to" and returned results about US citizens travelling to Brazil. BERT understood that "to" changes the meaning: this is a Brazilian travelling to the US.
Other common uses:
- Semantic search and retrieval. Sentence-BERT and similar models turn sentences into vectors so that texts with similar meaning are close together, even without shared keywords. This is the "retrieval" step behind many AI assistants that answer questions from a company's own documents.
- Classification at scale: routing support tickets, moderating content, sorting documents.
- Extracting information: pulling names, dates, amounts and drug names out of contracts, invoices or medical notes.
The BERT family, including Arabic
Google released BERT's code and trained models openly, and a large family of variants followed:
- RoBERTa (Facebook AI, 2019): the same design, trained longer on much more data and without next sentence prediction. Better results, and evidence that BERT had been under-trained.
- DistilBERT (Hugging Face, 2019): a compressed version about 40% smaller and 60% faster that keeps about 97% of BERT's language-understanding performance.
- ALBERT (Google, 2019): shares parameters across layers to make large models much smaller.
- Multilingual BERT: one model trained on Wikipedia in 104 languages.
- ModernBERT (2024): an updated encoder that reads up to 8,192 tokens at once and runs faster on modern hardware.
Arabic has its own dedicated BERT-style models, which generally do better on Arabic tasks than multilingual ones. They include AraBERT from the American University of Beirut, CAMeLBERT from NYU Abu Dhabi, MARBERT from the University of British Columbia, and QARiB from the Qatar Computing Research Institute here in Doha, which was trained on Arabic text including dialectal Arabic from social media. For anyone building language technology for this region, these are usually the right starting point.
BERT and GPT: which one when?
Both come from the transformer, but they are built for different jobs:
| BERT (encoder) | GPT-style (decoder) | |
|---|---|---|
| Reads text | In both directions at once | Left to right only |
| Trained to | Fill in hidden words | Predict the next word |
| Best at | Understanding: classify, search, extract | Generating: write, summarise, chat |
| Typical size | Hundreds of millions of parameters | Billions to hundreds of billions |
| Cost to run | Low, even on a CPU | High, usually on GPUs |
If the job is to put text into categories, find similar documents or extract fields, a fine-tuned BERT-style model is often faster, cheaper and more accurate than prompting a giant chatbot. Knowing when to use which model is exactly the kind of judgement that separates people who build AI from people who only use it.
Learn more
Our Introduction to Natural Language Processing course covers how text becomes numbers, and works up to fine-tuning transformer models like BERT on your own data. For the wider picture of how BERT fits among other architectures, see Neural Network Architectures: Transformers, GNNs and GANs.



