Qatar AI Institute
← Back to articles

25 September 2026 · By

What Is Logistic Regression? Uses in NLP and Computer Vision

What Is Logistic Regression? Uses in NLP and Computer Vision

Despite the name, logistic regression isn't used for regression. It's a classification algorithm: it answers yes/no questions such as "Is this email spam?", "Is this review positive?" or "Does this X-ray show a fracture?", and it gives a probability rather than just a yes or no.

It's one of the oldest tools in machine learning, and still one of the most used. It trains in seconds, it's easy to interpret, and it's often the first model a data scientist tries on a new problem. It also sits inside most deep learning systems: the final layer of a typical neural network classifier is, mathematically, a logistic regression. If you understand it, you understand the last step of much bigger models.

The idea in one picture

Logistic regression works in two steps.

  1. Score. Each input feature x gets a weight w. Multiply each feature by its weight, add them up, and add a constant called the bias b. The result is a score, z = w·x + b, which can be any number: −7, 0.3, 12.
  2. Squash. Pass the score through the sigmoid function, which turns any number into a value between 0 and 1. That value is the model's probability that the answer is "yes".
Graph of the S-shaped sigmoid curve. Large negative scores give probabilities near 0, large positive scores give probabilities near 1, and a score of 0 gives exactly 0.5. Class 0 examples sit along the bottom and class 1 examples along the top, with a few on the wrong side near the middle.
Figure 1. The sigmoid turns a score into a probability. A score of 0 means 50/50.

To make a decision, pick a threshold, usually 0.5: above it, predict "yes"; below it, predict "no". The threshold can be moved. A hospital screening tool might flag anything above 0.2, because missing a real case costs more than a false alarm.

How it learns

At the start, the weights are random and the predictions are useless. Training adjusts them using examples where the right answer is known:

  1. Make a prediction for each training example.
  2. Measure how wrong it was with a loss function called log loss (or cross-entropy). It punishes confident mistakes heavily: saying 99% "spam" for an email that isn't spam costs far more than saying 60%.
  3. Nudge every weight in the direction that reduces the loss (gradient descent), and repeat.

For logistic regression the loss has a single lowest point, with no false minima to get stuck in, so training reliably finds the best weights. In practice a small penalty on large weights (regularisation) is added so the model doesn't over-fit to quirks of the training data.

Reading the weights

A big advantage over deep networks is that you can look inside. Each weight tells you how a feature pushes the prediction:

  • a positive weight pushes towards "yes";
  • a negative weight pushes towards "no";
  • a weight near zero means the feature barely matters.

More precisely, each weight is the change in the log-odds when that feature goes up by one. A weight of 2.1 on the word "great" means each occurrence multiplies the odds of "positive" by e^2.1, about 8. This is why logistic regression is popular in medicine, finance and anywhere else where a decision has to be explained.

More than two classes

For problems with several possible answers, such as which of 10 digits an image shows or which of 20 topics an article covers, the same idea extends to multinomial logistic regression, also called softmax regression. It computes one score per class and uses the softmax function to turn the scores into probabilities that add up to 1. This is exactly what the last layer of an image classifier or a language model does.

Logistic regression in NLP

Computers can't read words, so text first has to become numbers. The classic approach is a bag of words: make a list of every word in the vocabulary, and represent a document by which words it contains (or how often, often weighted by a scheme called TF-IDF). Logistic regression then learns one weight per word.

Worked example: the review 'The food was great but the service was slow' is scored word by word. 'great' has weight +2.1, 'slow' −1.4, 'but' −0.3, 'food' +0.1, 'service' −0.2 and the bias is +0.2. They add up to 0.5, and the sigmoid of 0.5 is 0.62, so the model predicts 'positive' with 62% probability.
Figure 2. Sentiment analysis with logistic regression. Weights are illustrative.

The model has learned from thousands of labelled reviews that "great" is strong evidence for a positive review and "slow" is evidence against. For this mixed review it's only mildly confident (62%), which is a sensible answer.

Common NLP uses include:

  • Spam and phishing filters: fast enough to run on every message.
  • Sentiment analysis of reviews, surveys and social media.
  • Topic and intent classification: routing a support ticket to the right team, or working out what a chatbot user wants.
  • Language identification: deciding whether a text is Arabic, English or French, often using character sequences instead of whole words.

In older NLP papers you'll often see the name maximum entropy classifier ("MaxEnt"). It's the same model.

Today, logistic regression is also used on top of large language models. Instead of word counts, each text is turned into an embedding, a vector produced by a model such as BERT, and logistic regression is trained on those vectors. You get most of the language understanding of the big model with a classifier you can train in seconds on a few hundred examples. (The architecture behind BERT is explained in What Is a Transformer?)

Here's a complete sentiment classifier in Python with scikit-learn:

from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline

texts = ["The food was great", "Terrible and slow service", ...]
labels = [1, 0, ...]  # 1 = positive, 0 = negative

model = make_pipeline(
    TfidfVectorizer(ngram_range=(1, 2)),   # words and word pairs
    LogisticRegression(max_iter=1000),
)
model.fit(texts, labels)

model.predict_proba(["The food was great but the service was slow"])

Logistic regression in computer vision

An image is already numbers: each pixel is a brightness value. The simplest approach is to feed every pixel into logistic regression as a feature. On the classic MNIST dataset of handwritten digits, this alone gets roughly 92% of digits right. Modern convolutional networks score above 99%, but 92% from such a simple model shows how far a straight weighted sum can go.

Two approaches. Top: an 8 by 8 pixel image of the digit zero is multiplied by a learned weight map that is positive in a ring and negative in the centre, summed, and passed through the sigmoid to give a 97% probability that it is a zero. Bottom: a photo passes through a frozen pretrained CNN or vision transformer to produce a feature vector, and only a logistic regression on top of it is trained, giving a probability of a defect.
Figure 3. Two ways logistic regression is used on images. Numbers are illustrative.

The learned weights can be displayed as an image. For "is this a zero?", the weights are positive in a ring (ink there suggests a zero) and negative in the centre (ink there suggests some other digit). The model has learned a template.

Raw pixels break down on real photos, though. A cat is still a cat when it moves a few pixels to the left, but to a pixel-level model that's a completely different input. The modern solution is to let a deep network do the seeing and let logistic regression do the deciding:

  • Linear probes. Take a network already trained on millions of images, such as a CNN or a vision transformer, keep it frozen, and use its internal features as the input to a logistic regression. Researchers use this as a standard test of how good a vision model's features are. OpenAI's CLIP paper, for example, evaluated its image features by training logistic regression on them.
  • Small, specialised tasks. A factory spotting defects on a production line, or a clinic triaging scans, may only have a few hundred labelled images. That's far too few to train a deep network from scratch, but usually enough for a logistic regression on pretrained features.
  • The last layer of every classifier. When a convolutional network labels an image as one of 1,000 categories, its final layer is a softmax, which is multinomial logistic regression applied to features the network learned itself.

Where logistic regression falls short

The score is a weighted sum, so the boundary between "yes" and "no" is always a straight line (a flat plane or hyperplane when there are many features):

Scatter plot of two classes of points separated by a straight diagonal line where the probability is 0.5, with bands of shading showing confidence fading towards the line.
Figure 4. Logistic regression always separates classes with a straight boundary.
  • It can't learn curved boundaries on its own. If the classes are arranged in circles or tangled clusters, it will struggle unless you add features that straighten the problem out.
  • It only adds up evidence. It can't learn on its own that "not" flips the meaning of "good". Adding word pairs such as "not good" as features helps.
  • It depends on good features. This is the real reason deep learning took over: deep networks learn their features. Notice, though, that they still usually end with logistic regression.

Why it's worth learning properly

Logistic regression is the clearest way to understand ideas that run through all of machine learning: features and weights, turning scores into probabilities, loss functions, gradient descent and regularisation. Every one of them reappears in neural networks, just at a larger scale. Learn it well and deep learning becomes a lot less mysterious.

It appears in our list of 7 machine learning algorithms every beginner should know, and you'll build it from scratch in Machine Learning Fundamentals. To see it in action on text and images, see Introduction to Natural Language Processing and Advanced Computer Vision.

Related articles