Learn · Beginner
Logistic regression: turning evidence into a probability for a yes-or-no decision
Logistic regression is a model that learns how input features change the probability of a yes-or-no outcome. It combines the features into a weighted score and passes that score through a curve that stays between zero and one. It matters because many useful AI tasks end in a decision, and a probability lets you separate what the model predicts from what the system should do.
Suppose a service wants to identify messages that need urgent human review. Its inputs might include whether a customer reports lost access, how long an issue has remained unresolved, and whether a payment failed. Training examples pair those features with labels indicating whether review was actually necessary. Logistic regression learns which features increase or decrease the score.
The computation starts with a weighted sum: multiply each feature by its learned weight, add the results, and add an intercept. A positive weight pushes the prediction toward the positive outcome; a negative weight pushes it away. Feature scale matters: a weight on hours cannot be directly compared with a weight on a yes-or-no flag without considering the units.
The model then applies the logistic, or sigmoid, function: p = 1 / (1 + exp(-z)), where z is the score. A score of zero becomes probability 0.5. A large positive score approaches one, and a large negative score approaches zero. The curve compresses an unrestricted score into the range required for a probability.
An analogy is a dimmer attached to several weighted switches. Each switch changes the underlying setting, but the displayed brightness remains between fully off and fully on. Unlike a hard rule saying that any failed payment triggers review, the model can combine several modest clues. That makes it useful when no single feature reliably decides the outcome.
The historical foundation includes David Cox’s 1958 work on binary sequences. The important mathematical interpretation is that the weighted sum models log-odds. Odds are p divided by one minus p. Taking their logarithm makes a multiplicative change in odds into an additive change in score. A one-unit feature increase shifts log-odds by its weight, holding the other inputs fixed.
That interpretation describes association under the fitted model, not causation. If urgency correlates with account age, a learned coefficient does not establish that making an account older causes urgency. Missing variables, biased labels, and selection effects can change the apparent relationship. The model predicts from the examples it receives.
Training usually minimizes binary cross-entropy, which penalizes assigning low probability to the observed outcome. Predicting 0.01 for an event that happens is a bigger mistake than predicting 0.4. Gradient descent adjusts the weights to reduce that loss; Léon Bottou’s work on stochastic gradient descent explains how updates from small samples make large-scale learning practical. Regularization can discourage unnecessarily large weights and reduce overfitting.
The probability is only the first output. The action comes from a threshold. Sending every message above 0.5 to a reviewer may be reasonable when errors have similar costs. If missing a genuine emergency is much worse than reviewing a routine message, a lower threshold may be appropriate. If the review queue is overwhelmed, capacity also matters. These are operational choices, not properties automatically determined by training.
Evaluate both discrimination and calibration. Discrimination asks whether positive cases tend to receive higher scores than negative ones. Calibration asks whether cases scored at about 0.8 are positive about 80 percent of the time. A mathematically valid probability can still be badly calibrated, especially after the deployment population changes. Use an honest holdout set, and inspect important subgroups rather than just a single average.
Logistic regression draws a linear boundary in its input features. It can represent more complex relationships if you supply interactions or transformed features, but it does not discover arbitrary structure by itself. Its appeal is a small, fast, inspectable baseline. When a workflow has reliable structured features and enough representative labels, that baseline can be difficult to beat economically.
Today’s decision-model releases tackle similar bounded outputs using language-model backbones, often with natural language or images as inputs. Ground Truth’s Liquid d1 story shows the distinction in practice. A richer backbone can handle inputs that are hard to turn into hand-built features, but it does not remove the need for calibration, thresholds, or comparison with simpler models.
The useful habit is to keep three questions separate: what probability does the model assign, how trustworthy is that probability on current data, and what action follows given the cost of a mistake? Logistic regression makes that separation visible, which is why it remains an important foundation for understanding automated decisions.
David Cox, The Regression Analysis of Binary Sequences (1958)
Léon Bottou, Large-Scale Machine Learning with Stochastic Gradient Descent (2010)
Key questions
Why is logistic regression used for classification despite its name?
Does a logistic-regression probability of 0.8 mean the model is correct 80 percent of the time?
When should a workflow choose a threshold other than 0.5?
Cite this
APA
Ground Truth. (2026, October 10). Logistic regression: turning evidence into a probability for a yes-or-no decision. Ground Truth. https://groundtruth.day/learn/logistic-regression.html
BibTeX
@misc{groundtruth:logistic-regression,
title = {Logistic regression: turning evidence into a probability for a yes-or-no decision},
author = {{Ground Truth}},
year = {2026},
month = {oct},
url = {https://groundtruth.day/learn/logistic-regression.html}
}