News · 2026-10-10
Microsoft launches Decision-1 to score workflow choices without long generated answers
Microsoft introduced Microsoft-Decision-1 on October 9 as a hosted model that scores bounded choices for routing, classification, prioritization, verification, and other workflow tasks. The company reports evaluating it across 36 benchmarks and nearly 150,000 questions held blind from training. Its purpose is to return actionable probabilities in a single pass, rather than make software extract a decision from a long generated answer.
Key facts
- Microsoft reports a comparison covering 36 benchmarks and nearly 150,000 questions.
- The announcement is dated October 9, 2026.
- Decision-1 is available through Microsoft Foundry and OpenRouter; no downloadable checkpoint is established.
- The primary source is Microsoft’s Decision-1 announcement.
Many business tasks require a choice rather than a composition. An incoming message needs a queue. A record needs a risk band. A draft needs a rubric score. A proposed action needs a pass or fail. A conversational model can do these jobs, but producing prose around the answer adds delay and can make integration awkward. The workflow must parse the response and decide what to do when the explanation and selected label disagree.
Microsoft’s announcement positions Decision-1 as a “decision model.” It says the company post-trained a Qwen3.5-9B language-model backbone for single-pass scoring and that the system returns calibrated probabilities for yes/no, multiple-choice, and rating questions. The important specialization is the task and output contract. The source does not establish that Microsoft invented a new underlying architecture.
The analogy is an intake clerk using a marked form rather than writing an essay about every package. The form names the available destinations and records a confidence for each one. Software can use those fields immediately: send high-confidence routine cases onward, route uncertain cases to a person, or choose a threshold appropriate to the consequences of a mistake. The output remains a prediction, but it fits an operational decision boundary.
Calibration is critical to that boundary. A system that repeatedly assigns 90 percent confidence should be right about nine times out of ten in comparable cases. That is different from merely choosing the most likely label. A model can rank choices well while giving unreliable confidence values, making its scores risky inputs to automatic escalation or approval rules.
Microsoft’s evaluation is still Microsoft’s evaluation. The claim that the comparison data stayed blind from training is meaningful, but the dossier does not establish independent reproduction. Readers need the task mix, baselines, scoring rules, and operational costs before treating a headline comparison as evidence that one service is best for every decision workflow.
A separate local option makes the broader category tangible. H2O-Lightning-4B is built on Qwen3.5-4B and returns a choice, yes/no probability, or ordinal score with a probability distribution. H2O says a decision uses one forward pass and one output token, and that the model also supports images. Its weights are Apache-2.0 licensed and constitute a verified 9.1 GB weight download.
H2O’s card reports local measurements on an RTX PRO 4500 Blackwell with 32 GB of graphics memory. That is a measured configuration, not a published minimum requirement or validation of a 16 GB consumer card. Decision-1’s announcement describes hosted access, so it does not specify a local graphics-memory requirement for customers.
H2O says its version 1.1 checkpoint ranked first among open-weight systems on the October 7 JevBench version 1.6.1 composite. That composite is an equal-weight harmonic mean of intelligence, calibration, speed, and cost. Larger models can lead on raw intelligence while losing the composite. H2O’s result and Microsoft’s benchmark comparison are not entries on one shared leaderboard.
The JevBench repository helps clarify what typed-decision evaluation can include: probability vectors scored against unseen answers, calibration, threshold effects, option-order sensitivity, and paired comparisons. Ground Truth’s earlier Liquid d1 story shows this specialization already emerging across other providers.
The strongest case for Decision-1 is a cleaner contract between prediction and software action. The strongest counterargument is that familiar classification problems do not become reliable merely because a language model has been packaged for them. Developers still need comparison with simpler baselines, representative evaluation, suitable thresholds, and a route for uncertain cases.
The shipping story is a new hosted option for bounded decisions. The unresolved question is whether its probability quality and total operating cost justify replacing the current decision system on a particular task. A well-formed score is useful; it is not automatic evidence that the resulting action is correct.
Key questions
Can Microsoft-Decision-1 be downloaded and run locally?
How is a decision model different from a conversational answer?
Does H2O’s local model use the same evaluation as Decision-1?
Cite this
APA
Ground Truth. (2026, October 10). Microsoft launches Decision-1 to score workflow choices without long generated answers. Ground Truth. https://groundtruth.day/news/microsoft-decision-1-single-pass-workflow-scoring.html
BibTeX
@misc{groundtruth:microsoft-decision-1-single-pass-workflow-scoring,
title = {Microsoft launches Decision-1 to score workflow choices without long generated answers},
author = {{Ground Truth}},
year = {2026},
month = {oct},
url = {https://groundtruth.day/news/microsoft-decision-1-single-pass-workflow-scoring.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.