News · 2026-09-29
Jeff releases a small open model for fast typed decisions on local hardware
Jeff released an open, Jev-compatible family of small decision models that choose among typed options in a single forward pass instead of writing a response. The 0.8B version is a roughly 1.7 GB BF16 download, making the story a concrete example of local AI becoming useful for narrow, high-volume choices without pretending to replace a frontier reasoner.
Key facts
- The Jeff repository publishes serving code, training material, and the typed decision interface.
- The 0.8B model card lists about 1.7 GB of BF16 weights.
- The author reports 22 ms median latency on an RTX PRO 6000 and 28 ms on an M4 Max via MLX for a roughly 200-token decision.
- The author says the 0.8B training run took about two hours on one RTX PRO 6000; no cloud GPU use is reported.
A chat model is usually asked to write an answer, then an application parses the answer and decides what it means. Jeff flips that sequence. An application supplies the state plus named choices and descriptions; the model returns a probability per option, a selected option, and a confidence. The project supports choice, yes/no probability, and ordered-score requests. Its author describes the intended pattern concisely: “reason in code, decide with Jeff.”
Think of it as a compact dispatcher rather than a general consultant. A dispatcher can route a support ticket, choose a moderation action, select an intent, or decide whether a gate should open. It does not need to write a memoir about the decision. Avoiding free-form generation removes parsing and can make the loop fast enough for local use. The repository says Jeff is trained with supervised learning over option letters and then temperature-calibrated, rather than using a special proprietary architecture.
The project is unusually transparent about its limits. The published benchmark panel includes 4,599 questions, but its author says comparisons with Jev use different samples. On the harder 105-item JevBench tier, Jeff-0.8B reports 47.6% versus Jev’s published 73.3%; the 2B version reaches 53.3%. The author’s own message is not “a hobbyist cloned a frontier product.” It is that a small, locally controllable System-1 layer can cover a useful slice of decision work.
That distinction matters for hardware. The 1.7 GB download size is disk storage, not a published VRAM requirement. The official source gives measured configurations—RTX PRO 6000 and M4 Max—but does not state a minimum GPU-memory requirement, so one should not infer it from parameter count. The weights fit storage easily; actual runtime memory still includes framework overhead, context, and activations.
The strongest counterargument is practical and fair: many narrow classification problems are better solved with an embedding model plus linear head, an ordinary classifier, rules, or one-token logit scoring. Those systems can be smaller, easier to calibrate, and less sensitive to option wording. Jeff’s own results also show that a better static benchmark rank does not necessarily produce better behavior in its game harness.
That is why calibration—not merely raw accuracy—is central. A confidence score should not be mistaken for a guarantee. A team needs to test on its own options, order those options differently, establish abstention thresholds, and measure the cost of a false action. Jeff’s release is valuable because it makes this tradeoff inspectable: local weights and a clean interface lower the barrier to owning the decision loop, while shifting responsibility for its failure modes to the operator.
Key questions
What does Jeff do differently from a chat model?
How large is the downloadable Jeff 0.8B model?
Can Jeff replace a frontier reasoning model?
Cite this
APA
Ground Truth. (2026, September 29). Jeff releases a small open model for fast typed decisions on local hardware. Ground Truth. https://groundtruth.day/news/jeff-open-0-point-8b-local-decision-model.html
BibTeX
@misc{groundtruth:jeff-open-0-point-8b-local-decision-model,
title = {Jeff releases a small open model for fast typed decisions on local hardware},
author = {{Ground Truth}},
year = {2026},
month = {sep},
url = {https://groundtruth.day/news/jeff-open-0-point-8b-local-decision-model.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.