News · 2026-10-02
llama.cpp merges Qwen multi-token decoding with a reported 55% speed gain
llama.cpp merged multi-token prediction support for Qwen3.8-Flash-Next on October 1, enabling a speculative decoding path for the existing checkpoint. The pull-request author reported about 55% higher decoded throughput in a small test on one hardware and quantization setup, rather than a general performance guarantee.
Key facts
- Pull request 29761 was merged October 1, 2026.
- The author’s 24-prompt test reported 43.88 tokens per second versus 28.36 without the new path.
- The comparison used a DGX Spark and an IQ4_XS checkpoint quantization.
- The primary source is the merged llama.cpp pull request.
Local AI users often experience a model one generated token at a time. Even when an answer is good, a long wait can make an agent or interactive application feel impractical. This change targets that serving bottleneck by letting the runtime use a prediction component already associated with the model.
Multi-token prediction provides candidates for more than the immediate next token. Speculative decoding then uses a verification process to determine which proposed tokens can be accepted. A useful analogy is drafting several likely next words and checking them together, instead of stopping to compose and approve every single word separately.
The existing lessons on multi-token prediction and speculative decoding cover the broader distinction. Prediction supplies candidate continuations. The runtime’s checking and acceptance rules determine how those proposals become output. A larger number of proposals does not automatically create a speedup: verification cost and how many tokens are accepted also matter.
The merged implementation uses the architecture label Qwen4Exp. That label should not be promoted into the name of a separately launched Qwen4 model. Qwen’s first-party repository identifies Qwen3.8-Flash-Next as an open-weight multimodal mixture-of-experts checkpoint and an early architectural preview. The change is compatibility and execution support for that checkpoint.
The repository reports a large main model plus additional n-gram embedding parameters, with a much smaller subset active per token. Those figures describe architecture. They do not establish the storage size of a chosen quantized download or the memory required to run it. The dossier contains no verified file-size total for the tested IQ4_XS variant, so this article does not invent a disk-space figure.
The primary evidence identifies DGX Spark as the test machine, but does not establish a minimum video-memory requirement for this configuration. Its unified-memory design also should not be casually relabeled as a conventional discrete-card requirement. The required runtime memory is unstated in the reviewed evidence; caches and working state matter alongside loaded weights. Our lessons on quantization and offloading explain why parameter counts alone cannot answer the hardware question.
The reported comparison is concrete. Across 24 prompts in seven categories, throughput rose from 28.36 to 43.88 tokens per second overall. That is 1.55 times the baseline, or about 55% more decoded tokens in the same time. It is not a statement that latency fell by 55%, and it does not measure every part of an application’s waiting time.
The pull request author described the evaluation as “a small portion” of the speed benchmark. That limitation belongs beside the number. Hardware, quantization, prompt mix, generation length, and acceptance behavior can change the result. A model that responds faster after generation begins can still spend substantial time reading a large prompt or loading into memory.
Maintainer Georgi Gerganov merged the patch. That establishes acceptance into the codebase, not independent certification of the benchmark’s representativeness. The discussion records implementation activity; the supplied LocalLLaMA feed confirms that users noticed the change, but it does not provide a reliable survey of real-world performance or broad sentiment.
The strongest counter-argument is therefore about transfer. A favorable small run may not match another user’s machine, model file, or workload. That does not make the merge unimportant. It gives builders a concrete execution path to test, using their own repeated workloads and comparable settings, instead of waiting for a new model release to improve usability.
This story illustrates why runtime work deserves its own place in AI news. Better use of a shipped checkpoint can change the experience without changing its main weights. The verified event is the merged support; the headline speed figure remains the author’s measured result on one setup. Independent reproduction and verified resource requirements are the next pieces needed to turn a promising local improvement into dependable deployment advice.
Key questions
Is Qwen4Exp the name of a newly released model?
How broad is the reported speed improvement?
What did the merged code add?
Cite this
APA
Ground Truth. (2026, October 2). llama.cpp merges Qwen multi-token decoding with a reported 55% speed gain. Ground Truth. https://groundtruth.day/news/llama-cpp-qwen-flash-next-mtp-merge.html
BibTeX
@misc{groundtruth:llama-cpp-qwen-flash-next-mtp-merge,
title = {llama.cpp merges Qwen multi-token decoding with a reported 55% speed gain},
author = {{Ground Truth}},
year = {2026},
month = {oct},
url = {https://groundtruth.day/news/llama-cpp-qwen-flash-next-mtp-merge.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.