Ground Truth.
AI, checked against the source.

News · 2026-10-02

llama.cpp merges Qwen multi-token decoding with a reported 55% speed gain

llama.cpp merged multi-token prediction support for Qwen3.8-Flash-Next on October 1, enabling a speculative decoding path for the existing checkpoint. The pull-request author reported about 55% higher decoded throughput in a small test on one hardware and quantization setup, rather than a general performance guarantee.

Key facts

Local AI users often experience a model one generated token at a time. Even when an answer is good, a long wait can make an agent or interactive application feel impractical. This change targets that serving bottleneck by letting the runtime use a prediction component already associated with the model.

Multi-token prediction provides candidates for more than the immediate next token. Speculative decoding then uses a verification process to determine which proposed tokens can be accepted. A useful analogy is drafting several likely next words and checking them together, instead of stopping to compose and approve every single word separately.

The existing lessons on multi-token prediction and speculative decoding cover the broader distinction. Prediction supplies candidate continuations. The runtime’s checking and acceptance rules determine how those proposals become output. A larger number of proposals does not automatically create a speedup: verification cost and how many tokens are accepted also matter.

The merged implementation uses the architecture label Qwen4Exp. That label should not be promoted into the name of a separately launched Qwen4 model. Qwen’s first-party repository identifies Qwen3.8-Flash-Next as an open-weight multimodal mixture-of-experts checkpoint and an early architectural preview. The change is compatibility and execution support for that checkpoint.

The repository reports a large main model plus additional n-gram embedding parameters, with a much smaller subset active per token. Those figures describe architecture. They do not establish the storage size of a chosen quantized download or the memory required to run it. The dossier contains no verified file-size total for the tested IQ4_XS variant, so this article does not invent a disk-space figure.

The primary evidence identifies DGX Spark as the test machine, but does not establish a minimum video-memory requirement for this configuration. Its unified-memory design also should not be casually relabeled as a conventional discrete-card requirement. The required runtime memory is unstated in the reviewed evidence; caches and working state matter alongside loaded weights. Our lessons on quantization and offloading explain why parameter counts alone cannot answer the hardware question.

The reported comparison is concrete. Across 24 prompts in seven categories, throughput rose from 28.36 to 43.88 tokens per second overall. That is 1.55 times the baseline, or about 55% more decoded tokens in the same time. It is not a statement that latency fell by 55%, and it does not measure every part of an application’s waiting time.

The pull request author described the evaluation as “a small portion” of the speed benchmark. That limitation belongs beside the number. Hardware, quantization, prompt mix, generation length, and acceptance behavior can change the result. A model that responds faster after generation begins can still spend substantial time reading a large prompt or loading into memory.

Maintainer Georgi Gerganov merged the patch. That establishes acceptance into the codebase, not independent certification of the benchmark’s representativeness. The discussion records implementation activity; the supplied LocalLLaMA feed confirms that users noticed the change, but it does not provide a reliable survey of real-world performance or broad sentiment.

The strongest counter-argument is therefore about transfer. A favorable small run may not match another user’s machine, model file, or workload. That does not make the merge unimportant. It gives builders a concrete execution path to test, using their own repeated workloads and comparable settings, instead of waiting for a new model release to improve usability.

This story illustrates why runtime work deserves its own place in AI news. Better use of a shipped checkpoint can change the experience without changing its main weights. The verified event is the merged support; the headline speed figure remains the author’s measured result on one setup. Independent reproduction and verified resource requirements are the next pieces needed to turn a promising local improvement into dependable deployment advice.


Primary source, verified: read the paper →

Key questions

Is Qwen4Exp the name of a newly released model?

Qwen4Exp is the architecture label used in this llama.cpp implementation. The actual checkpoint is Qwen3.8-Flash-Next, which Qwen describes as an early preview of architecture intended for Qwen4.

How broad is the reported speed improvement?

The reported 1.55-times throughput comes from 24 prompts on one DGX Spark setup with a particular quantization. It is not an independently reproduced or universal gain.

What did the merged code add?

The pull request adds support for using the checkpoint’s multi-token prediction head for speculative decoding. It improves a runtime path for an existing released checkpoint rather than launching another model.
Cite this

APA

Ground Truth. (2026, October 2). llama.cpp merges Qwen multi-token decoding with a reported 55% speed gain. Ground Truth. https://groundtruth.day/news/llama-cpp-qwen-flash-next-mtp-merge.html

BibTeX

@misc{groundtruth:llama-cpp-qwen-flash-next-mtp-merge,
  title  = {llama.cpp merges Qwen multi-token decoding with a reported 55% speed gain},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {oct},
  url    = {https://groundtruth.day/news/llama-cpp-qwen-flash-next-mtp-merge.html}
}

Topics: open-source · inference · local-ai · speculative-decoding · qwen

Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.