Ground Truth.
AI, checked against the source.

News · 2026-10-05

Strata users report roughly 100 tokens a second on a gaming GPU—with substantial RAM

Strata users report roughly 100 generated tokens per second from Qwen3.8-Flash-Next on an RTX 4090, using compressed weights and substantial system RAM. The result makes a large sparse model more practical on consumer hardware, but it does not put the complete model inside the card’s 24 GB of GPU memory or establish comparable prompt-processing speed.

Key facts

The eye-catching part is the mismatch between the model’s headline size and the graphics card. The useful explanation is the memory system surrounding that card. Strata is an inference engine maintained by Niko1221, publicly displayed as “Niko,” rather than a new model or compression format. Its server and interface expose local application connections, including compatibility with familiar model-service interfaces. The release history makes this a shipping software story, although installation alone does not verify its performance on another computer.

Qwen’s own accounting matters. The model card lists 125 billion main-model parameters, six billion activated for each token, a separate 51-billion-parameter n-gram embedding table and a four-billion-parameter prediction head. Qwen labels that head “MTP,” for multi-token prediction: a mechanism related to proposing several upcoming tokens rather than repeatedly doing all the work for just one. The frequently repeated 125-billion figure describes the backbone. A 176-billion total omits the separately listed prediction head.

The model is a mixture of experts. Imagine a large workshop where only a few specialists work on each job. You must retain access to all the specialists, but you do not need every specialist at the same workbench for every operation. Strata keeps frequently used experts close to the GPU, holds others in system RAM and overlaps some missed-expert computation with GPU work. A large lookup table can remain on storage. Its technical explanation describes a coordinated arrangement of memory, computation and transfers.

That distinction changes the hardware claim. One detailed report uses a 24 GB RTX 4090 and 64 GB of system RAM; a separate Hacker News commenter reports 124 tokens per second with 128 GB of RAM but gives less information about the test. These are measured configurations reported by users, not a published universal minimum. Strata says it loads roughly 35–55 GB into RAM at startup and recommends at least 64 GB for its larger model sizes. GPU-memory needs depend on configuration; the dossier establishes successful claimed runs on a 24 GB card, not a complete minimum requirement for every setting.

The quantization card provides a storage figure readers can act on: IQ3_XXS totals 75.8 GB, with 47.0 GB of transformer weights and a 28.8 GB n-gram shard. Disk size is separate from both resident RAM and GPU memory. Mapping a shard from disk means pages can be fetched as required; it does not make the files disappear or guarantee that a slow drive will perform well. This is the practical territory covered by offloading and streaming weights.

The speed evidence also has two distinct phases. Issue #307 reports fast output generation but much lower prompt-processing rates in that Windows setup. The deeper dossier also records issue #308, with whole-request and peak-generation figures that should not be blended into a single benchmark. A configured 128K context is capacity, not evidence that the measured prompt contained 128,000 tokens. Readers comparing machines should keep prefill and decode separate and compare the same prompt, output length, compression and settings.

Quality is the unresolved half. The quantization authors report near-baseline averages on selected tasks, while community posts describe coding weaknesses, tool-call failures and unrelated text in reasoning. Those complaints are evidence of reported failures, not a diagnosis of compression or engine defects. A separately documented tensor-reading bug is not proof of the complaint’s cause.

The strongest counterargument is therefore practical: fast text is useful only if representative tasks still succeed. Strata has credible first-person reproductions and a plausible engineering mechanism, but no controlled independent suite in the dossier establishes dependable coding, long-prompt latency or concurrent service. The advance is fast local generation from a larger whole-system memory arrangement; the next test is useful work per minute under disclosed conditions.


Primary source, verified: read the paper →

Key questions

Does the entire Qwen model fit in the RTX 4090’s 24 GB?

No; Strata distributes the workload across GPU memory, system RAM and storage rather than fitting the complete model on the card.

How large is the IQ3_XXS download used in the reported run?

The quantization card lists 75.8 GB, comprising 47.0 GB of transformer weights and a 28.8 GB n-gram shard.

Does the 100-token-per-second figure measure prompt processing?

No; it is a reported generation rate, and neither a configured 128K context nor that rate establishes performance on a 128K-token prompt.
Cite this

APA

Ground Truth. (2026, October 5). Strata users report roughly 100 tokens a second on a gaming GPU—with substantial RAM. Ground Truth. https://groundtruth.day/news/strata-reports-fast-qwen-generation-on-a-gaming-gpu.html

BibTeX

@misc{groundtruth:strata-reports-fast-qwen-generation-on-a-gaming-gpu,
  title  = {Strata users report roughly 100 tokens a second on a gaming GPU—with substantial RAM},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {oct},
  url    = {https://groundtruth.day/news/strata-reports-fast-qwen-generation-on-a-gaming-gpu.html}
}

Topics: local-ai · inference · open-weights · quantization

Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.