consumer-gpu
Everything on Ground Truth tagged “consumer-gpu” — 1 item.
NInfer Tool
A focused inference engine that runs Qwen 3.6 models on one RTX 5090 with a 262,000-token context using an INT8 key-value cache, reporting roughly 188 tokens per second at 250,000 tokens of context. Methodology, seeds and limits are published openly.