kimi
The Full 2.8-Trillion-Parameter Kimi K3 Now Runs on Sixteen Desktop Boxes News
An operator has the complete Kimi K3 checkpoint running across sixteen GB10 mini-workstations wired through a single 400G switch, producing roughly 21 to 25 tokens per second for one user, on hardware with a verifiable floor around $57,200.
Kimi K3 runs in 8 gigabytes of RAM, at 33 seconds per token News
A hand-written C engine generates text with Moonshot's 2.8-trillion-parameter Kimi K3 using a peak of 8.24 gigabytes of RAM and no GPU, by reading the model's four-bit experts directly off disk, at a rate of roughly one token every 33 seconds.
Kimi K3 topped a fullstack coding board at maximum effort News
Moonshot's open-weight Kimi K3, served at its highest reasoning setting, took first place on Code Arena's July 23 WebDev snapshot over Claude Fable 5 and GPT-5.6 Sol, though the live board has since moved it to second.
Axios: U.S. Officials Revive an Effort to Discourage Chinese Open-Weight AI News
Axios reports that internal U.S. efforts to restrict Chinese open-weight AI models have revived after Kimi's rise, but no ban, rule, or executive order has been announced.
A Free Model That Splits Your Work Across 300 Helpers News
Moonshot AI's Kimi K2.6 is a frontier-grade model anyone can download, and its headline trick is fanning a single job out to hundreds of helpers working in parallel.
Unsloth Kimi-K3-GGUF Tool
Converted local-inference builds of Moonshot's Kimi K3: a 1.51 TB four-bit UD-Q4_K_XL file, a 1.56 TB eight-bit build, and BF16/F16/F32 multimodal projector files that preserve an image-input path. Datacenter-scale hardware still required.
Kimi-K3-DSpark Tool
Inferact's draft model for Kimi K3. It proposes seven tokens at a time for K3 to verify and accept or discard, and its block-diffusion backbone shares K3's attention-cache layout so no second cache format is needed. This is the component behind the 21-25 tokens per second measured on a sixteen-node GB10 cluster running the full K3 checkpoint.