Ground Truth.
AI, checked against the source.

← All topics

kimi

Everything on Ground Truth tagged “kimi” — 7 items.

The Full 2.8-Trillion-Parameter Kimi K3 Now Runs on Sixteen Desktop Boxes News

An operator has the complete Kimi K3 checkpoint running across sixteen GB10 mini-workstations wired through a single 400G switch, producing roughly 21 to 25 tokens per second for one user, on hardware with a verifiable floor around $57,200.

Kimi K3 runs in 8 gigabytes of RAM, at 33 seconds per token News

A hand-written C engine generates text with Moonshot's 2.8-trillion-parameter Kimi K3 using a peak of 8.24 gigabytes of RAM and no GPU, by reading the model's four-bit experts directly off disk, at a rate of roughly one token every 33 seconds.

Kimi K3 topped a fullstack coding board at maximum effort News

Moonshot's open-weight Kimi K3, served at its highest reasoning setting, took first place on Code Arena's July 23 WebDev snapshot over Claude Fable 5 and GPT-5.6 Sol, though the live board has since moved it to second.

Axios: U.S. Officials Revive an Effort to Discourage Chinese Open-Weight AI News

Axios reports that internal U.S. efforts to restrict Chinese open-weight AI models have revived after Kimi's rise, but no ban, rule, or executive order has been announced.

A Free Model That Splits Your Work Across 300 Helpers News

Moonshot AI's Kimi K2.6 is a frontier-grade model anyone can download, and its headline trick is fanning a single job out to hundreds of helpers working in parallel.

Unsloth Kimi-K3-GGUF Tool

Converted local-inference builds of Moonshot's Kimi K3: a 1.51 TB four-bit UD-Q4_K_XL file, a 1.56 TB eight-bit build, and BF16/F16/F32 multimodal projector files that preserve an image-input path. Datacenter-scale hardware still required.

Kimi-K3-DSpark Tool

Inferact's draft model for Kimi K3. It proposes seven tokens at a time for K3 to verify and accept or discard, and its block-diffusion backbone shares K3's attention-cache layout so no second cache format is needed. This is the component behind the 21-25 tokens per second measured on a sixteen-node GB10 cluster running the full K3 checkpoint.