Ground Truth.
AI, checked against the source.

← All topics

local-ai

Everything on Ground Truth tagged “local-ai” — 40 items.

Jeff releases a small open model for fast typed decisions on local hardware News

Jeff is an open 0.8B decision-model project that returns calibrated option choices instead of prose, illustrating where a small local model can be useful without claiming frontier reasoning ability.

Kev brings Jev-like local decision models to Qwen, while audits challenge universal calibration claims News

Kev released open local models for typed probability decisions, as published evaluations and an out-of-distribution audit show that confidence calibration varies sharply by task.

Apple's M5 Ultra makes large local AI practical by capacity, not by beating an RTX 5090 News

A 256 GB M5 Ultra Mac Studio ran large sparse models and long contexts locally in a detailed review, but an RTX 5090 remained faster whenever the tested workload fit in its VRAM.

Bonsai 2 ships a 27B local model in a 5.95 GB download News

Prism's ternary Bonsai 2 compresses its smallest 27B language-weight pack to 5.95 GB, but the headline is a disk-size claim rather than a full runtime or VRAM requirement.

A 753 billion parameter model ran on a single workstation GPU News

A serving system called FreeToken reports running frontier-scale sparse models on ordinary personal hardware, including a 753 billion parameter model on one workstation GPU and a 35 billion parameter model on a laptop with 8 gigabytes of video memory.

The open world model ships inference and keeps the training code News

AlayaWorld released inference code and pretrained weights for an interactive world model with long-horizon memory, but the training code is still an unchecked box, the license is a community license, and running it requires a gated Google model plus a ByteDance depth model.

DeepSeek's new open model is 1.6 trillion parameters and runs 49 billion of them per token News

DeepSeek published DeepSeek-V4-Pro on Hugging Face with 1.6 trillion total parameters, 49 billion activated per token, and a one-million-token context window, making it the largest openly downloadable model of the current frontier wave.

China's biggest memory maker is booked through 2027 News

ChangXin Memory Technologies has reportedly sold out its DRAM output through the end of 2027 as PC brands rushed to secure supply, and consumer memory prices have stayed near their highs since.

llama.cpp shipped DSpark for DeepSeek V4 Flash, and almost everyone called it the wrong name News

llama.cpp release b10228 merged speculative-decoding support for DeepSeek V4 Flash, but the new 0731 checkpoint embeds DSpark and ships no MTP at all, so the widely repeated "MTP support landed" advice points users at a head their model does not contain.

Quantizing V4 Flash's KV cache in llama.cpp changes which tokens it picks News

A community experiment found that llama.cpp's generic 8-bit KV cache leaves DeepSeek V4 Flash's average perplexity almost unchanged while altering which tokens make the model's shortlist about one time in eight, because the model's attention makes discrete block-retrieval decisions that a small numerical error can flip.

Offloading and streaming: running a model bigger than your memory Lesson

Offloading and streaming let a machine run a model far larger than its memory by keeping only the parts needed right now in fast memory and fetching the rest from system RAM or disk on demand, trading speed for capacity.

Kimi K3 runs in 8 gigabytes of RAM, at 33 seconds per token News

A hand-written C engine generates text with Moonshot's 2.8-trillion-parameter Kimi K3 using a peak of 8.24 gigabytes of RAM and no GPU, by reading the model's four-bit experts directly off disk, at a rate of roughly one token every 33 seconds.

DeepSeek's low effort setting writes more than its high setting, because the dial is just a prompt News

A reproduction across both a local copy and DeepSeek's hosted API found V4 Flash consuming substantially more tokens on its low reasoning-effort setting than on high, and the released encoder explains why: low injects no instruction at all while high prepends a paragraph demanding exhaustive deliberation.

A 284-billion-parameter model with a 3-gigabyte working set, and a 96-gigabyte disk bill News

An open-source engine called Mference runs DeepSeek V4 Flash on a 24 GB Mac with an effective memory footprint of about 3 gigabytes by streaming each token's experts off the SSD, but the checkpoint still occupies 90 to 98 gigabytes of disk and the test ran at a 4,000-token context.

A New Campaign Argues You Have a Right to Run AI on Your Own Computer News

A grassroots advocacy site, Right to Local Intelligence, is campaigning against proposed state laws it says could require a license just to download and run open AI models, framing local AI as the next personal computer.

Ollama nearly doubles Gemma's speed on Macs by guessing ahead News

A free local-AI tool now runs Google's Gemma model far faster on Apple computers using a trick where a small model drafts words and the big one checks them in bulk.

A model that rivals the frontier now squeezes onto a single high-end desktop News

Aggressive compression shrinks GLM 5.2 by more than 80 percent while keeping most of its accuracy, putting a near-frontier model within reach of local hardware.

llama.cpp b10228 Tool

The release that adds DeepSeek V4 Flash's embedded DSpark speculative-decoding head, plus a converter that can split the draft tensors into a separate GGUF. Gains are workload-dependent: roughly 2x decode on large multi-GPU setups, and a measured slowdown on a 24 GB card with CPU offload.

h3-studio Tool

Local front end for MiniMax H3 built on ComfyUI 0.30.0 or newer, with a VRAM meter, idle GPU release and a service setup so the video model hands the card back to other workloads. Its notes document H3 needing roughly 15.5 GB in the NVFP4 build, which is the practical ceiling on a 16 GB card.

bitnet.cpp Tool

Microsoft's official inference framework for 1.58-bit ternary language models, built on llama.cpp with optimized CPU and GPU kernels for running very heavily compressed models on ordinary hardware.

Voodoo Dynamic Quant Tool

MIT-licensed tooling that learns per-tensor mixed-precision allocations and exports llama.cpp-compatible GGUF quantizations.

Von 1.0 Tool

An Apache-2.0 395M ModernBERT-based model and runtime for local structured decisions and HTTP serving.

Unsloth Tool

Toolkit and documentation for running and fine-tuning large open models faster and on smaller hardware, including aggressive dynamic quantization recipes that shrink models like GLM 5.2 by 80-plus percent while keeping most of their accuracy. The practical on-ramp to running near-frontier models privately.

TurboQuant-MLX Tool

Quantization tooling for MLX with published size and speed measurements, including a 3-bit path that takes a 120-billion-parameter model from about 63GB down to 48GB on consumer Macs.

SwiftLM Tool

An MLX-based runtime for Apple silicon whose --stream-experts mode reads mixture-of-experts weights straight off an NVMe SSD, letting a machine run models several times larger than its RAM.

SHADOW-50M Tool

A proof-of-concept 44M-parameter CPU model with exact arithmetic circuits, disk-backed retrieval and WebAssembly support.

Ornith-1.0-9B GGUF Tool

The quantized build of the smallest member of the MIT-licensed Ornith-1.0 family, an open agentic coding line post-trained on top of Gemma 4 and Qwen 3.5. The Q4_K_M file is 5.63 gigabytes, which puts it within reach of a single consumer GPU.

Ollaya Tool

A public local runtime for typed, inspectable decision-model inference with a developer-oriented API.

Ollama 0.31 Tool

Run open models on your own computer; the new version nearly doubles Gemma's speed on Apple Silicon using multi-token prediction, on by default.

MicroLLM Lab Tool

A WebGPU browser lab for running and comparing several tiny quantized language models without an API key or local Python stack.

MiMo-V2.6-Distill-Qwen-9B Tool

Xiaomi's 9B Qwen3.5-based supervised distill for coding, agent, visual-coding, and cybersecurity research.

Mference Tool

Runs DeepSeek V4 Flash on Apple silicon by keeping the shared core, attention and cache resident while streaming each token's routed experts off the SSD. Publishes an unusually honest memory budget: about 3 GB working set, 90 to 98 GB on disk, tested at a 4,000-token context on a 24 GB Mac, with no quality parity test yet.

Kev Tool

Apache-2.0 local Qwen-based models that return typed probability decisions for yes/no, choice, and rating tasks.

Jeff Tool

Open code and downloadable small models for structured option selection, probabilities, and low-latency local decision loops.

Gemma-4 12B Coder (GGUF) Tool

A fine-tuned, locally-runnable version of Google's Gemma-4 model specialized for programming tasks, packaged in a format that runs efficiently on everyday consumer hardware.

GLM-5.2 Tool

A flagship openly-available language model with a very large context window for long documents and code. Free to download and run yourself, with compressed versions for more modest hardware.

GLM 5.2 (GGUF, runnable locally) Tool

Zhipu AI's open, MIT-licensed mixture-of-experts model with a roughly million-token context, now packaged as ready-to-run quantized files you can host on your own machine. Strong on agent and coding workflows; this week it beat Claude on a narrow security benchmark at a fraction of the cost.

DeepSeek-V4-Pro Tool

A downloadable 1.6-trillion-parameter mixture-of-experts model that activates 49 billion parameters per token, with a one-million-token context window under an MIT license. Serious server hardware required, but the weights are yours.

Bonsai 2 Tool

Prism's downloadable ternary 27B model and runtime for local text, tool-calling and optional vision workflows.

AlayaWorld Tool

Inference code and pretrained weights for an autoregressive world model with real-time camera control, prompt switching and long-horizon memory consistency, using an explicit 3D cache for spatial recall plus a compressed frame-history embedding. Training code is not included and the weights ship under a community license.