News · 2026-10-03
Redis creator antirez's ds4 runs DeepSeek-class models on a 128GB Mac, and it is back in the spotlight
ds4 is a compact inference engine written in C by Salvatore Sanfilippo, the creator of Redis who goes by antirez, that runs a handful of very large open-weight models on a single high-end desktop. Its recommended starting point is DeepSeek V4 Flash compressed to about 2 bits, an 81GB download that upstream suggests running on a 96GB or 128GB machine. The project is from May; an October 2 Hacker News post brought it back to wide attention.
Key facts
- Speed: about 34 generated tokens per second at a 32,000-token context on a 128GB M5 Max Mac, by the project's own measurement.
- When: first version in early May 2026; the October 2 Hacker News post linked a community guide, not a new release.
- Who: Salvatore Sanfilippo, MIT-licensed, with AI coding agents credited for heavy assistance.
- Primary sources: the ds4 repository and antirez's post "A few words on DS4".
Most people who run open models locally use general-purpose tools like llama.cpp or Ollama, which try to support thousands of models. ds4 goes the other way. It supports a short list of models it has tested, each with its own carefully chosen file, and tunes the code path for each one. The model guide currently covers DeepSeek V4 Flash and V4.1 Flash, GLM 5.2, GLM 5.3 and GLM 5.3 Flash, and Qwen3.8 Flash Next. It runs on Apple's Metal, NVIDIA's CUDA (including the DGX Spark and multi-GPU systems) and AMD's ROCm.
Why antirez built it
In his launch post, antirez explains that three things came together: strong open-weight models, a DeepSeek Flash checkpoint that tolerated an aggressive mix of 2-bit and 8-bit compression, and years of community work on local inference. Together they made serious local use possible on 96GB to 128GB machines. He describes using a local model for work he would previously have sent to Claude or GPT, and ends with a line that sums up the project: "AI is too critical to be just a provided service."
How it fits a huge model on a desk
DeepSeek V4 Flash is a mixture-of-experts model: it has many specialist sub-networks, and only a few are used for each word. ds4 compresses those rarely-used experts hardest, to about 2 bits per number, while keeping the parts used on every step at higher precision. Think of packing for a long trip by vacuum-sealing the clothes you will wear once and leaving the everyday ones loose. The result for DeepSeek V4 Flash at the "Q2" setting is about 81 GiB of weights.
For machines without enough memory, ds4 can keep the most-used experts in fast memory and read the rest from the SSD as needed, the same idea behind offloading and streaming weights. It also saves the work done on a long prompt so a resumed session does not recompute it, using the KV cache. Memory needs vary by model: the Qwen3.8 Flash Next path keeps about 42 GiB resident and leaves a 95 GiB lookup table on disk, which is how the project documents a 64GB Mac route for it. The full GLM 5.3 at 2 bits is about 197 GiB and needs more memory or streaming.
What "fast" means here
The project's performance guide reports DeepSeek V4 Flash generating about 34 tokens per second at a 32,000-token context on a 128GB M5 Max, falling to about 28 at 65,000 tokens. That is roughly a brisk reading pace once the prompt has been processed. The same table shows about 14 tokens per second on a DGX Spark. These numbers measure raw generation speed with one user, not how well the model does real coding work.
Ground Truth has covered the same models arriving in mainstream tools, including GLM 5.3 Flash support in llama.cpp and a llama.cpp fix for DeepSeek V4 Flash tool calls. ds4 is the specialist alternative.
The caveat
Heavy compression costs something, and nobody has published an independent quality comparison. On the new Hacker News thread, one commenter said the compressed DeepSeek checkpoint "isn't very good," while another said ds4's files beat other popular compressed versions in their own tests. Neither is a controlled evaluation. In the original May discussion, a commenter reported their RTX 5090 ran their workload faster than a similarly priced Mac. Memory figures are for the model weights only; long conversations need extra room on top. And the "new" framing in some social posts is wrong: this is a five-month-old project getting renewed attention through the community guide at dwarfstar.sh.
Key questions
Is ds4 a new release?
How much memory do I need to run ds4?
Is ds4 just a wrapper around llama.cpp?
Cite this
APA
Ground Truth. (2026, October 3). Redis creator antirez's ds4 runs DeepSeek-class models on a 128GB Mac, and it is back in the spotlight. Ground Truth. https://groundtruth.day/news/antirez-ds4-runs-deepseek-class-models-on-a-128gb-mac.html
BibTeX
@misc{groundtruth:antirez-ds4-runs-deepseek-class-models-on-a-128gb-mac,
title = {Redis creator antirez's ds4 runs DeepSeek-class models on a 128GB Mac, and it is back in the spotlight},
author = {{Ground Truth}},
year = {2026},
month = {oct},
url = {https://groundtruth.day/news/antirez-ds4-runs-deepseek-class-models-on-a-128gb-mac.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.