News · 2026-10-09
Saluki offers a Qwen-derived coding and tool model in a 7.89 GB file
Saluki’s publisher released a Qwen-derived model as a 7.89 GB download and reports retaining 96 percent of average benchmark performance across nine tests. The file runs with the llama.cpp runtime, making the release a concrete option for local experimentation. Its own results vary by task, and the model card does not establish a verified minimum running-memory requirement in the dossier.
Key facts
- The selected IQ2-mix GGUF download is 7.89 GB; the optional vision add-on is separate.
- Repository history checked October 9 is consistent with an October 8, 2026 upload.
- The ConwayResearch model card attributes Saluki to Underdog and identifies Qwen3.8-27B as its base.
- The authoritative source is Saluki’s Hugging Face model card.
The publisher’s headline is “96% across 9 benchmarks.” That is an aggregate retention claim, not a universal guarantee of quality. The difference matters because local users choose a model for particular jobs: a short structured tool call, a coding task, a long explanation, or a mathematics problem. An average can hide a loss exactly where one user needs reliability.
Saluki is a compressed version of a larger model. Quantization stores numerical weights in a more compact representation. Think of replacing a detailed map with a carefully simplified edition. The simplified map can preserve the roads most travelers need while losing fine detail. Different kinds of trips reveal different weaknesses.
The model card compares its 7.89 GB file with a 54 GB reference size. Those are the publisher’s listed storage figures, not measurements of the memory needed during inference. The actual downloadable variant is identified as an IQ2-mix GGUF file, a format with a public specification. Model-file formats explain why the format and variant matter: a model name by itself does not specify which set of files a reader will install.
The release also offers an optional vision component. Readers downloading that add-on need to account for it separately, rather than assume the main file includes every feature. The model card reports an Apache 2.0 license and credits Qwen and an intermediate quantization from ISTA-DASLab. It does not provide a detailed recipe for Saluki’s own tuning, so the observed differences cannot be confidently assigned to one particular training or compression step.
The task-level results are more informative than the headline. The publisher reports higher scores on function calling, parallel tool calls, and instruction following, while recording lower scores on several mathematics and coding comparisons. That pattern does not establish that compression improves a model in general. It says the released artifact has a particular measured profile under the publisher’s tests.
A concrete example comes from the function-calling comparison. Saluki passed 88 of 120 frozen tasks, compared with 84 for the full-model baseline. That is a difference of four tasks in a modest evaluation, not a broad demonstration of superiority. The card itself calls the test limited and says several full-model comparison results come from public evaluations using different harnesses.
A tool call is the model selecting an operation and supplying its arguments, such as choosing a calendar function with a date and event title. Getting those arguments right can be more useful to an agent than writing an eloquent paragraph. Our coverage of local runtime tool hosting explains why serving a model and operating its tools are separate responsibilities. A test of isolated calls is not the same as evaluating a complete agent workflow with retries, permissions, changing state, and unexpected tool responses.
The hardware question also needs care. The dossier establishes the file’s disk size, but no primary-source minimum or recommended GPU memory. That requirement is unstated in the verified material. The download should not be relabeled as a 7.89 GB graphics-card requirement: running a model also involves context state, working memory, software overhead, and any additional components. No estimate is substituted for a published or measured configuration.
For a developer, the useful experiment is therefore narrow. Use the same prompts and harness for Saluki and a relevant baseline, record correct tool actions and completed tasks, and examine failures rather than averaging them away. This is the ordinary discipline of AI benchmarking, applied to a small downloadable artifact instead of a frontier leaderboard.
The supplied LocalLLaMA capture records interest in the size-reduction and retention claim, but contains no replies that establish independent performance. Community pickup is a reason to test the model, not a reason to skip testing it.
Saluki’s immediate contribution is practical accessibility: a named, licensed artifact with a verified selected-file size and compatibility with an established local runtime. Its strongest limitation is the same one that accompanies many compressed-model launches. Publisher comparisons are useful starting evidence, but mixed harnesses, a small tool-call sample, and missing running-memory guidance leave readers responsible for validating the trade-off on their own workload.
Key questions
How large is the Saluki download?
Does 96 percent retention mean Saluki is equally good at every task?
How much GPU memory is required to run Saluki?
Cite this
APA
Ground Truth. (2026, October 9). Saluki offers a Qwen-derived coding and tool model in a 7.89 GB file. Ground Truth. https://groundtruth.day/news/saluki-compresses-qwen-into-789-gb-download.html
BibTeX
@misc{groundtruth:saluki-compresses-qwen-into-789-gb-download,
title = {Saluki offers a Qwen-derived coding and tool model in a 7.89 GB file},
author = {{Ground Truth}},
year = {2026},
month = {oct},
url = {https://groundtruth.day/news/saluki-compresses-qwen-into-789-gb-download.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.