Ground Truth.
AI, checked against the source.

← All topics

models

Everything on Ground Truth tagged “models” — 29 items.

DeepSeek releases a 168 GB MIT-licensed multimodal V4 checkpoint News

DeepSeek's V4-Flash-Vision-Exp is an MIT-licensed 168 GB downloadable multimodal model whose strongest comparisons remain vendor results under DeepSeek's own harness.

Artificial Analysis changed its leaderboard's ruler, not just its rankings News

Artificial Analysis Intelligence Index v4.2 doubles the share of held-out/private data to 40% and removes saturated GPQA Diamond, making its methodology shift the story as much as any score.

Anthropic ships the same model behind two different safety boundaries News

Anthropic says Claude Fable 5.1 and restricted Mythos 5.1 share underlying capability, making safeguards and access policy—not a new weight set—the central product difference.

Meta's Muse Spark 1.3 caught GPT-5.6 on one scoreboard and still trails Claude News

Meta released Muse Spark 1.3 on September 2, 2026, and Artificial Analysis scored its public tier at 61 on its Intelligence Index, level with OpenAI's GPT-5.6 Sol, while the same measurement puts Anthropic's Fable 5.1 four points ahead of Meta's best variant.

Google's new Flash model scores higher and costs more to finish a job News

Google released Gemini 3.8 Flash on September 2, 2026, and the per-token price is unchanged, but the model deliberately spends about 30% more output tokens per task, pushing measured cost per task from roughly $0.40 to $0.58.

DeepSeek gave its cheapest model eyes and did not change the price News

DeepSeek shipped an experimental vision version of its V4-Flash model that accepts images by base64, URL, or file upload, and bills it at exactly the same rate as the text-only model.

Gemini Omni 1.1 Flash can extend a scene instead of restarting it News

Google's updated video model reads up to ten seconds of a clip's prior context before continuing it, up from one second, and adds keyframe control, cheap 360p drafts and 4K upscaling through the Gemini API.

fal post-trained MiniMax H3 and kept the weights News

Inference company fal released H3 Max, a post-trained version of the open-weight MiniMax H3 video model that renders a five-second 768p clip in under three seconds -- available only as a hosted API, with no weights published.

Qwen put a 20-million-entry n-gram table inside a model News

Alibaba's Qwen released Qwen3.8-Flash-Next, a preview of the architecture behind Qwen4, whose headline idea is scaling parameters through a 20-million-entry table of word pairs and triples that can be offloaded off the GPU -- 51 billion parameters that never need to be computed, only looked up.

Google's new transcription model edits what you said News

Gemini 3.5 Transcribe removes filler words, silently resolves speakers' self-corrections, and can make function calls out of the transcription layer -- which makes it excellent for voice agents and unusable as a verbatim record.

GLM-5.3-Flash was Ox Alpha, and it ran on Chinese chips News

Z.ai released GLM-5.3-Flash under an MIT licence and confirmed it is the anonymous \u201cOx Alpha\u201d model that topped OpenRouter for a week -- served, the company says, entirely on a cluster of Chinese AI accelerators at per-token cost comparable to NVIDIA hardware.

Ant published Ling-3.0-flash's weights under plain MIT, with no rider attached News

Ant Group's InclusionAI lab published the full weights for its 124-billion-parameter Ling-3.0-flash model on Hugging Face this week under an unmodified MIT license, with no acceptable-use policy, revenue threshold, or branding requirement anywhere in the release.

Claude Opus 5 posts a verified four-fold lead on the hardest adaptation benchmark News

Anthropic released Claude Opus 5 on July 24, and the independent benchmark owner ARC Prize verified it at 30.16% on ARC-AGI-3, roughly four times the previous best published result, while the model's API price stayed identical to Opus 4.8.

Ant's Ling-3.0-flash goes live free: 124 billion parameters, 5 billion doing the work News

Ant released Ling-3.0-flash on July 23, a 124-billion-parameter model that activates only about 4% of itself per token, with a 256,000-token context and free access on OpenRouter and Vercel's AI Gateway.

Poolside's Laguna S 2.1 is a small open coding agent with big benchmark claims News

Poolside released Laguna S 2.1, a public-weight coding model with an unusually low 8 billion active parameters that runs locally on a single high-end machine, but its claims of beating DeepSeek V4 Pro come from the company's own benchmark table and one early hands-on tester found it fabricates facts when evidence runs out.

Nanbeige4.2-3B reuses one 22-layer stack twice to punch above its size News

A Chinese lab released Nanbeige4.2-3B, a small open-weight model that runs its 22 transformer layers twice in sequence to get 44 layers of depth from one set of weights, posting benchmark numbers rivaling models three times its size, though the results are vendor-reported and a widely repeated 'beats 4x its size' claim does not survive clean accounting.

Gemini 3.6 Flash: Google ships a faster worker, not a bigger brain News

Google released Gemini 3.6 Flash into general availability, and independent benchmarks show it streams output nearly twice as fast as 3.5 Flash and costs less per task while scoring the same on a leading intelligence index, though it still takes a conspicuous 11-plus seconds to start responding.

A Chinese open-weight model is now shipping inside GitHub Copilot News

GitHub has made Moonshot AI's Kimi K2.7 Code generally available as a selectable model in GitHub Copilot, hosting the Chinese-developed open weights on US Azure infrastructure so prompts never reach Moonshot, while a viral claim that Microsoft is secretly testing the newer Kimi K3 remains unconfirmed.

DeepSeek's new open models give everyone a million-word memory by default News

DeepSeek previewed two free-to-download V4 models that can read a million tokens at once, no longer as a premium add-on but as the standard setting.

The best free AI model just landed — but almost nobody can run it at home News

A powerful open model anyone can legally download has reignited the open-vs-closed debate — but it's so large that 'open' now means 'open if you own a small server.'

Open vs. closed AI models — what "open weights" really means Lesson

Some AI models you can only rent through a company's interface; others you can download and run yourself. That difference — open weights vs. closed — shapes privacy, research, cost, and who controls the technology.

A powerful open model lands and reignites the open-vs-closed debate News

A Chinese lab released a flagship model anyone can download and run, with a huge memory for long documents — and a viral claim that it makes things up less than a top closed model.

Qwen3.8-Flash-Next Tool

Alibaba's preview of the architecture behind Qwen4: 125 billion parameters with 6 billion active, a 20-million-entry n-gram embedding table, and a 262k context extensible to a million tokens. Weights are 360 GB in bf16 under the Qwen Community License 1.0, which allows commercial use and fine-tuning but requires a separate licence to run a model-as-a-service or an AI coding-assistant business.

Ox Alpha Tool

An anonymous reasoning model with a 1,048,576-token context window and 131,072 max output tokens, accepting text, image, and video with tool and JSON support. Free to try in the browser, with no disclosed creator.

Ollama Tool

Download and run open AI models locally with a single command. The easiest on-ramp to running your own model.

LM Studio Tool

A friendly desktop app to find, download, and chat with open models on your own machine — no command line needed.

Hugging Face Tool

The main hub for finding, downloading, and trying open AI models and datasets — the field's town square.

GLM-5.3-Flash Tool

Z.ai's 320-billion-parameter multimodal model with 18 billion active per token and a one-million-token context, released under the MIT licence -- one of the most permissive terms any model this size has shipped under. The download is 328 GB of already fp8-quantized weights, and it runs locally through SGLang, vLLM or TokenSpeed.

GLM-5.3 on OpenRouter Tool

Z.ai's GLM-5.3 with a 1 million token context window and always-on reasoning, billed at $1.40 per million input tokens and $4.40 per million output, with cheaper cache reads. Tuned for long-horizon software engineering and vulnerability discovery.