Ground Truth.
AI, checked against the source.

← All topics

coding

Everything on Ground Truth tagged “coding” — 45 items.

Z.ai changed only the post-training, and the model learned to find exploits News

Z.ai released GLM-5.3 on August 14 using the same base model as GLM-5.2, with every gain coming from post-training, and the largest jump was in finding and exploiting software vulnerabilities.

A new terminal benchmark drops the best agent from 84 percent to 34 News

Terminal-Bench 3.0 launched with 74 tasks across seven domains, and the top agent scores 34.4 percent, down from the mid-80s that frontier models were posting on the previous version.

xAI shipped Grok 4.6 into Cursor at two dollars a million input tokens News

xAI released Grok 4.6 on August 12, claiming a score of 61 on the Artificial Analysis Intelligence Index, tying GPT-5.6 Sol Max, and priced it at two dollars per million input tokens.

Walmart put a token allowance on its in-house coding AI News

Walmart replaced unlimited access to its in-house AI coding tool with a fixed per-employee token allotment, and its CTO says the reason is duplicated requests rather than the bill.

Stack Overflow took 1,490 questions in July News

Stack Overflow received 1,490 new questions in July 2026, down from 6,414 in July 2025 and 176,610 in July 2014 -- a 118-fold collapse in the public programming corpus that trained today's coding models.

DeepSeek re-trained V4 Flash without touching the architecture and its coding-agent score went from 7 to 54 News

DeepSeek published new MIT-licensed weights for V4 Flash on July 31 that change only the post-training, lifting the model's score on a real-world software-engineering agent test from 7.3 to 54.4 out of 100.

13.6% of SWE-bench Verified pairs a bug report with a patch that does not match it News

A systematic audit of SWE-bench Verified, the benchmark used everywhere to rank AI coding ability, found that 68 of its 500 tasks link an issue to a pull request that fixes something else, adds unrelated work, or only partly addresses the report.

Kimi K3 topped a fullstack coding board at maximum effort News

Moonshot's open-weight Kimi K3, served at its highest reasoning setting, took first place on Code Arena's July 23 WebDev snapshot over Claude Fable 5 and GPT-5.6 Sol, though the live board has since moved it to second.

Poolside's Laguna S 2.1 shipped with a broken chat template - and the fixes explain the reviews News

Days after releasing its open coding model, Poolside has been repairing it in public: the base chat template shipped with reasoning disabled by default and a 32,768-token generation cap, and its own quantised builds needed re-releases to fix tool calls and thinking.

Poolside's Laguna S 2.1 is a small open coding agent with big benchmark claims News

Poolside released Laguna S 2.1, a public-weight coding model with an unusually low 8 billion active parameters that runs locally on a single high-end machine, but its claims of beating DeepSeek V4 Pro come from the company's own benchmark table and one early hands-on tester found it fabricates facts when evidence runs out.

A Chinese open-weight model is now shipping inside GitHub Copilot News

GitHub has made Moonshot AI's Kimi K2.7 Code generally available as a selectable model in GitHub Copilot, hosting the Chinese-developed open weights on US Azure infrastructure so prompts never reach Moonshot, while a viral claim that Microsoft is secretly testing the newer Kimi K3 remains unconfirmed.

SpaceXAI ships Grok 4.5, trained on trillions of Cursor coding sessions News

SpaceXAI released Grok 4.5 on July 8, its first model as a public SpaceX subsidiary, trained on trillions of Cursor developer-interaction tokens and priced aggressively at $2 per million input tokens, though that rate only holds below 200K context.

OpenAI ships GPT-5.6 and bets on efficiency, not raw intelligence News

OpenAI publicly launched GPT-5.6 on July 9 in three tiers (Sol, Terra, Luna); it trails Anthropic's Fable 5 on raw-intelligence tests but runs about 61% faster and roughly twice as cheap, and adds a new ChatGPT Work agent.

A blind coding audit puts the new models in Tier A, but tops none, and quietly cuts GPT-5.5 by 11 points News

An independent blind-audited coding benchmark placed GPT-5.6 Sol (92) and Grok 4.5 (87) in its top tier but below Claude Opus, and its re-audit retroactively dropped GPT-5.5 from 96 to 85, exposing how unstable single-run model scores are.

OpenAI says a leading coding benchmark can no longer tell the best models apart News

OpenAI published an analysis concluding that SWE-Bench Pro, a widely-cited coding benchmark, has hit a roughly 70% noise ceiling where higher scores may reflect quirks rather than real skill, and retracted its recommendation to use the benchmark to rank frontier models.

Grok 4.5 arrives claiming Opus-class quality at a third the price News

SpaceXAI released Grok 4.5, a 1.5-trillion-parameter model priced at $2 per million input tokens and $6 per million output, undercutting frontier rivals roughly threefold while claiming comparable coding quality.

Anthropic's own data says the best coders gain the most from AI News

By studying hundreds of thousands of real coding sessions, Anthropic found that experienced engineers get more out of AI assistants, not less, a direct challenge to the idea that AI levels the playing field.

Can an AI Agent Reproduce Real Science? A New Test Says: Rarely News

A new benchmark points coding agents at the actual computational results behind ninety papers in top journals. The strongest models matched the published science on fewer than one in five.

A Coding AI Ran Through Uber's Yearly Budget in Four Months News

Uber gave Claude Code to about 5,000 engineers, who loved it. By April the company had burned through its entire 2026 AI budget, exposing how badly old software pricing fits new agent tools.

The AI That Now Writes Most of Its Maker's Code News

Anthropic says more than 80 percent of the code it ships is now written by its own model, Claude, and the more interesting numbers are about judgment.

A Free Model That Splits Your Work Across 300 Helpers News

Moonshot AI's Kimi K2.6 is a frontier-grade model anyone can download, and its headline trick is fanning a single job out to hundreds of helpers working in parallel.

AI coding skill in Python doesn't carry over to other languages News

A widely-trusted coding benchmark was Python-only. Expanding it to a dozen languages revealed that models acing Python often stumble badly elsewhere — Python skill isn't general coding skill.

agent-skills Tool

Addy Osmani's collection of production-grade, reusable skills for AI coding agents, trending near the top of GitHub this week.

Xiaomi MiMo-V2.5-DFlash Tool

Xiaomi's official DFlash release on Hugging Face -- a 1-trillion-parameter mixture-of-experts model (42B active) under an MIT license, with FP4 quantization and parallel decoding for high inference throughput.

Qwen3.6 (open weights) Tool

Alibaba's stable Qwen3.6 release: open-weight general chat and coding models you can self-host, the same family at the center of this week's open-vs-closed pricing debate.

Poolside Laguna S 2.1 (GGUF) Tool

Open-weight 118B mixture-of-experts coding agent activating about 8B parameters per token, under the permissive OpenMDW-1.1 licence, in GGUF plus FP8, NVFP4 and INT4 builds. Use the current re-released Q4/Q8 files - the initial ones shipped with a broken chat template.

Poolside Laguna S 2.1 Tool

A public-weight, 118B-total mixture-of-experts coding model with only ~8B active parameters that runs locally on a single 128GB machine via a 75GB Q4 GGUF, built for long-horizon agentic software work under the permissive OpenMDW-1.1 license.

Ouroboros Tool

An agent harness that improves its own tools, prompts and core implementation through reviewed commits, which then become the runtime for its next task. Public code, with benchmark campaigns run on frozen snapshots so the numbers mean something.

LongCat-2.0 Tool

Meituan's 1.6T-parameter MoE model tuned for coding and agentic work, MIT-licensed weights plus a cheap hosted API (launch promo $0.30/$1.20 per million tokens) that self-hosts to avoid data-jurisdiction concerns.

Kimi K3 Tool

Moonshot AI's 2.8-trillion-parameter flagship with a 1M-token context window, tuned for agentic coding and knowledge work; it topped a frontend-coding leaderboard. Usable now via kimi.com chat and an OpenAI-compatible API, with open weights due July 27.

Kimi K2.7 Code Tool

Moonshot AI's trillion-parameter mixture-of-experts coding agent, with only 32B active per token, a 256K context, and vision input, now selectable inside GitHub Copilot and downloadable under a Modified MIT license.

Kimi K2.6 weights (Hugging Face) Tool

The actual Kimi K2.6 model weights, published under a modified-MIT license for anyone to download, run, and build on; large enough that full-strength use needs a multi-GPU node.

Kimi (Kimi K2.6) Tool

Moonshot AI's web assistant and agent, running the open-weight Kimi K2.6 model; free to use in the browser for chat and long-horizon agent tasks, with the weights also downloadable for self-hosting.

Grok 4.6 Tool

xAI's new frontier model, tuned for long-running agents and available day one in Cursor, Grok Build, and the xAI API. Two dollars per million input tokens and six per million output, with a faster variant at double the price.

Grok 4.5 Tool

SpaceXAI's new 1.5-trillion-parameter model, available in Grok Build, Cursor, and the API at $2 per million input / $6 per million output tokens, with a full public release on July 9.

Gemma-4 12B Coder (GGUF) Tool

A fine-tuned, locally-runnable version of Google's Gemma-4 model specialized for programming tasks, packaged in a format that runs efficiently on everyday consumer hardware.

GPT-5.6 (Sol / Terra / Luna) Tool

OpenAI's newest model family, tuned for cheap, fast, reliable agentic work, with programmatic tool calling, a multi-agent beta, persisted reasoning, and a high-reliability 'pro' mode.

GLM-5.2 on Baseten Tool

The top trending open-weight model served as a fast hosted endpoint, reported at 280+ tokens/sec on Blackwell-class hardware -- an open model you can call like a closed one.

GLM 5.2 (GGUF, runnable locally) Tool

Zhipu AI's open, MIT-licensed mixture-of-experts model with a roughly million-token context, now packaged as ready-to-run quantized files you can host on your own machine. Strong on agent and coding workflows; this week it beat Claude on a narrow security benchmark at a fraction of the cost.

DeepSeek V4 Pro (API) Tool

A strong open-weight reasoning and coding model now offered through DeepSeek's own API at a permanently cut, low per-token price, undercutting frontier closed models for high-volume work.

Cursor Design Mode Tool

Edit a running web app by clicking elements, drawing on the page, or describing the change out loud, and Cursor rewrites the underlying code with the app hot-reloading as it goes. Visual context instead of file paths.

Cursor Tool

The AI-native code editor whose real-world developer interaction data trained Grok 4.5; a mature, widely-used tool for agentic coding across many models.

Claude Sonnet 5 Tool

Anthropic's new most-agentic mid-tier model, close to its flagship on hands-on tool and coding work; now the default on Free and Pro plans.

Claude Code auto mode Tool

A permission mode that replaces per-action approval prompts with a separate classifier model which blocks escalation, unrecognized infrastructure and actions driven by injected content; becomes the default on 14 August 2026.

Claude Code Tool

Anthropic's command-line coding agent that reads a whole codebase, edits files, runs tests and fixes failures on its own; it is the tool behind Anthropic's disclosure that Claude now authors most of its production code.