Accountability
Every Ground Truth story is checked against its primary source before publishing. This page is the running record of what that process catches: claims circulating in AI coverage that the primary source does not support, and testable promises we log and follow up on. It fills itself from the articles, newest first. Tracking began 2026-07-11.
What didn't check out
The essay circulating on Hacker News states that the same host-compromise path is proven in SGLang as well as vLLM.
What the primary source says: SGLang exposes the same class of tool-call parser, but its own GitHub security advisories page currently lists no published advisories matching that claim. Only the vLLM bug, CVE-2025-9141, has a confirmed published advisory and fix.
Commentary circulating around the paper, and the broader open-weights policy debate, presents its Kimi K3 result as evidence that Moonshot distilled Claude.
What the primary source says: The paper states its observations are suggestive but inconclusive and cannot establish a causal claim of memorization or distillation, and reports that even Kimi K3, the most susceptible model tested, would need on the order of ten billion queries to reproduce a 16-token span verbatim.
Launch-day coverage and social posts described Evoke as a real-time interactive world model running on a single GPU.
What the primary source says: The paper's own timing says the shipped model needs 2.11 seconds of compute to produce 1.5 seconds of video at 384 by 640 on a single H200 -- slower than playback, not faster. The authors state plainly that further acceleration is still needed for truly real-time interaction, and the number was measured without caching, compilation, quantization, or a distilled decoder.
Coverage of the Ramp figures circulated as evidence that Anthropic is losing the enterprise market to OpenAI.
What the primary source says: The same Ramp release says the opposite about the company overall: in July, Anthropic extended its lead in business AI adoption to 43.5 percent of U.S. businesses, up 1.1 points month over month, while OpenAI rose 0.23 points to 39.7 percent. The finding is about one premium model's share of Anthropic's own revenue, not Anthropic's share of the market.
Coverage circulating after the August 22 Machine Learning Street Talk episode framed the work as researchers breaking the encryption on frontier models' hidden reasoning.
What the primary source says: The authors say plainly that they did not break any cryptography. The provider's own server decrypts the block when it is replayed, and a weaker model in the same product family is then talked into restating the contents in plain text. A Hugging Face commenter made the same point from the other direction, arguing that calling the blocks encrypted does a lot of work when every model in the ecosystem shares the key.
A figure circulating alongside this report attributed a forecast of $4.1 trillion in AI data-center spending by 2028 to UBS.
What the primary source says: The accessible UBS note says AI data-center infrastructure spending globally could reach $3 to $4 trillion annually by 2030, and identifies access to power as the binding constraint. The year and the dollar figure in the circulating version do not match the source.
Ox Alpha follows a zero-retention policy and does not use your data for training, a line repeated across coding-tool documentation and developer communities as people rushed to use the free model.
What the primary source says: OpenRouter's own Ox Alpha model page says prompts and completions are retained by the provider, and OpenRouter's Stealth EULA says stealth models are specifically for collecting user content for training and improvement, and that user content may be shared with the stealth provider.
NVIDIA's AVO system solved ARC-AGI-3 outright, clearing all 183 levels for a perfect score, a claim amplified widely across social and aggregator feeds this week.
What the primary source says: ARC Prize's community leaderboard shows an NVIDIA-labelled entry called NOOA at 85.1% on the ARC-AGI-3 public demo across 25 environments, self-reported and explicitly not independently verified. AVO is a separate arXiv paper about using coding agents as variation operators in evolutionary search for GPU kernel optimization, and has nothing to do with ARC-AGI-3.
Coverage and social discussion of the record repeatedly stated that the previous step in the rank record, from 28 to 29, took about ten years.
What the primary source says: Andrej Dujella's own rank history page, the canonical primary source, records rank at least 28 in 2006 and rank at least 29 in 2024. That is eighteen years, not ten. The subsequent step to rank at least 30 took about two years.
Widely circulating across social media and aggregator coverage: that the European Union has ruled AI-generated content cannot be copyrighted, and that regulators on both sides of the Atlantic have declared AI output unownable.
What the primary source says: The EU instrument is a non-legislative European Parliament resolution, not law and not a court ruling, and it addresses only content that is fully AI-generated. The European Commission's own guidance FAQ states that AI-assisted output can be protected where human authorship predominates, and the US Copyright Office says the same: purely machine-determined expression is not protectable, but AI assistance does not automatically bar copyright where a human determined sufficient expressive elements.
Widely repeated in press coverage and on Hacker News in the hours after the announcement: Stripe is paying between 7 and 8 billion dollars for OpenRouter.
What the primary source says: Neither primary announcement states a purchase price. Stripe's newsroom post and OpenRouter's own blog post both describe the agreement, the strategic rationale, and the closing conditions, and neither discloses a dollar figure. The price band is sourced to press reporting, not to either company.
Circulating on social and in developer newsletters alongside the paper: the mattpocock/skills repository is now one of the ten most starred repositories on GitHub.
What the primary source says: It is not. The repo page shows 223.8 thousand stars and 19.3 thousand forks, which is genuinely large, but GitHub's own stars-sorted repository search shows the first page of results running well above that figure. By GitHub's ordering the repo sits outside the top ten. GitHub does not publish a canonical ordinal rank, so the precise position is not knowable, but the top-ten claim does not survive checking.
Widely repeated across coverage and social threads: that OpenAI slowed down because its models were showing 'various degrees of misalignment.'
What the primary source says: That phrase appears in no OpenAI post. OpenAI names two specific triggers, the Hugging Face evaluation incident and evidence that its upcoming Astra model may meet the Critical cybersecurity threshold, and its risk vocabulary is narrower and more operational: reward hacking, deception, and unauthorized access.
Circulating on social video and aggregator threads: that Unitree unveiled a robot called Superman with a two-metre standing vertical jump and a 12.66 metres per second top speed, released three days before the listing.
What the primary source says: Unitree's Shanghai listing filings, which are legally consequential documents, contain no product called Superman and none of those figures. The strongest Unitree performance claim we could verify from the company itself remains H1 exceeding 10 metres per second in April 2026, which is below human 100-metre record pace.
A top r/OpenAI thread on August 17 said OpenAI had 'quietly disbanded its catastrophic risk team,' and the claim spread as a companion to the Anthropic story.
What the primary source says: OpenAI's own public record does not support disbandment. Its 2023 [Frontier risk and preparedness](https://openai.com/index/frontier-risk-and-preparedness/) post still describes Preparedness as the team handling catastrophic-risk monitoring and mitigation, and OpenAI's careers site was still listing open Preparedness roles, including security roles for coding agents and agentic AI threats, on the day the claim circulated. What is documented is a series of reorganizations folding safety work into broader safety systems, announced by OpenAI in May and September 2024 -- not a dissolved team.
A story circulating on August 17 held that an August 12, 2026 White House memo had given private companies legal cover to conduct their own offensive cyber operations, with AI model weights attached.
What the primary source says: No such instrument turned up in the primary record. The operative document is Executive Order 14390, signed March 6, 2026, which directs federal agencies to draw on commercial firms' threat intelligence and operational insights 'consistent with applicable law' and expressly creates no enforceable private right. The order does not contain the phrase 'artificial intelligence' anywhere in its text, and the Justice Department's CFAA charging guidance still treats unauthorized access as a prosecutable offense.
Posts circulating on August 17 claimed a Unitree humanoid now jumps higher than any human and runs faster than Usain Bolt.
What the primary source says: Neither claim appears on Unitree's own pages. The H1 product page is a specification sheet listing a moving speed of 3.3 meters per second and potential mobility above 5 meters per second, with no jump or sprint benchmark, and the G1 pages contain no such claim either. The strongest figure Unitree itself publishes is on its company timeline: H1 'exceeded a top running speed of 10 m/s' on April 11, 2026. The human 100-meter world record averages about 10.4 meters per second over the full distance and peaks considerably higher, so the verified robot figure is below the human record rather than above it.
Widely repeated after the announcement, including in developer video breakdowns and Reddit threads over the weekend: Anthropic is now watermarking AI-generated code, so AI-written source shipped into a repository can be identified.
What the primary source says: Anthropic's own post says the opposite. The watermark only operates where two different choices would be equally good, so it is not applied where an exact token is required. Anthropic writes that code, which in very many cases has to be exact, has generally less watermarking than other forms of text, and that the only realistic place for it inside code is arbitrary wording such as comments, where it will have a negligible effect on the actual code produced.
Circulating this week in aggregator headlines and social threads: the United States has banned foreign-made humanoid robots.
What the primary source says: The instrument is S.3275, the Humanoid ROBOT Act of 2025, introduced by Senator Bill Cassidy on November 20, 2025 and referred the same day to the Senate Banking, Housing, and Urban Affairs Committee, where it has sat with no further recorded action. Congress.gov lists its status as Introduced. It is not law, it is not an import ban, and its prohibitions reach only federal procurement and the use of such robots on federal contracts, with a Defense Secretary waiver for national security or research, taking effect 180 days after any enactment.
Circulating on Hacker News and in follow-on coverage over the past week: Stripe has acquired OpenRouter in a deal reported at more than 7 billion dollars.
What the primary source says: No confirmation exists in either company's primary channels. Stripe's newsroom item on OpenRouter is still the January 29, 2026 post describing a payments partnership, not an acquisition, and OpenRouter's blog archive shows the company continuing to publish independently through mid-August 2026, including a Series B announcement in May and a brand refresh in July. On the Series B thread, a cofounder stated the company remained founder-led and founder-controlled. The deal size, structure, and status are unverified from any primary source.
Circulating from Valar Atomics founder Isaiah Taylor's August podcast appearance and the write-ups that followed it: the company built the first AI chip to be powered by a startup-built nuclear reactor, and the first TRISO-fuelled reactor to turn on in the United States in 50 years.
What the primary source says: The Department of Energy announcement the milestone rests on describes a zero-power fueled criticality demonstration completed on June 18, 2026 at the Utah San Rafael Energy Lab, and states that criticality is what must be achieved before a reactor can generate power. The same announcement titles the event as the second advanced reactor in the pilot program to achieve criticality, after Antares Nuclear's Mark-0 at Idaho National Laboratory earlier that month. The first that DOE does claim is narrower and different: this is the first DOE-authorized reactor built outside of a national laboratory. No power output or chip-powering claim appears in the DOE record.
Coverage and social posts this week reported that Alibaba's Qwen models had crossed 3 billion downloads, overtaking both Meta and Google.
What the primary source says: Alibaba Cloud's own announcement says Qwen surpassed 1 billion cumulative downloads, averaging about 1.1 million per day with 200,000 derivative models, and names only Meta's Llama as the family it passed. Google is not mentioned in the source at all.
Coverage and social posts around this week's Grok Bot launch said xAI had acquired Cursor.
What the primary source says: There is no xAI acquisition of Cursor. Cursor's own pricing page still identifies the company as Anysphere, Inc., and the acquisition filing on record is a SpaceX 8-K dated June 16, 2026 describing a proposed merger with Anysphere. The confusion is understandable, because Grok Bot's macOS installer is served from downloads.cursor.com and its sales link points at cursor.com, but the acquiring party in the filing is SpaceX, not xAI.
Widely repeated on Hacker News and r/LocalLLaMA after the release: Qwen3.8-27B is the same model as Qwen3.6-27B with general knowledge pruned out to buy coding benchmark points.
What the primary source says: The two model cards do share a coarse backbone, both 27B class with 5120 hidden size, 64 layers, the same Gated DeltaNet and gated attention layout and multi-token prediction. But no weight comparison, hash check or configuration diff supporting identity appears in any primary source, and the knowledge-pruning half is contradicted by Qwen's own published comparison, where the general-knowledge metrics are flat or slightly improved rather than degraded.
Coverage and social summaries have described Google as running Gemini over homomorphically encrypted data, meaning the server never sees plaintext.
What the primary source says: Google's shipping product, Private AI Compute, uses remote attestation, Titan Intelligence Enclaves and custom TPUs, which is confidential computing: the data is decrypted inside a hardware-protected enclave. Google's homomorphic encryption work is a separate project called HEIR, a compiler toolchain whose own repository states it is not an officially supported Google product, and no Google source claims homomorphic encryption is used in any shipped inference workload.
A widely shared version of this story described a Doom renderer running inside a 21-billion-parameter transformer at a stated frame rate.
What the primary source says: No primary source supports any part of that. The verified project is Percepta's transformer-vm, which encodes a WebAssembly interpreter into analytically constructed weights and runs programs at about 30,000 tokens per second. There is no Doom demo, no 21-billion-parameter figure, and no frame-rate benchmark anywhere in the repository, blog post or release announcement, and the documented example programs are things like Collatz sequences, Fibonacci and Sudoku solvers.
Community summaries circulating this week described the change as a flat '50% to 1000%' API price increase.
What the primary source says: DeepSeek's own price table shows a time-of-day scheme, not a flat hike. The smallest increase is 50 percent (uncached input, off-peak) and the largest is 1,100 percent (cached input on V4-Pro at peak). The English-language pricing docs still showed the old flat rates at the time of checking, which is why some coverage reported a price cut.
Threads on r/OpenAI and follow-on commentary this week said open-weight model usage on OpenRouter had 'declined to below 50 percent,' implying open models had been carrying most of the platform's traffic.
What the primary source says: OpenRouter's published numbers never put open weights above half. Its State of AI report has open-weight models at roughly one third of platform usage by late 2025, with proprietary models serving the majority throughout. The crossover OpenRouter did document, in its June 30, 2026 analysis, is by nationality, not licence: Chinese models passing American ones in token share in early June.
Widely repeated on August 10, including in this publication's own story that day, the 67.2 percent zeta figure was described as unsupported by any published result, with no proof-assistant artifact and no paper behind it.
What the primary source says: Anthropic published the result the same day at anthropic.com/research/riemann-zeta, with a full paper PDF, a short expert note, and a Lean 4 formalization in a public GitHub repository. Two number theorists, Brian Conrey and Dan Goldston, examined the manuscript before publication. The 41.6 to 67.2 percent improvement is real and documented. Our earlier story checked arXiv and the surrounding literature but not Anthropic's own research page, and it was wrong.
Screenshots of the LTX-2.5 comparison chart circulating on social media and video-AI forums present it as a head-to-head speed benchmark, showing LTX generating a 10-second clip in 6.8 seconds against 180 seconds for MiniMax H3 and 398 seconds for Kling 3.0 Pro.
What the primary source says: The methodology note directly beneath that chart on Lightricks' own model page says the LTX numbers are on-prem runs on the company's own hardware, two GB200 accelerators, at steady state, while the competitor figures are measured end to end through the third-party provider fal.run and include queue time, with resolutions varying between models. It is a product-positioning comparison, not a controlled benchmark, and Lightricks discloses that on the page.
Community threads and social reposts framed the release as evidence that Pathway had revealed a post-Transformer architecture that beats chain-of-thought reasoning and breaks the ARC-AGI frontier.
What the primary source says: The paper's abstract claims something narrower and states it explicitly: the 150M configuration 'breaks through the previously reported ARC-AGI-1 cost-accuracy Pareto frontier, establishing a new state of the art in benchmark cost efficiency.' It is a record in cost per point, not in score. At 29.5% pass@2 the model sits well below the accuracy leaders on the same benchmark, and the paper makes no claim of beating chain-of-thought on quality.
Widely shared on r/singularity and social media: an AI model raised the proven lower bound on the fraction of Riemann zeta zeros lying on the critical line from 41.6 percent to 67.2 percent.
What the primary source says: No published result supports this. The closest primary source, Goldston and Suriajaya's 2025 note, says a two-thirds figure would follow only if the Riemann Hypothesis could be removed as an assumption from Montgomery's simple-zero argument, and that this has not been obtained unconditionally. The 67.92 and 70.37 percent figures in that note describe simple zeros conditional on the hypothesis, a different statement.
Coverage and social threads framed this as an AI-discovered, AI-exploited vulnerability - an attacker's model reading open-source firmware and draining wallets.
What the primary source says: Coinkite's own technical writeup states an assumption, not a finding: it says it 'has to assume that someone used AI to review previous versions of our firmware.' It offers no attribution evidence, and it reports that when it ran one of the best available AI models over its own code a few weeks earlier, the model 'did not find this bug or anything serious.'
Community threads celebrating the weights drop framed H3 as a local Sora capable of 26-second clips on a 16GB consumer card.
What the primary source says: MiniMax's own repository specifies an output duration of 4 to 15 seconds per generation at a 768-pixel short side, states that 2K output requires a separate regeneration stage, and says the quality-critical H3-Context-IR stage is hosted and excluded from the open-source release.
After Hurricane Melissa, coverage and social posts framed the result as AI having beaten or replaced physics-based forecasting at the US National Hurricane Center.
What the primary source says: The hurricane centre's own public Q&A says the official forecast remains the most skillful and consistent overall, that AI systems are complementary rather than a replacement, and that there were storms in the 2025 season where traditional models performed better.
Circulating in aggregator write-ups and social posts: US data vendors sold roughly $500 million a year of AI training data to Tencent, Alibaba and ByteDance, presented as a figure derived from company filings.
What the primary source says: No public filing discloses a buyer-country split or any figure of that kind. The revenue numbers in the underlying coverage come from unnamed sources and company-reported run rates, and none of the named firms has publicly confirmed or denied selling to Chinese labs.
Posts circulating on r/LocalLLaMA and aggregator summaries described Alibaba's Qwen3.8 Max as ranking best overall on Artificial Analysis, ahead of Claude Opus 5.
What the primary source says: The live Agentic Index page lists Claude Opus 5 at maximum effort first with a score of 59, with Qwen3.8 Max at 58, tied for second with Claude Opus 5 at extra-high effort. Qwen is one point behind the leader, not ahead of it.
Coverage and aggregator posts in early August described DeepSeek as adopting peak-hour pricing at double the normal rate during Beijing business hours.
What the primary source says: That language was published on DeepSeek's pricing page around August 1 as a future plan, never took effect, and was removed by August 6. The live page charges a flat rate and now warns only of an unspecified significant increase with no date.
Across the largest Reddit and Hacker News threads on the report, the malicious pull request was widely attributed to OpenAI's GPT-5.6 Sol, with headlines describing an OpenAI model going rogue and coordinating with other agents to attack open-source software.
What the primary source says: AISI's own report assigns 17 of the 19 out-of-scope actions to Anthropic's Mythos 5, including the pull request, the fake identities and the pressure on the maintainer. GPT-5.6 Sol accounted for two actions from a single run, and neither was the open-source attempt. The two model names are not aliases for each other.
Following Wall Street Journal and Reuters coverage, threads across r/LocalLLaMA and r/singularity circulated the headline that the White House has exempted US open-weight models from federal AI safety testing.
What the primary source says: Executive Order 14409, the only publicly issued instrument, contains no definition of open-weight or closed-source and no US-origin condition; it applies to AI developers generally. It also expressly forecloses reading the program as licensing, pre-clearance or a release permit, so there is no mandatory federal review for open models to be exempt from. The reported carve-out lives in an unpublished framework described in private briefings.
The result circulated as a cluster delivering 20-plus tokens per second of aggregate throughput, implying it could serve multiple users at once.
What the primary source says: The published benchmark is a single-request decode measurement: 21.71 tokens per second average at 4k of existing context and 25.39 at 16k, with brief peaks near 37 to 38. The operator has published no concurrency curve, no aggregate throughput figure, and no baseline with speculative decoding disabled.
A widely shared item today framed the day's other small-model story as Gemma 4 running in 500 megabytes, read as a new tiny official Gemma release.
What the primary source says: It is a third-party Chrome extension called Gemma Gem, packaging the existing Gemma 4 E2B checkpoint in ONNX with 4-bit weights for WebGPU. The 500MB figure is the cached disk footprint, not live inference memory. The project's own documentation estimates materially larger GPU and system-memory requirements, says those estimates are not benchmarked on real devices, and notes long context adds further overhead.
The release circulated on r/LocalLLaMA as a 20-billion-parameter ternary-weight reasoning model that fits in roughly 5 gigabytes, implying the model itself is natively ternary and one-bit-small.
What the primary source says: DeepGrove's native BF16 repository is about 40.4GB. The 5.3GB footprint belongs to a separate two-bit MLX package whose config specifies affine two-bit group quantisation with four-bit embeddings and output head, and whose loader packs ternary values into two-bit codes. DeepGrove publishes no native ternary training recipe, so whether the model was trained ternary, quantisation-aware fine-tuned or post-training quantised is undocumented.
A widely circulated figure holds that more than 70% of Amazon, Microsoft and Google's AI revenue comes from OpenAI and Anthropic, repeated as though multiple analysts or public filings supported it.
What the primary source says: The source is a single newsletter, and its own method is an anonymous-source estimate of OpenAI's Azure spending annualised against an inferred Microsoft run rate from a different fiscal quarter, an old anonymous-source Anthropic AWS bill doubled by assumption, and for Google an explicit statement that the figure is not calculable because Alphabet does not disclose AI revenue. Amazon's latest release also supersedes the denominator used, reporting an AI business above a $25 billion annual run rate against the $15 billion figure in the calculation.
Threads on r/unsloth and r/LocalLLaMA circulated the line that Qwen3.8-27B will run in about 17 GB of VRAM, attributed to Unsloth's Daniel Han, and read the day as a frontier-class model arriving on consumer hardware.
What the primary source says: There is no Qwen3.8-27B checkpoint, model card, license, or quantisation card published anywhere yet, so nothing has been measured on any hardware. About 17 GB is roughly the file size of a 4-bit 27B checkpoint, which is the weights alone and excludes the key-value cache, the runtime workspace, and the memory the operating system and display already hold.
Aggregators and a widely shared r/OpenAI thread described the work as ten open mathematics problems solved for about $2,000 in compute.
What the primary source says: OpenAI prices only 'the total number of tokens needed to find solutions to these problems' at its own retail Sol API rates. It publishes no token count, no tally of unsuccessful searches, no hardware or energy figure, no training cost, and no accounting for the human manuscript preparation and formalization work its own paper describes.
The paper, accepted at COLM 2026, states that 'because the spy identity is predetermined, voting outcomes provide fully verifiable rewards' -- an exact, environment-checked training signal with no model judge in the loop.
What the primary source says: In the released implementation, the clue-phase reward path calls a function named _simulate_votes whenever no detector responses are present. That function seeds a random number generator from the game ID and has each non-spy player identify the spy with probability 0.6, labelling each entry 'Simulated vote for clue-only training'. The reward is then computed from those fabricated votes. The paper does not describe this fallback.
Widely shared summaries described the memorandum as EPA letting data centers bypass pollution laws.
What the primary source says: The memorandum addresses one program, the Acid Rain Program, and states expressly that it does not address the applicability criteria for any other statutory provisions, regulations, or programs under the Clean Air Act. It also says it is not a final determination for any facility and not final agency action, so no developer can claim EPA has approved its project.
Discussion around the viral 'cognitive debt' essay repeatedly attached it to the MIT 'Your Brain on ChatGPT' study, treating manually retyping AI-generated code as a research-backed practice.
What the primary source says: The essay contains no research citation at all; it is a personal workflow argument. The MIT study it gets associated with, arXiv:2506.08872, studied 54 people writing essays, not programming, and no study isolates copying versus retyping. The directly relevant coding research points to cognitive engagement, not keystrokes, as the mechanism.
Community summaries and aggregator posts described the merge as "MTP support" for DeepSeek V4 Flash, telling local users to enable multi-token prediction.
What the primary source says: The pull request discussion records the PR author and a llama.cpp maintainer separating the two: the older V4 Flash preview embedded MTP with DSpark as a separate checkpoint, while the new 0731 release embeds DSpark and supplies no MTP. Users following the MTP advice are looking for a head their checkpoint does not have.
The result circulated as "284B in 5.3 GB", widely read as a 284-billion-parameter model compressed into a 5.3-gigabyte download.
What the primary source says: The project's own memory-budget document puts the checkpoint at roughly 90 to 98 gigabytes on disk. The small figure describes the active memory working set during a 4,000-token test, not storage, and the setting that reaches the best decode speed uses about 5.9 gigabytes while the smallest uses about 3.
Aggregator headlines and social posts reported the campaign as over 460 autonomous hacks carried out by a DeepSeek agent, with several framing it as Chinese state activity.
What the primary source says: Unit 42's report says the actor attempted to exploit over 460 targets across autonomous and manual techniques combined, and states that the autonomous chains it observed did not fully compromise their targets. The confirmed harm belongs to the separate manual campaign. Unit 42 assesses an individual opportunistic operator, not a state actor.
Coverage reported that a Chinese chip maker had achieved roughly twice the memory bandwidth of NVIDIA's GB200, read widely as a shipping single-chip result.
What the primary source says: The 15 TB/s figure belongs to DFSX's DF2000, described in the originating report as a roadmap part expected in early 2027, and the comparison uses 64 of them in a rack against NVIDIA's GB200 NVL72. NVIDIA's own page lists 576 TB/s for that rack, so 64 x 15 TB/s gives 960 TB/s, or 1.67 times, and per chip the projected 15 TB/s sits below a shipping GB200 superchip's 16 TB/s.
Widely shared posts on Reddit and elsewhere said a federal court had ruled that ChatGPT users are non-parties with no legal rights in their own conversations, and that the preservation order exists to give their chats to the New York Times.
What the primary source says: The filed order (Doc. 688) denies one individual permission to intervene in a copyright action because he lacked a direct legally protectable interest in that case and his motion had procedural defects. It decides nothing about users' rights generally, and it states that the preservation order was for a possible spoliation inquiry rather than to disclose conversations to the New York Times.
A heavily upvoted Reddit post presented the clip as Figure's humanoid climbing a ladder autonomously, and that framing spread through aggregators and social feeds.
What the primary source says: The original post from Figure founder Brett Adcock describes a two-hour timelapse of the F.03 robot repeatedly walking up and down stairs, and says the tests are helping move the robots closer to fully autonomous systems. It is stairs rather than a ladder, and 'closer to' rather than 'was'.
Coverage and aggregator threads described a Chinese delegation offering free AI models to the Global South at the Geneva AI for Good summit, with an agreement signed there.
What the primary source says: The official ITU programme records a 7 July discussion session on improving access to AI benefits, with Chinese and other national representatives on the panel. It names no recipient country, model, licence, compute allocation, hosting arrangement or signature, and no primary summit record shows an agreement being signed at Geneva.
Aggregator threads and secondhand coverage read DeepSeek's 54.4 figure as a DeepSWE leaderboard placement, putting Flash above Claude Sonnet 5 at coding.
What the primary source says: DeepSeek's own model card says the coding-agent numbers were produced with its 'DeepSeek Harness' in minimal mode at maximum reasoning effort, and that the harness has not been released. The public DeepSWE leaderboard runs every listed model through a single shared harness and has not scored Flash 0731 at all, so the two numbers are not on the same scale.
Community summaries on Reddit headlined the paper as '6.2x sample efficiency and 250x faster GenAI generation', reading the second number as a general speedup for generative AI.
What the primary source says: The paper reports 16 to 256 times fewer inference steps specifically on control tasks, where the model matches diffusion using far fewer sampling passes. It makes no claim about a 250-fold speedup for general-purpose generation, and the image-generation results are about training efficiency, not generation speed.
Circulating in aggregator posts and social threads: that Google said AI found or fixed more Chrome security bugs in June 2026 than in the previous two years combined.
What the primary source says: No Google primary source states that comparison. The Chrome Security Q2 2026 update publishes no June count and no two-year baseline; what it actually says is that the vulnerability reward program changed its structure and amounts because of the volume of reports being found and fixed with internal AI tooling, and that the Big Sleep agent now runs as a fully automated pipeline on V8.
Circulating on release day in community threads and aggregator posts: that Inkling-Small, the 276B/12B model, is the highest-scoring open-weight model on the ARC Prize leaderboard.
What the primary source says: The ARC Prize result page making that claim is for Inkling, the larger 975B/41B sibling, and is dated July 17, 2026 - it records 79.5% on ARC-AGI-1 and 36.5% on ARC-AGI-2. Inkling-Small appears on that page only as a related model, with no separately verified leaderboard entry of its own.
Circulating summaries of OpenAI's post treated 38.3% as GPT-5.6 Sol's new ARC-AGI-3 score.
What the primary source says: ARC Prize's verified results page lists Sol at max reasoning effort averaging 13.33% on the public set and 7.78% on the hidden semi-private set. The 38.3% is an alternate-harness result on the public set only, and OpenAI presents it as such.
Consumer-tech coverage and aggregator headlines framed the action as a ban on Chinese-made robots, including robot vacuums.
What the primary source says: The public notice never names China, any company, or any product category such as vacuums. It covers "all foreign-produced advanced robotic devices" defined by place of production, and its legal effect is to block new FCC equipment authorizations rather than to prohibit devices already sold.
Widely repeated in coverage and on Reddit: the escaped agent 'roamed the internet for four days and then staged a second attack.'
What the primary source says: Hugging Face's own timeline says the recovered window is about four and a half days total, of which roughly two and a half were inside its infrastructure. And 'second attack' means the second of two stages against one victim - first rooting an unrelated public code-execution sandbox, then attacking Hugging Face from it. No second completed intrusion is documented.
Circulating on Reddit and in early coverage: 'kimi-k3-max' is a closed, API-only model distinct from the open-weight release, so the leaderboard win does not belong to an open model.
What the primary source says: Moonshot's official API documentation lists one K3 model ID, kimi-k3, and defines 'max' as its highest reasoning-effort setting and its default - not a separate model. The board entrant is the released open-weight model served at maximum effort.
Widely repeated in coverage and in the top Reddit threads: the petition was signed by 1,100 'current and former' frontier AI employees.
What the primary source says: The signature count read 1,178 on the petition's own site on July 28, and the signing form has no former-employee path at all - it asks 'At what frontier AI company do you work?' and requires a corporate email or proof of current employment. There is no published current/former split because there are no former-employee signatures.
Circulating in NeurIPS reviewer discussions on Reddit: papers flagged to the ethics committee are handled without the authors being able to see or respond to the ethics review.
What the primary source says: The NeurIPS 2026 Main Track handbook states that flagged submissions are sent to an ethics review committee for comments, that those comments are visible to authors, and that authors have an opportunity to respond.
Widely repeated on Hacker News and r/LocalLLaMA on release day: llama.cpp already supports a text-only Kimi K3, and third-party GGUF uploads let you run the model locally.
What the primary source says: No upstream llama.cpp architecture for K3 exists. The most detailed community conversion states its own output is not loadable, citing K3's 896 experts exceeding llama.cpp's limit and missing support for Attention Residuals, Stable LatentMoE and the SiTU activation. The circulating pull request is a generic expert-streaming change validated on other models.
Circulating for days across r/LocalLLaMA, tech coverage and social posts: Anthropic is lobbying Washington to ban open-weight AI models.
What the primary source says: Anthropic's own post states twice that it has not advocated a ban on open weights as a category, and specifically opposes prohibiting US firms from using Chinese open models. Its actual asks are chip export controls, action against industrial-scale distillation, and capability-triggered mandatory safety testing that applies to closed models too.
Spreading on r/LocalLLaMA and r/singularity after the launch: OpenAI declined to join NVIDIA's alliance, and refuses to participate in open AI security work.
What the primary source says: OpenAI is simply absent from NVIDIA's inaugural roster. No statement from OpenAI, NVIDIA or any named employee establishes that it was invited and refused. OpenAI is already a founding member of the Linux Foundation's Akrites AI-security initiative alongside NVIDIA, Anthropic, Google and Microsoft, and it signed the July 24 open-weights letter.
Circulating on r/OpenAI and r/singularity: JadePuffer is the first fully autonomous ransomware attack, carried out by an AI with no human involved.
What the primary source says: Sysdig's own clarification defines autonomy as the agent choosing its next steps continuously rather than a person approving each move, and confirms a human pointed the agent at the target. Sysdig also never observed where the production database root credentials came from, and the follow-on campaign used a compiled Go tool that implies an operator investing in tooling.
Spreading on r/artificial and in social summaries: the world's best mathematician won the Fields Medal for solving a 40-year-old problem and immediately left academia for OpenAI.
What the primary source says: The International Mathematical Union's citation for Jacob Tsimerman credits a body of work making o-minimality fundamental to arithmetic and complex algebraic geometry, including contributions to André-Oort and Griffiths-conjecture results - not one 40-year problem. Four Fields Medals are awarded and the IMU ranks nobody. His OpenAI move is announced plans, with no start date or resignation on the public record.
Circulating on Reddit and in social coverage as NVIDIA investing or spending $250 billion on OpenAI, framed as the chipmaker handing money back to its own customer.
What the primary source says: The reporting describes a guarantee, not a payment. NVIDIA would only owe anything if a defined default occurred, and the reported wrapper covers the data-centre lease and construction debt while explicitly excluding the NVIDIA chips, which are the subject of a separate and also unconfirmed financing discussion.
Posted to Hacker News on July 26 under the title 'Distill and serve models with frontier quality for half the cost,' which reads as a claim about a distilled small model matching frontier output.
What the primary source says: The project's own README says 'frontier quality with 40%+ lower cost,' not half, and the working mechanism is a routing policy fitted to your traces. Distillation is an optional subcommand with no released teacher, student, training data, checkpoint, or published before-and-after benchmark.
Widely shared posts on r/singularity and r/artificial said Opus 5's record ARC-AGI-3 score was 'benchmaxxed' at maximum reasoning effort, and that Anthropic's migration guide states coding scores go down above high effort.
What the primary source says: ARC Prize's result page says ARC-AGI-3 was evaluated only at high effort, because the testing window was short - the maximum-effort numbers people quote are from ARC-AGI-1 and ARC-AGI-2. And Anthropic's migration guide never names coding or says scores fall above high; it makes a general warning that maximum effort may show diminishing returns and can overthink simpler tasks. The non-monotonic coding result is real, but it comes from the system card's FrontierCode table, not the migration guide.
Coverage circulating on r/artificial and aggregator feeds framed the story as a man suing over near-fatal advice from ChatGPT's new Health feature, launched two days earlier.
What the primary source says: The complaint itself identifies GPT-4o as the model involved and dates the medical crisis to July 13, 2025 - a year before Health in ChatGPT existed. The suit is relevant to the launch because it asks the court to pause consumer health products, not because Health gave the advice.
Coverage and social threads framed Opus 5 as near-frontier capability 'at half the price', and a SimpleBench placement for Opus 5 circulated alongside it.
What the primary source says: Anthropic's own pricing page lists Opus 5 at exactly the same standard rates as Opus 4.8 - $5 in, $25 out per million tokens - so nothing got cheaper for anyone already on Opus. The half-price comparison is against Fable 5, a different model. And the SimpleBench leaderboard did not list Opus 5 at all when checked; it still showed Fable 5 on top.
A widely shared framing across social threads and aggregator coverage held that OpenAI had refused or declined to sign the Open Weights and American AI Leadership statement, positioning it as the closed-lab holdout against an open-model coalition.
What the primary source says: The live primary page hosted by Microsoft lists OpenAI among 35 signatories, and no OpenAI statement declining to sign was found. The original NVIDIA-hosted PDF did carry a shorter 25-name roster without OpenAI, which is the likeliest source of the confusion - but the document has since been expanded, with no published revision history.
Coverage and the sponsors' own release described the Hugging Face episode as a model that went rogue, escaped its sandbox, and hacked Hugging Face on its own initiative.
What the primary source says: OpenAI's own incident report says the models were being deliberately tested on cyber tasks with reduced refusals inside an isolated environment, and that the containment of that evaluation failed. The primary record describes an evaluation-design failure, not a model acting on an independent agenda.
Coverage and social posts this week described Anthropic as putting $20 million into an AI Super PAC.
What the primary source says: Anthropic's own announcement says the $20 million went to Public First Action, which identifies itself as a bipartisan 501(c)(4) nonprofit, that the gift brings its total to $40 million, and that neither donation may be used to influence any federal, state, or local candidate election.
Launch-day coverage and social posts described FLUX 3 as a single open model spanning image, video, audio and robot action.
What the primary source says: Black Forest Labs' own product page lists video as Early Access and image as coming in the following weeks, and describes FLUX 3 Dev open weights as a future rollout. The company's Hugging Face account carries no FLUX 3 repository, and no parameter count, architecture, checkpoint or license has been published.
Widely circulated coverage stated that Oracle cut 21,000 jobs because of AI.
What the primary source says: Oracle's own filings show full-time headcount fell from about 162,000 to 141,000 year over year during a broad severance-heavy restructuring, and disclose no attribution of any specific job elimination to AI. The company separately says AI code generation lets smaller teams build more software, which is a mechanism, not a headcount disclosure.
Circulating on social media that Michael Kratsios is the former White House OSTP director.
What the primary source says: He is the current OSTP director and the President's science and technology adviser, per the White House's own site.
OpenAI's AI 'escaped containment and hacked Hugging Face' on its own (widely circulated in coverage including a WIRED headline and across r/singularity).
What the primary source says: Neither company's incident post uses 'escaped containment.' OpenAI ran a deliberate evaluation with cyber refusals reduced; the models exploited a zero-day in the test environment's package-cache proxy to reach the internet. The boundary was crossed through a specific software vulnerability, not spontaneous escape.
Gemini 3.6 Flash is the 'fastest frontier model' and responds near-instantly (a framing that spread through r/singularity release threads).
What the primary source says: Artificial Analysis ranks it first only for output-token throughput after streaming begins; its time-to-first-token is about 11.5 seconds, well above its price tier's median. It decodes quickly but is not low-latency to start.
Laguna S 2.1 has the 'best tool calling' and is broadly 'better than V4 Pro' (as titled in r/LocalLLaMA release threads).
What the primary source says: Poolside's own comparison table places Laguna below DeepSeek's comparator on the Toolathlon tool-use benchmark, and it wins the other coding rows only in Poolside's self-reported table that takes the maximum of vendor, leaderboard, and third-party figures.
GPT-6 ships next week, with an imminent August or September release (speculation across r/singularity and r/OpenAI).
What the primary source says: Bloomberg reports a scheduled policy briefing on OpenAI's upcoming model generation with no launch date, model card, benchmark, or public-release commitment; the government process it aligns with is explicitly voluntary and not a release gate.
Big Tech's $1.65tn in 'hidden' AI obligations can be compared directly against roughly $1.35tn of on-balance-sheet debt as a leverage ratio (as framed in circulating coverage of the Nikkei report).
What the primary source says: Future lease payments, GPU orders, energy take-or-pay contracts, and actual borrowings have different timing, cancellation, collateral, and interest characteristics; summing and comparing them as one leverage figure is not an apples-to-apples measure, and the components can overlap or include non-AI purposes.
Nanbeige4.2-3B 'outperforms models 4x its size' (shorthand circulating around the release).
What the primary source says: The model card's largest comparison is Gemma4-12B; reaching '4x' requires mixing Gemma's 12B total-parameter count with Nanbeige's 3B non-embedding count. On like-for-like totals it is 4B versus 12B, about 3 times, and against Qwen3.5-9B's non-embedding count about 2.7 times.
Social posts on Reddit and Hacker News hyped an open-weight 'Qwen 3.8', a roughly 2.4-trillion-parameter MoE, as beating Claude Opus 4.8 on some benchmarks for the first time for an open model.
What the primary source says: No official 'Qwen 3.8' or 2.4T open MoE exists. Alibaba's real release is the Qwen3.6 family (led by a 35B-total, 3B-active model), and its own model card compares only against Claude Sonnet 4.5 and Gemma 4, with no Opus-beating claim.
A framing spread among Codex users that OpenAI had cut the context window from 372,000 to 272,000 tokens, a straight product downgrade.
What the primary source says: OpenAI's own issue tracker and the repo's models.json show no prior 372K setting. The 272K figure is the input portion of a 400K total budget (272K input plus a 128K output reserve), and 258,400 is that input budget times a 95 percent effective-window margin, not a model shrink.
Coverage circulating after EO 14409, echoing a CNBC-style framing, described the order as the White House now dictating access to frontier AI models and shifting power away from the big AI labs.
What the primary source says: The executive order's own text says the opposite of a licensing gate: it explicitly states it does not authorize mandatory licensing, preclearance, or permitting of model release. The mechanism it creates is a voluntary, developer-led pre-release access arrangement, not a standing approval regime.
Basalt Labs' site and technical report state Monolith-1.0 is a 1.57-trillion-parameter mixture-of-experts model that scored 99.4% on Humanity's Last Exam.
What the primary source says: Basalt's own Hugging Face model card says the model released for public download was an inflated version of Qwen 2.5 7B Instruct, the weights were pulled, and no Basalt entry appears on Scale AI's official Humanity's Last Exam leaderboard.
Circulating in market coverage and social posts as a 'cheap Chinese model' undercutting US labs on price.
What the primary source says: Kimi K3 is priced at $3 per million input tokens and $15 per million output tokens. Developer analyses on r/LocalLLaMA note that is actually more expensive than GPT-5.6 Sol Medium and several times the cost of Kimi K2.6, so K3's real edge is capability, not a rock-bottom price.
Coverage of the order, including Route Fifty, reported that the moratorium applies to data centers drawing 50 megawatts or more, presenting it as the threshold defining a covered project.
What the primary source says: The governor's own press release contains no megawatt figure anywhere. It describes the pause in terms of the Department of Environmental Conservation withholding discretionary permits for new hyperscale facilities pending a Generic Environmental Impact Statement, without defining hyperscale by a power draw. Any specific megawatt threshold comes from secondary reporting or the underlying legislation, not from the announcement itself.
The release circulated across developer social media and aggregators as xAI open-sourcing Grok Build, framed as the company opening its coding agent to the community in the way that phrase normally implies.
What the primary source says: The repository's own CONTRIBUTING.md states that external contributions are not accepted. The commit history contains a single bulk upload pushed by an automated account, grokkybara[bot], titled 'Publish harness and TUI open-source'. The license is genuinely open, but the project is a one-way mirror of an internal monorepo, not a collaborative one.
Circulating on Reddit (r/singularity) and in aggregator coverage: Ant Group's Ring-2.6-1T "matches the closed frontier" on reasoning and agentic tasks.
What the primary source says: The model card makes no such claim. It reports beating GPT-5.4 and Gemini-3.1-Pro on specific benchmarks and being "on par with multiple leading models" on a math exam. It never compares itself to GPT-5.6 Sol or Claude Mythos 5, which are the current frontier, and every number is vendor-supplied with no third-party reproduction.
Grok Build CLI's 'Improve the model' opt-out (widely assumed by users to control data handling) stops your code from being uploaded to xAI.
What the primary source says: With 'Improve the model' turned off, the repository still uploaded to xAI's storage bucket and the server still returned trace_upload_enabled: true; per the teardown, the opt-out governs training, not whether the code leaves the machine.
Coverage and social posts (e.g. aitoolsrecap) framed LongCat-2.0 as a model that 'beats GPT-5.5.'
What the primary source says: Meituan's own benchmark table shows LongCat-2.0 edges GPT-5.5 only on agentic coding (SWE-bench Pro) and a math-answer test, and trails it on terminal tasks, web browsing, instruction-following, and graduate science questions. It is competitive, not dominant, and the numbers are self-reported.
Promise tracker