Ground Truth.
AI, checked against the source.
← 2026-08-022026-08-03later →

Qwen3.8-Max Shipped as a Paid API, Not as Open Weights

2026-08-03

Alibaba put Qwen3.8-Max live as a hosted API at $2 per million input tokens and $6 per million output tokens, a fifth cheaper than the model it replaces, while the open weights it promised for Max and a 27B sibling have not shipped.

qwen · alibaba · model-release · pricing · open-weights · local-llm · api

MiniMax Shipped H3's Weights and Kept the Best Part Hosted

2026-08-03

MiniMax released the weights for its H3 video-and-audio generation model, and its own model card says the input-processing stage that is critical to output quality is not included in the release and the 2K output stage is not open-sourced at all.

minimax · video-generation · open-weights · licensing · model-release · multimodal

Three Days On, Nobody Has Publicly Compiled OpenAI's Ten Proofs

2026-08-03

OpenAI's repository of Lean proofs for ten mathematics results has 434 stars and 39 forks but exactly one commit, no pull requests, and no issues, and no third party has published a build log showing the proofs check.

openai · mathematics · formal-verification · reproducibility · lean · evaluation

OpenAI Rebuilt Voice So the Model Itself Decides When to Talk

2026-08-03

OpenAI's engineering posts on GPT-Live describe removing the separate turn detector from the audio path entirely and cutting session startup from six network round trips to one, treating a voice conversation as a live media system rather than a model feature.

openai · voice-ai · full-duplex · systems · latency · infrastructure

NVIDIA's Open Full-Duplex Voice Model Wants an 80GB GPU

2026-08-03

NVIDIA released an 11-billion-parameter speech model that listens and speaks at the same time and calls tools mid-conversation, and its own documentation requires a GPU with at least 80 GB of memory and lists more than a dozen failure modes.

nvidia · voice-ai · full-duplex · open-weights · speech · tool-use

An RL Trainer That Invents Its Reward When the Judge Says Nothing

2026-08-03

The published code for SpyRL, a reinforcement learning method built on the promise of fully verifiable rewards, silently substitutes randomly generated votes with a hard-coded 60 percent accuracy rate whenever no judge outputs are present.

cybersecurity · supply-chain · ai-security · reinforcement-learning · reproducibility · research-integrity

The Cheap 284B Rig Is Really 768GB of Server Memory

2026-08-03

A builder running DeepSeek V4-Flash at 33 tokens a second on two RTX 3090s is holding about 6.6GB of weights per card and roughly 170GB per instance in system memory on a four-socket enterprise server, which is where the model actually lives.

deepseek · local-llm · mixture-of-experts · hardware · inference · quantization

A Munich Court Found Suno's Models Memorised Six Songs

2026-08-03

The Regional Court of Munich I largely granted GEMA's claims against Suno, holding that six well-known works were reproducibly stored in Suno's models and could be extracted through simple prompts, and assigning responsibility to Suno rather than its users.

copyright · music-generation · suno · regulation · training-data · germany

EPA Says an Off-Grid Plant Built for One Data Center Escapes the Acid Rain Program

2026-08-03

An EPA guidance memorandum states that a fossil-fuel power plant with no physical connection to the utility grid, built to serve only an adjacent private data center, falls outside the federal Acid Rain Program and its permit, allowance, and monitoring requirements.

policy · energy · data-centers · regulation · epa · infrastructure

Robot Policies That Predict the Touch Before They Make It

2026-08-03

Two matched robotics releases from NeoteAI and Fudan give manipulation policies a sense of touch that anticipates contact rather than reporting it, winning all nine real-robot tasks in the authors' benchmark against strong vision-only baselines.

robotics · tactile-sensing · vision-language-action · world-models · manipulation · open-source

A New Benchmark Asks Whether a Coding Agent Can Stop Asking

2026-08-03

CAPA tests whether an assistant that has watched one developer resolve the same ambiguity before can write the intended code without asking again, and finds that the best model still needs a clarification round on four sessions in ten.

benchmarks · coding-agents · agent-memory · personalization · evaluation · llm

Two Essays About AI and Your Brain, and One Actual Study

2026-08-03

A randomized experiment found that AI assistance impaired conceptual understanding, code reading, and debugging while delivering no significant average speed gain, which supports the concern behind this week's viral developer essays but not their proposed fix.

developer-productivity · research · ai-assistants · education · coding-agents · skills

← 2026-08-022026-08-03later →