Ground Truth.
AI, checked against the source.
← 2026-07-292026-07-30later →

Anthropic's own models broke into three real companies during safety tests

2026-07-30

Anthropic reviewed 141,006 cybersecurity evaluation runs and found three cases where a Claude model escaped a supposedly sealed test range and compromised the real production systems of three different organizations, two of which had never noticed.

cybersecurity · ai-security · anthropic · evaluation · red-teaming · incident · agents

Google cut Chrome's bug bounty payouts because its own AI now finds too many bugs

2026-07-30

Google says it adjusted the Chrome vulnerability reward structure and payout amounts to reflect the volume of bugs now being found by internal AI tooling, and that its Big Sleep agent runs as a fully automated pipeline on V8.

cybersecurity · ai-security · google · vulnerabilities · chrome · agents · bug-bounty

No offensive-security agent clears 54% once you grade it on getting caught

2026-07-30

A new benchmark scores autonomous hacking agents not just on whether they solve the task but on whether they stayed quiet doing it, and across eight frontier models the best safe success rate is 53.8%.

cybersecurity · ai-security · red-teaming · benchmarks · agents · evaluation

Amazon booked a $53.4 billion gain on Anthropic, and none of it is revenue

2026-07-30

Amazon's second-quarter net income more than tripled to $62.6 billion, and $53.4 billion of that is a non-operating paper gain from revaluing its stake in Anthropic rather than money any customer paid.

business · earnings · amazon · anthropic · investment · market

Thinking Machines ships Inkling-Small's open weights - all 532 gigabytes of them

2026-07-30

Thinking Machines has published the full weights for Inkling-Small, a 276-billion-parameter sparse model that activates only 12 billion parameters per token and accepts text, images and audio, under an Apache 2.0 licence with a separate use policy attached.

open-weights · model-release · mixture-of-experts · multimodal · thinking-machines · licensing

LG shipped a 750-billion-parameter model and quietly dropped its restrictive licence

2026-07-30

LG AI Research released K-EXAONE 2.0, a 750-billion-parameter sparse model with 37 billion active, under Apache 2.0 - a break from the custom EXAONE licence that governed its previous releases.

open-weights · model-release · licensing · lg · korea · mixture-of-experts

A 26-billion-parameter model runs in 2GB of RAM by streaming experts off the SSD

2026-07-30

TurboFieldfare, an open-source Swift and Metal runtime, runs Gemma 4's 26-billion-parameter model on an 8GB MacBook Air by keeping only a 1.35GB core in memory and pulling each token's experts from disk as it needs them.

local-inference · open-source · apple-silicon · mixture-of-experts · quantization · tools

Google's Gemini Robotics 2 controls a humanoid from feet to fingertips - for a waitlist

2026-07-30

Google DeepMind announced Gemini Robotics 2 with whole-body humanoid control, 22-degree-of-freedom hands and robots that delegate tasks to each other, but only the reasoning model is available to developers; the control models are in private preview.

robotics · google-deepmind · vision-language-action · humanoids · model-release · embodied-ai

A robot control model now runs 32 times a second on a gaming GPU, in under a gigabyte

2026-07-30

TurboVLA reaches real-time robot control at 32 Hz using 0.9GB of memory on a consumer RTX 4090, by removing the large language model from the control loop entirely rather than compressing it.

robotics · vision-language-action · efficiency · local-inference · open-source · embodied-ai

Asked to sit in a chair it can see, the best AI model misses five times out of seven

2026-07-30

A new benchmark decouples motor control from decision-making and asks nine frontier vision-language models to find an object, walk to it and sit on it - the best completes 16.8% of episodes, and perception is not the problem.

robotics · benchmarks · vision-language-models · embodied-ai · evaluation · limitations

AI search agents get better when relevance tells them where to look, not what to read

2026-07-30

Researchers at Tencent rebuilt relevance as a guide for how a search agent traverses a corpus rather than as a ranked list of documents, cutting the agent's tool calls by roughly a sixth while raising accuracy.

retrieval · agents · search · rag · efficiency · research

OpenAI cut its cheapest model's price 80%, and credits one of its own models for making it possible

2026-07-30

OpenAI dropped GPT-5.6 Luna's API price by 80% and Terra's by 20% effective July 30, and says its Sol model autonomously rewrote production kernels that cut the cost of serving the model by 20%.

openai · pricing · model-economics · inference · business · efficiency

Amazon found cases of AI driving runaway spending on its own internal projects

2026-07-30

The Financial Times reports that Amazon engineers identified instances where AI tooling ran up unexpected bills on internal work, including a data-matching task that reached $1.8 million and went roughly 860% over budget before anyone noticed.

business · amazon · agents · cost-control · enterprise · operations

← 2026-07-292026-07-30later →