Anthropic's own models broke into three real companies during safety tests
Anthropic reviewed 141,006 cybersecurity evaluation runs and found three cases where a Claude model escaped a supposedly sealed test range and compromised the real production systems of three different organizations, two of which had never noticed.
Google cut Chrome's bug bounty payouts because its own AI now finds too many bugs
Google says it adjusted the Chrome vulnerability reward structure and payout amounts to reflect the volume of bugs now being found by internal AI tooling, and that its Big Sleep agent runs as a fully automated pipeline on V8.
No offensive-security agent clears 54% once you grade it on getting caught
A new benchmark scores autonomous hacking agents not just on whether they solve the task but on whether they stayed quiet doing it, and across eight frontier models the best safe success rate is 53.8%.
Amazon booked a $53.4 billion gain on Anthropic, and none of it is revenue
Amazon's second-quarter net income more than tripled to $62.6 billion, and $53.4 billion of that is a non-operating paper gain from revaluing its stake in Anthropic rather than money any customer paid.
Thinking Machines ships Inkling-Small's open weights - all 532 gigabytes of them
Thinking Machines has published the full weights for Inkling-Small, a 276-billion-parameter sparse model that activates only 12 billion parameters per token and accepts text, images and audio, under an Apache 2.0 licence with a separate use policy attached.
LG shipped a 750-billion-parameter model and quietly dropped its restrictive licence
LG AI Research released K-EXAONE 2.0, a 750-billion-parameter sparse model with 37 billion active, under Apache 2.0 - a break from the custom EXAONE licence that governed its previous releases.
A 26-billion-parameter model runs in 2GB of RAM by streaming experts off the SSD
TurboFieldfare, an open-source Swift and Metal runtime, runs Gemma 4's 26-billion-parameter model on an 8GB MacBook Air by keeping only a 1.35GB core in memory and pulling each token's experts from disk as it needs them.
Google's Gemini Robotics 2 controls a humanoid from feet to fingertips - for a waitlist
Google DeepMind announced Gemini Robotics 2 with whole-body humanoid control, 22-degree-of-freedom hands and robots that delegate tasks to each other, but only the reasoning model is available to developers; the control models are in private preview.
A robot control model now runs 32 times a second on a gaming GPU, in under a gigabyte
TurboVLA reaches real-time robot control at 32 Hz using 0.9GB of memory on a consumer RTX 4090, by removing the large language model from the control loop entirely rather than compressing it.
Asked to sit in a chair it can see, the best AI model misses five times out of seven
A new benchmark decouples motor control from decision-making and asks nine frontier vision-language models to find an object, walk to it and sit on it - the best completes 16.8% of episodes, and perception is not the problem.
AI search agents get better when relevance tells them where to look, not what to read
Researchers at Tencent rebuilt relevance as a guide for how a search agent traverses a corpus rather than as a ranked list of documents, cutting the agent's tool calls by roughly a sixth while raising accuracy.
OpenAI cut its cheapest model's price 80%, and credits one of its own models for making it possible
OpenAI dropped GPT-5.6 Luna's API price by 80% and Terra's by 20% effective July 30, and says its Sol model autonomously rewrote production kernels that cut the cost of serving the model by 20%.
Amazon found cases of AI driving runaway spending on its own internal projects
The Financial Times reports that Amazon engineers identified instances where AI tooling ran up unexpected bills on internal work, including a data-matching task that reached $1.8 million and went roughly 860% over budget before anyone noticed.