Ground Truth.
AI, checked against the source.
← 2026-07-282026-07-29later →

OpenAI says GPT-5.6 Sol autonomously rewrote the code that serves it, cutting serving costs 20%

2026-07-29

OpenAI published an engineering account on July 29 saying GPT-5.6 Sol, working through Codex, autonomously rewrote its production GPU kernels and redesigned its own draft model, contributing to a 20% cut in end-to-end serving cost and a 15% gain in token-generation efficiency.

openai · inference · agents · efficiency · kernels · gpt-5-6

Two API settings tripled OpenAI's ARC-AGI-3 score without touching the model

2026-07-29

OpenAI reported on July 29 that enabling retained reasoning and compaction lifted GPT-5.6 Sol from 13.3% to 38.3% on the ARC-AGI-3 public task set while using six times fewer output tokens, an identical model scoring three times higher because of harness settings.

benchmarks · openai · agents · evaluation · arc-agi · context-windows

The FCC just added every foreign-made advanced robot to its national security Covered List

2026-07-29

On July 28 the FCC added all foreign-produced advanced robotic devices and foreign-produced power inverters to its Covered List, blocking them from new equipment authorizations, on national security determinations that cite remote commandeering and surveillance risk rather than naming any country or company.

cybersecurity · supply-chain · policy · robotics · hardware · regulation

Meta's AI build swallowed 98% of its cash flow in a single quarter

2026-07-29

Meta reported second-quarter 2026 revenue of $60.80 billion, up 28%, but $31.08 billion of capital spending left just $784 million in free cash flow, and the company narrowed 2026 capex guidance upward to $130-145 billion.

industry · meta · compute · capex · earnings · data-centers

A $500 fine-tune of a 9B open model beat all five frontier models it was tested against

2026-07-29

A consultancy reinforcement-trained a 9-billion-parameter open model on a simulated product-catalog review workflow for about $500 of GPU time, and it outscored the best of five frontier configurations while costing $0.50 per thousand listings against $34.

fine-tuning · open-weights · reinforcement-learning · cost · enterprise · agents

A frozen 12B model answers already-solved problems at zero generation tokens

2026-07-29

A technical report describes a 12-billion-parameter model whose weights never change but which answers new instances of nine previously solved problem families with no generated tokens at all, scoring 180 out of 180 by executing verified stored procedures instead of reasoning again.

inference · agent-memory · verification · efficiency · research · determinism

A handheld gripper and a head camera can now train robots with no robot demonstrations

2026-07-29

Researchers report that raising the fidelity of handheld human demonstrations removes the need for any robot teleoperation on the target task, with policies trained on handheld data alone reaching parity with robot-taught baselines on four two-armed tasks.

robotics · manipulation · datasets · imitation-learning · research · vla

NVIDIA shipped a drop-in kernel that nearly halves video generation time

2026-07-29

NVIDIA released code on July 28 for Sol-Attn, an attention kernel that decides which parts of a long video to compute exactly while approximating the rest inside a single pass, reporting up to 2.1 times faster video generation with no retraining and no weight changes.

video-generation · inference · nvidia · open-source · sparse-attention · efficiency

A new distillation method lets the teacher model interrupt the student mid-thought

2026-07-29

Researchers found that when a student model starts reasoning down a wrong path, its teacher's next word tends to be a redirection like But or Wait, and turned that disagreement into an automatic trigger for the teacher to briefly take over.

distillation · training · reasoning · research · on-policy · llm-training

Two papers attack the same waste: coding agents rediscovering the same repository every session

2026-07-29

CodeNib builds reusable lexical, semantic and structural views of a repository per commit and cuts an agent's exploration tokens by 50 to 87%, while a companion benchmark finally measures the file-finding stage that patch-success scores hide.

coding-agents · retrieval · developer-tools · research · benchmarks · context-engineering

A new benchmark grades video models on film craft instead of whether clips look nice

2026-07-29

FilmBench scores text-to-video and reference-to-video models against professional cinematic criteria such as camera language, shot continuity and performance, using prompts reverse-engineered from professionally selected film clips, with the dataset and toolkit released publicly.

video-generation · benchmarks · evaluation · open-source · research · filmmaking

Researchers built a model whose dangerous knowledge can be switched off like a module

2026-07-29

A method called GRAM routes risky training data into small auxiliary modules that can be turned on or off after training, so one model can approximate several models each trained without a different category of dangerous data, tested from 50 million to 5 billion parameters.

cybersecurity · ai-security · alignment · access-control · red-teaming · open-weights

← 2026-07-282026-07-29later →