AI papers — 2026-09-29
Jump to one of 85 papers
- Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation
- YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
- Post-Training Leaves Behavioral Shadows on Unrelated Decisions
- Self-Evolving Coding Agents: From Digital Programs to Physical-World Intelligence
- Duplex-MPE: Benchmarking Multi-Party Interaction in Full-Duplex Dialogue
- MassAlloc Attention: Let Attention Allocate Its Own Compute
- TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces
- How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining
- Improving Test-Time Scaling with Adaptive Looped Transformers
- Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
- Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning
- EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks
- CompoWorld: Compositional Environment Scaling for General Agents
- Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
- Recursive Harness Distillation across Agents for Robot Manipulation
- Surprising Success, Repeated Failure: Entropy-Guided Credit Assignment for Exploration in LLM Reasoning
- WorldPlay2: Extending Real-Time Interactive World Models in Control and Horizon
- DepthBench: Measuring How Residual Connections Enable More Computational Depth
- Diffusion Reward Models
- Learning to Learn from Context: Synthetic Training from Perturbed Public Documents
- Skill2Env: Capability-Oriented Environment Synthesis from Skills for General Agents
- QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents
- SpatialSpeak: QA-Native Reconstruction with Local and Global Context for Spatial Chain-of-Thought Reasoning
- RoboFoundry: System-as-Policy Evolution for Self-Learning Embodied Agents
- CoWindow Attention: Full Causal Coverage Is a Collective Property
- REALM: A Coarse-to-Fine Generative Framework for Embodied Reactive Listening
- AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research
- WideSWE: Can Coding Agents Coordinate Changes Across Repositories?
- Just MLPs: Efficient Visual State Reconstruction for Multimodal Language Models
- Change the Product, Keep the Parameters: Associative Algebra Layers for Transformers
- In-Flight KV Cache with Clean Anchors for Faster Autoregressive Video Diffusion
- Structured Residual Connectivity Matters for Diffusion Transformers
- Precise Editing and Flexible Referencing for Interactable Worlds
- FlowTool: Controlling Tool Parameter in Image Retouching via Flow Matching
- SolveEdit: Benchmarking Visual Problem Solving in Generative Models
- SciGen-Verifier: A Multimodal Reasoner for Explainable Verification in Scientific Image Generation
- Program-Verified Self-Evolution for Vision-Language Models
- Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It
- Imprint Reader: From Weight-Update Readout to Behavioral Intervention
- TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining
- InfiniHand: Streaming World-Space Hand Motion Estimation from Egocentric Video
- Residual Transferability in Neural Image Watermarking
- BaRe-Mem: Bayesian Reliability Memory for Robust and Adaptive Agent Consultation
- Fewer Tokens, More Self-Teaching: On-Policy Self-Distillation for Extreme Visual Token Reduction
- Hard Vision, Easy Vision: What GPT-6 Astra Reveals Across Computer Vision
- Nereus: Adaptive Parallelism for LLM Post-Training
- An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning
- GeoVerse: World-Consistent Novel View Synthesis in Geometric Latent Space
- RenderRank: Learning to Rerank Text with Compressed Visual Tokens
- Rethinking Automated Voice Similarity by Shifting from EER to Embedding Geometry
- Relic: From Multi-Agent Collaboration to Persistent Organizational Capability
- Measuring Collapse and Correction in Homogeneous-Panel LLM Debate
- WaveFront Decoding: Parallelized Self-Speculative Decoding for Looped Language Models
- Beyond Timestamps: Decision-Aligned On-Policy Distillation for Long-Horizon Agents
- When Do Model Internals Help? Exploring the Role of Representation Engineering in LLM Safety
- DroneWAM: Efficient World Action Model for Drone Visual Navigation
- Draft-KV: Learning Useful Latent Communication Between Language Models
- TokenCast: Forecasting Token Consumption During LLM Agent Execution
- KVCMAS: Efficient KV Cache Correction for Shared Context in Multi-Agent Systems
- PReCache: Efficient KV Cache Sharing for Multi-LoRA Agents via Low-Rank Precomputation and Neutral Reconstruction
- ColNanoVDR: Document-Free Query Distillation for Multi-Vector Visual Document Retrieval via Optimal Transport
- EvolvingAvatar: Interactive 3D Head Generation That Adapts as Conversations Unfold
- Adaptive Consistency Graph for Long-Horizon Agents
- DISCO: Distributed Long Context Scaling with Grounding-Reasoning Disaggregation
- G^2PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation
- SMAT: Simple and Efficient Merge-Aware Training
- KernelZero: Co-Evolving Proposer and Coder for Continuously Improved GPU Kernel Generation
- On-Policy Self-Distillation for Multi-Turn Image Editing
- Why Deterministic PRM Guidance Underperforms in Discrete Diffusion Reasoning
- SkillDRE: Dual-Stage Red-Team Evolution of Agent Skills via Pre-Execution and Runtime Feedback
- VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis
- ControlScope: Workflow Revision and Reliability in LLM Agents
- Playing to Par: Reinforcement Learning for Provably Optimal Quadrilateral Block Decompositions
- What Masking Geometry Works Best for EEG Foundation Models?
- Safe Error Correction for Language Models: Frozen-Base Adjustment with Capability Preservation
- NanoForecast v0.5: Competitive Time Series Forecasting Through Training Pipeline Optimization
- Not All Objectives Are Born Equal: Priority-Constrained Descent for Hierarchical Multi-Objective Optimization
- When Privacy Moves ML-Mediated Decisions On Device: Information and Incentive Misalignment in Auctions
- How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure
- Training and Inference Dynamics of PLDR-LLMs: Row-Map Collapse, Renormalization, and Predictive Reduction
- How Does "English (US)" Become the Default? Triangulating Structural Bias Towards American English Across the LLM Pipeline
- AdaGuard: An Adaptive Guard Model with User-defined Policies
- Routing Drift Alone Does Not Diagnose Failure in Merged MoE LLMs
- PyroAdapt: Adapting Wildfire Prediction under Spatial Heterogeneity and Temporal Shift
- Rolling-WAM: World Action Models with Rolling Imagination
Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation →
Training a language model with reinforcement learning can produce separate experts in maths, coding or instruction-following, but users want one model that does all three. The authors found that when experts jointly teach a single student, the instruction-following expert's feedback is far more spread out and drowns out the maths expert. Rescaling each expert's feedback by its measured spread produced a consistently better all-round student.
Technical breakdown
- Problem: Multi-teacher on-policy distillation (MOPD) merges specialist models by routing each prompt to the appropriate expert, but unbalanced feedback scales prevent effective capability transfer.
- Method: Domain-Normalized MOPD (DN-MOPD) rescales each domain's distillation advantages using bounded multipliers based on the measured spread of teacher–student log-ratios within each domain. The method maintains label-based routing while applying a clipped ratio of pooled to domain-specific log-ratio standard deviation (clip(σall/σd, 0.25, 4)) to each batch's distillation signal.
- Key results:
- DN-MOPD improves six-task average over MOPD by 1.17–2.47 percentage points at 9B, 4B and 2B across two evaluation budgets
- Instruction-following feedback is 2.3–4.4 times more dispersed than mathematics feedback; DN-MOPD rebalances these scales
- Fixed weights near DN-MOPD's measured multipliers perform comparably at 9B and 4B, showing per-batch calibration matters mainly at smaller scales
- Mathematics gain recovery of most of the lost advantage in MOPD (5.95 pp gain on MATH-500 at 4B)
- Why it matters / caveats: Integrating domain specialists requires controlling not just which expert teaches but also how strongly its feedback counts; most gains come from limiting instruction-following dominance rather than amplifying mathematics alone. Benefits depend on teacher–student configuration, and the clipping bounds frequently sit at limits rather than equalizing all domain scales.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality →
Music AI either writes notation without producing a recording, or generates audio without an explicit composition. The authors built a single system that first writes a readable score setting out melody, harmony and structure, then expands it into a full song recording. Expert listeners preferred songs made this way, and because the score is readable, people can edit the composition and have the audio regenerated.
Technical breakdown
- Problem: Symbolic models produce readable, editable compositions but stop before finished audio; audio models produce complete songs while leaving composition implicit, making editing and controlled generation difficult.
- Method: YuE2 unifies symbolic and audio generation through progressive musical commitment: a single AR–NAR Mixture-of-Transformers first writes an ABC-notation score specifying melody, harmony, rhythm and form, then expands it into 25-Hz MERT2 semantic tokens, and finally realizes acoustic latents for stereo audio. The model is trained on approximately 346,000 hours of music with symbolic supervision from SheetSage2 transcription and semantic tokens from a multi-view target synthesis approach combining MuQ and Qwen2-Audio encoders.
- Key results:
- YuE2 reaches SongBench Global Avg score of 6.7316 on WildSongBench, highest among evaluated public systems; best-of-8 reaches 6.9632
- Experts prefer symbolic planning for overall quality (49.3% vs 34.6% without planning) and musicality (45.0% vs 29.4%)
- Audio quality mean tie-adjusted preference of 58.9% across six proprietary systems
- SheetSage2-AR leads 12 of 15 benchmark–metric pairs in full-song transcription; MERT2 surpasses prior results on 14 of 15 MARBLE metrics
- Why it matters / caveats: Symbolic planning significantly improves perceived song quality and enables controlled editing and agentic revision; however, expert listening shows nearly balanced preferences against Suno v5, suggesting proprietary systems remain competitive despite unified architecture.
Post-Training Leaves Behavioral Shadows on Unrelated Decisions →
When a company privately fine-tunes a language model, the changes show up even in its choices on completely unrelated text. The authors picked prompts where the original public model was almost undecided between two ordinary words, asked the tuned model for just one word each time, and trained a copy on those choices. The copy picked up real skills in coding, science and reading, raising concerns about leakage.
Technical breakdown
- Problem: Private post-training updates invisibly change model behavior on task-unrelated inputs, creating a "behavioral shadow" that is difficult to observe but potentially carries capability-relevant information.
- Method: Active Taskless Distillation (ATD) selects task-unrelated prompts where the public base model is nearly indifferent between two ordinary words (near-tie selection with |qi−0.5|≤0.02), queries a privately trained teacher for one word per prompt, and trains a same-ancestor student using only cross-entropy loss on these prompt–word pairs without target-task examples, teacher logits, or target-task data.
- Key results:
- On HumanEval+, ATD student reaches 51.22% pass@1 with 5,664 single-word observations, gaining 5.34 pp over exact nuisance-matched control
- Transfer generalizes across code, science, commonsense reasoning and reading comprehension with gains of 0.81–5.03 pp across seven tasks
- Mean gains positive across Qwen2.5, Qwen3 and Llama model families; fixed-depth LoRA matching shows 4.80 pp effect across five independently constructed acquisitions
- Code-trained and science-trained shadows selectively improve their matched endpoints (+5.34 and +2.52 pp respectively) with near-zero off-diagonal effects
- Why it matters / caveats: Private post-training updates leak observable signals through task-unrelated decisions, and learning from these observations can transfer capability without target-task data, teacher logits or parameters; however, transfer depends on whether the capability is already expressible by the student model (memorized or cipher-based teachers do not transfer).
Self-Evolving Coding Agents: From Digital Programs to Physical-World Intelligence →
Robots trained to map what they see directly to movements often fail when the layout or camera angle changes, partly because their plan and progress are hidden inside the actions. The authors instead have the robot write out its understanding of the scene and its plan as runnable code, which can be checked, revised and reused. The approach improved success on unfamiliar tasks and on a real robot.
Technical breakdown
- Problem: Vision-language-action (VLA) and world-action models lose more than half their success under small viewpoint or layout changes because task state, execution, and recovery are implicitly encoded in action chunks rather than explicitly represented and verifiable.
- Method: Physical Coding represents task state and execution as executable programs: Code as World records objects, relations, observations, constraints and progress predicates; Code as Policy organizes planning, tool calls, action execution, verification and recovery. HexaAnything integrates language models, state observation through SVG-based code generation, independent verifiers, and action tools (perception, planning, control, VLA/WAM policies) into a unified execution loop where verified traces become reusable data and memory.
- Key results:
- On RoboCasa365, HexaAnything improves Composite-Unseen success over XR-1 VLA; HexaModel trained on Harness traces beats the base on every split
- On PhyBench and dual-arm AgileX robot, the agent autonomously completes physics experiments and most tabletop tasks, often faster than published results
- Explicit workflow, verification and recovery are demonstrated while holding the underlying action model fixed, with first data-to-model update showing traces internalize physical execution
- Why it matters / caveats: Representing execution as executable, verifiable code enables inspection, revision and persistent learning across episodes, moving beyond implicit end-to-end policies; however, full autonomous co-evolution of models, representations, environments and embodiments remains future work, with current results showing preliminary evidence at the Harness, tool and data-to-model levels.
Duplex-MPE: Benchmarking Multi-Party Interaction in Full-Duplex Dialogue →
Voice assistants are usually tested one-on-one, but in a group conversation an assistant must also judge when to stay quiet or stop talking. The authors built a test set of group conversations where a request is made either directly to the assistant or only hinted at, and measured five existing open speech systems. The systems talked often but frequently answered wrongly or failed to keep silent.
Technical breakdown
- Problem: Current benchmarks evaluate turn-taking and interruption handling but do not assess when an assistant should answer, remain silent or stop speaking in shared multi-party conversations where any speaker can request help.
- Method: Duplex-MPE provides 2,000 scenarios with three to four human speakers and one assistant, each with one unresolved request paired across explicit and implicit addressing of the same task. Four metrics measure fresh response initiation, conditional answer accuracy, silence preservation and stopping when a human resolves a request, using continuous multi-party audio without transcripts or turn boundaries as input.
- Key results:
- MiniCPM-o 4.5 leads on three scored capabilities among evaluated speech systems (Moshi, FLM-Audio, Voila, Freeze-Omni)
- Gemini 3.1 Pro responds 64.3 percentage points more often to explicit than implicit requests on transcripts; speech systems show no significant paired response-rate difference
- Frequent speech does not imply appropriate participation; inaccurate answers can coexist with high response rates
- Why it matters / caveats: Silence is a correct output for spoken assistants in multi-party settings, requiring them to follow shared conversation while deciding separately whether participation is needed; current open-weight speech systems struggle with this selective participation task.
MassAlloc Attention: Let Attention Allocate Its Own Compute →
Standard attention makes a model do full work for every pair of words, even though most pairs barely matter. The authors designed an attention method that still looks at every allowed pair but only does the expensive follow-up work where the contribution is meaningful, using a single tolerance setting for both training and use. It matched ordinary attention's quality while running markedly faster on long inputs.
Technical breakdown
- Problem: Full softmax attention assigns negligible normalized mass to much of the causal score space yet dense kernels execute complete post-score computation (softmax, value loading, attention accumulation) for every legal causal tile.
- Method: MassAlloc Attention (MALA) preserves score access to every legal causal tile while using normalized contribution to allocate post-score computation. Forward uses evolving online-softmax normalizer to bound contributions; backward reuses finalized normalizer to derive nested retained support using only standard attention state. A length-normalized tolerance τ/Lq governs post-score allocation across queries, heads, layers and inputs without materializing explicit masks.
- Key results:
- At 8K, MALA achieves mean omitted mass of 0.0188% versus 0.0182% oracle under matched total post-score work
- On 128K tokens with tensor parallelism, MALA reduces forward and backward training latency by 2.2× and 3.0× and inference latency by 1.6× relative to FullAttn
- At 14B parameters, MALA reduces total training FLOPs by 2.5% during 4K pre-training and 23.1% during 32K long-context training while closely tracking FullAttn perplexity
- Reaches 89.67% accuracy on associative recall at 8K compared with 89.97% for FullAttn
- Why it matters / caveats: Allocating attention computation based on normalized contributions preserves operator fidelity and model quality while reducing attention cost, with savings concentrated in low-contribution regions rather than sacrificing core capabilities; efficiency gains scale with context length and long-context pretraining requirements.
TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces →
Agents can finish a task while behaving badly along the way, and fixed test suites miss the specific problems seen in real use. The authors built a system that mines recorded deployment sessions for a behaviour a developer names, then turns those moments into tests that ask a model what it would do next. Leading models failed most of these tests.
Technical breakdown
- Problem: Fixed benchmark suites do not test behaviors encountered in real deployment, and manually reviewing sessions at scale to construct targeted benchmarks is prohibitively expensive.
- Method: TraceDance constructs targeted benchmarks from deployment traces using Anchor-and-Confirm: programmable CPU-based anchors scan structured traces to retrieve candidates, then a Flash LLM confirms the requested behavior. For unknown behaviors, an Anchor Synthesis Loop synthesizes, validates and revises specifications. Decision-point continuation evaluates LLMs by regenerating next turns from recorded context before observed behaviors using behavior-specific rubrics without environment replay.
- Key results:
- From 252,557 sessions, produces 107 benchmarks with 4,125 instances, fulfilling 95.3% of build-target requests
- Human annotators confirm the requested behavior in 84% of sampled instances; automated grader agreement with human pass/fail judgments is comparable to inter-annotator agreement
- Nine frontier LLMs achieve mean pass rate of only 26.7%, with rates ranging from 22.9% to 33.5%
- Analysis reveals only 8.1% pass rate when a check is required before proceeding (versus 67.9% for well-formed tool calls), identifying concrete behavioral weaknesses
- Why it matters / caveats: Converting deployment problems into targeted benchmarks enables evaluation of specific behavioral weaknesses not captured by fixed suites, supporting recursive self-improvement loops; TraceDance targets behaviors with observable signals for programmable retrieval and cannot assess behaviors requiring full-session LLM or human inspection.
How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining →
Most systems that handle images bolt a separately trained vision component onto a language model. The authors ran matched training runs to compare that design with one that learns straight from raw pixels, tracking how each improves as computing power grows. The simpler pixel-based design lags at small scale but improves faster, and is projected to catch up within budgets labs already use.
Technical breakdown
- Problem: Most multimodal models use pretrained visual encoders providing strong visual priors, but encoder-free architectures have unclear scaling behavior relative to encoder-based approaches.
- Method: Controlled scaling study comparing encoder-free and encoder-based MLLMs sharing the same sparse MoE decoder ladder, data mixture, optimization setup and visual-token granularity. Fitting separate scaling laws for text and multimodal objectives using isoFLOP profiles and compute-optimal allocation laws (Mopt(C)∝Ca, Dopt(C)∝Cb with a+b=1) to quantify efficiency gap evolution with scale.
- Key results:
- Encoder-free training shifts compute-optimal allocation from a=0.464 to a=0.570 on multimodal objective, favoring larger models
- Text loss frontiers nearly overlap between architectures (exponent −0.0973 vs −0.0979); multimodal frontiers diverge (exponent −0.3778 vs −0.2998)
- Predicted crossover on multimodal objective at approximately 1022 FLOPs under compute-optimal allocation
- Decoder learns vision-specific adaptation: bidirectional attention among visual tokens increases, visual processing shifts to earlier layers, expert routing for visual tokens becomes concentrated
- Why it matters / caveats: Encoder-free models require more training compute within measured range but their faster-improving multimodal frontier predicts efficiency crossover within practical pretraining budgets (recent flagship models ~1025 FLOPs), positioning encoder-free as promising; crossover timing varies by topic, arriving earlier on language-heavy topics.
Improving Test-Time Scaling with Adaptive Looped Transformers →
Some language models save memory by running the same layers repeatedly, but spending extra passes on every word wastes effort since many words do not need them. The authors trained a model alongside a small decider that learns, word by word, whether another pass would actually improve the prediction. It got more accuracy out of the same computing budget, and kept improving as more passes were allowed.
Technical breakdown
- Problem: Fixed-depth looped transformers achieve steeper accuracy–compute scaling slopes than non-looped baselines but underperform at matched test-time compute because they spend extra iterations on every token, though many tokens gain little from looping.
- Method: TaH2 jointly post-trains backbone and iteration decider through lookahead depth supervision: depth labels are derived online from changes in prediction loss, and cost-sensitive loss supervises each depth decision. An updater provides input injection between iterations, and stopping probabilities weight predictions across executed depths. Shares parameters across iterations while enabling adaptive allocation of iterations per token.
- Key results:
- On AIME24–26 at 1.7B, TaH2 improves accuracy–compute slope by 53% over non-looped baseline (2.74 vs 1.79 pp per doubling of decoding FLOPs)
- Exceeds Standard's peak accuracy by 3.4 points at matched decoding FLOPs with evaluation extended to 32K tokens
- Gains persist across iteration depths: grows from +2.8 points at M=2 to +3.9 points at M=8 while baselines plateau
- Improvements generalize to 4B and 8B scales and beyond math to code, QA and tool use
- Why it matters / caveats: Online depth supervision enables models to allocate extra computation adaptively to tokens that benefit from looping, improving both test-time scaling efficiency and attainable accuracy; approach requires post-training looped models rather than training from scratch.
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL →
When training coding agents by rewarding whatever passes the tests, every passing attempt is treated as equally good, so sloppy fixes are reinforced alongside clean ones. The authors added a grader that compares all passing attempts for the same problem and ranks them on things like minimal, well-targeted changes, then shifts credit toward the better ones. Training became more stable and produced better code.
Technical breakdown
- Problem: GRPO assigns identical advantages to all test-passing trajectories within each rollout group, overlooking differences in implementation quality, code complexity, and adherence to task requirements.
- Method: Gagar (Groupwise Agentic Grading for Advantage Redistribution) uses dynamic sampling to retain mixed-outcome groups, then an agentic grader jointly inspects trajectories and ranks passing candidates across five dimensions: approach suitability, implementation precision, minimality of changes, avoidance of unintended side effects and codebase consistency. Sum-preserving redistribution downweights lower-quality trajectories and proportionally rescales advantages of all passing trajectories to restore their original sum.
- Key results:
- On DeepSWE v1.1 with MiMo-V2.6-Flash (310B total, 15B active), Gagar improves later-stage pass rates and training stability while reducing interaction turns and token usage
- After mixed-task RL, DeepSWE v1.1 avg@3 scores reach 67.9 for Flash and 71.9 for Pro (MiMo-V2.6-Pro 1.02T total, 42B active)
- Blinded rubric-based evaluation shows quality-aware advantage redistribution improves problem-solving behavior
- SFT-trained grader reduces average grading time from 2,000 s to 600 s per group versus Claude Opus 5
- Why it matters / caveats: Combining test-based verification with groupwise agentic grading improves code agent quality, stability and developer experience by reinforcing sound strategies and merge-ready implementations; credit redistribution preserves total positive advantage while shifting it toward higher-quality solutions.
Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning →
Models that both view and create images could in principle fix their own pictures, but a revision can only be judged after it is drawn. The authors trained such a model with reinforcement learning across whole loops of critique-and-redraw, so the written critique and the redrawing improve together. The model became much better at repairing its own images, and the gains carried over to untrained tests.
Technical breakdown
- Problem: Unified multimodal models can view and render images but do not reliably repair their own generations through self-reflection.
- Method: UMM-Reflection applies reinforcement learning to multi-round reflection trajectories within a single unified model. It uses group-relative advantage computed over sibling trajectories sharing the same initial image, and one trajectory-level advantage updates both reflection tokens and flow-based revisions.
- Key results:
- GenEval improves by 12.05 points over supervised fine-tuning, transferring to WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63).
- Conditional repair rate rises from 20.59% (SFT) to 64.94% (RL).
- Reflection SFT+RL outperforms direct RL on the generator alone.
- Why it matters / caveats: Demonstrates that combining reflection with RL makes reliable which revisions the model already produces; lacks mechanisms for generating entirely new repair strategies.
EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks →
Robots doing long tasks need to remember what they have seen and update it as things change, which current systems do poorly. The authors built a test suite of interactive episodes probing four specific memory weaknesses, plus a memory system that files experience separately as places, events and scenes. Their system helped every model they tried, though overall performance remains weak.
Technical breakdown
- Problem: Embodied agents struggle to maintain memory over long-horizon interaction, exhibiting four key deficiencies: weak fine-grained visual memory, unreliable dynamic world-state tracking, failure to record world state from interaction outcomes, and limited generalization from prior experience.
- Method: EMem-Bench comprises 2,554 episodes across four task families testing these memory capabilities. Embodied-Memorizer (EMem) organizes multimodal experience into spatial, event, and scene memories; an 8B policy learns to write, retrieve, and use these memories.
- Key results:
- Gemini-3-Flash reaches 64.2% average success rate, below frontier models.
- EMem improves Qwen3-VL-8B by 20.6 points, GPT-5.4-mini by 23.6 points.
- All three memory types contribute substantially; dynamic tracking depends most on spatial memory.
- Why it matters / caveats: Identifies fundamental bottlenecks in embodied memory but current approaches remain weak; improvements transfer modestly to general benchmarks.
CompoWorld: Compositional Environment Scaling for General Agents →
Real work often spans several separate online services, yet training environments for agents usually cover only one at a time. The authors assembled a library of reusable simulated services and automatically wired them together so tasks require passing information between them. An open model trained on these tasks improved broadly on agent tests and outperformed some leading commercial systems at multi-service workflows.
Technical breakdown
- Problem: Agents must coordinate across multiple services with cross-service dependencies, yet scaling typically generates isolated tasks within single environments.
- Method: CompoWorld composes 448 services (10,130 tools) into environments via random-walk dependency graphs. It uses a factorized hybrid control interface, typed Python schemas for service states, and Completion-Focused Rubric Reward that emphasizes criteria with lower pass rates, combined with GRPO.
- Key results:
- AutomationBench: 32.33% pass rate, exceeding GPT-5.4 (27.67%) and Claude Opus 4.6 (25.50%).
- Average 9.17-point gain over Qwen3.6-35B-A3B across eight benchmarks.
- Composed environments substantially outperform single-environment scaling at matched data budget.
- Why it matters / caveats: Demonstrates compositional environment scaling is effective but service quality and composability remain limiting factors; RL improvements are modest relative to SFT.
Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge →
Small reasoning models are cheap to run, but giving them more time to think does not help when the real gap is missing knowledge. The authors show extra self-checking mostly reshuffles answers the model could already reach, then trained small models to notice when they are stuck for lack of information and ask a stronger model a targeted question. This beat larger models at lower cost.
Technical breakdown
- Problem: Small reasoning models fail not only from weak reasoning but from knowledge bottlenecks where external information is required; merely prompting for reflection does not reliably help.
- Method: FlyBy diagnoses bottlenecks through counterfactual interventions, distinguishing execution bottlenecks (where reflection helps) from knowledge bottlenecks (where external information is needed). It trains a small model to selectively query stronger models at knowledge bottlenecks using cost-aware RL with multi-depth query actions.
- Key results:
- FlyBy-4B achieves 45.96% pass@8 on hard problems, surpassing Qwen3-14B (41.64%) at 2.7× lower cost.
- FlyBy-8B reaches 51.81% pass@8, outperforming Qwen3-14B by 10.17 points while remaining cheaper.
- Cost efficiency remains robust across a broad range of GPU and API pricing.
- Why it matters / caveats: Identifies fundamental computational bottleneck types but relies on external models for practical gains; availability and privacy concerns limit deployment scenarios.
Recursive Harness Distillation across Agents for Robot Manipulation →
Robot control models often need outside help to diagnose and correct failures, but calling a powerful model throughout a task is expensive. The authors had a strong model gather experience with a robot policy and write it up as a playbook of when and how to intervene, then refine that playbook using a cheaper model's results. The playbook lifted both models' success without retraining anything.
Technical breakdown
- Problem: Light agents cannot effectively operate vision-language-action policies without guidance, while relying on capable models is costly throughout execution.
- Method: Recursive Harness Distillation lets a strong agent distill interaction experience into a playbook (structured intervention guidance), then iteratively refines it using a light agent's execution feedback. The playbook specifies when to intervene via instruction editing, attention guidance, and action correction.
- Key results:
- Real-world manipulation: playbook improves light agent success from 37.3% to 64.0%.
- SimplerEnv Bridge: light agent with playbook reaches 66.7%, outperforming strong agent without playbook (54.2%), and strong agent with playbook reaches 79.2%.
- Instruction editing contributes most to gains among intervention sites.
- Why it matters / caveats: Demonstrates effective distillation of policy-specific knowledge but requires interaction with fixed VLAs and does not update model parameters.
Surprising Success, Repeated Failure: Entropy-Guided Credit Assignment for Exploration in LLM Reasoning →
When training models to reason with only right-or-wrong feedback, it is hard to know which steps deserve credit. The authors note that lucky successes at uncertain moments rarely repeat while confident mistakes keep recurring, so they reward uncertain steps in correct answers and penalise confident steps in wrong ones, while easing up on uncertain steps in failures. This improved reasoning and encouraged more varied attempts.
Technical breakdown
- Problem: Success under uncertainty is less repeatable while confident failures recur; standard entropy methods apply the same preference regardless of outcome, potentially suppressing exploration.
- Method: Entropic Advantage Policy Optimization (EAPO) treats success and failure asymmetrically: it reinforces high-entropy decisions in successful responses and penalizes low-entropy decisions in failed responses while preserving uncertain alternatives.
- Key results:
- Qwen3-4B-Base: mean 31.0% avg@32 accuracy, surpassing EntropyAdv (25.4%) and 80/20 (25.3%) by substantial margins.
- Qwen3-4B reasoning backbone reaches 72.4% avg@32, surpassing strongest baseline RLRT.
- EAPO generates higher answer diversity and more epistemic markers, broadening problem coverage.
- Why it matters / caveats: Simple entropy-based method that scales without auxiliary models but relies on policy entropy as proxy for actual reasoning bottlenecks.
WorldPlay2: Extending Real-Time Interactive World Models in Control and Horizon →
Interactive video worlds must react instantly to varied controls while staying consistent over long stretches, which is hard because the history grows quickly. The authors separate frame-by-frame movement controls from higher-level settings such as scene, character and events, and compress past frames into a compact memory shared by the fast and slow versions of the model. The result stays coherent over very long runs while remaining responsive.
Technical breakdown
- Problem: Interactive world models must handle heterogeneous controls spanning disparate semantic granularities while maintaining long-horizon consistency in real time.
- Method: WorldPlay2 couples a factorized hybrid control interface (frame-aligned actions for camera/locomotion, structured semantic control for scene/character/events) with compressed memory tokens and Stable Forcing distillation (few-step initialization, full-rollout replay, clip-wise score evaluation).
- Key results:
- Maintains long-horizon consistency over 1200+ frames while supporting versatile multi-turn interactions.
- Generalizes across different scenes and characters with superior performance compared to existing methods.
- Supports flexible interactive controls including navigation and complex semantic events.
- Why it matters / caveats: Addresses control heterogeneity but evaluation relies on qualitative examples; quantitative metrics on long-horizon consistency are limited.
DepthBench: Measuring How Residual Connections Enable More Computational Depth →
Adding more layers to a model does not reliably make it better, and it has been unclear whether recent fixes genuinely help. The authors ran matched training runs across many architectures, holding size and recipe fixed while shifting the balance between depth and width. Most normalisation-based fixes gave little benefit in deep, narrow shapes, while certain redesigns of the skip connections kept helping, pointing to those as the key factor.
Technical breakdown
- Problem: Increasing depth does not reliably improve Transformers under fixed parameter budgets; it remains unclear whether architectural improvements truly enable effective computational depth.
- Method: DepthBench systematically varies width-depth aspect ratios while holding model size and training fixed across 10 architectures. It controls learning rates per configuration and analyzes layer-level utilization.
- Key results:
- Pre-LN and normalization variants show little benefit or degradation at extreme depth (aspect ratio 9.1).
- HC and Full AttnRes consistently reduce loss even at aspect ratio 9.1 (70 layers, 640 hidden), reaching 400M parameters.
- Deeper architectures reduce hardware efficiency despite modeling gains.
- Why it matters / caveats: Identifies residual connection design as critical but does not clarify which designs are preferable for practical applications; systems-level efficiency trade-offs limit real-world adoption.
Diffusion Reward Models →
Models that score how good an answer is usually output a single number, even though people can reasonably judge the same answer in several different ways. The authors instead trained a scorer that produces a whole spread of plausible ratings, sampling several at once. It matched or beat conventional scorers, revealed disagreement that single numbers hide, and produced better-behaved models when used to guide training.
Technical breakdown
- Problem: Reward models collapse multimodal human preference into point estimates, losing distributional structure and uncertainty information.
- Method: DRM replaces scalar value heads with a Diffusion Transformer that models the conditional reward distribution p(r|x,y) without parametric assumptions. It supports both multi-attribute regression and pairwise preference data via masked MSE and Bradley-Terry objectives.
- Key results:
- Matches or surpasses baselines on five benchmarks (RewardBench v2, PPE, RMB, RM-Bench, JudgeBench) under matched backbone and training scale.
- Uncertainty-aware rejection and lower-confidence-bound aggregation improve reward-model decisions.
- RLHF downstream experiments show improved policy performance using DRM as training-time reward.
- Why it matters / caveats: Captures multimodal reward structure but inference requires sampling multiple times; practical benefit for RLHF training is modest.
Learning to Learn from Context: Synthetic Training from Perturbed Public Documents →
Language models are weak at working from a long document handed to them rather than from memorised knowledge, and hand-written training data is expensive. The authors lightly rewrote high-quality public documents, renaming entities and altering numbers so memorisation would not help, then automatically generated questions and kept only those genuinely requiring the document. A mid-sized model trained this way rivalled a far larger one.
Technical breakdown
- Problem: LLMs struggle with context learning despite exposure to long documents during pretraining; human annotation for long-document tasks is prohibitively expensive.
- Method: A synthesis pipeline perturbs public documents (entity renaming, numeric perturbation), generates questions requiring reasoning, answers with the document as context, and admits only samples genuinely dependent on the document via gap check (comparing rubric pass rates with and without context).
- Key results:
- SFT on 10k synthetic samples from 3.5k documents improves Qwen3.6-35B-A3B from 13.7% to 22.8% on CL-bench.
- SFT+RL reaches 24.6%, comparable with HY3 (23.5%) and Qwen3.8-2.4T (23.9%).
- Transfers broadly to long-context understanding, instruction following, and reasoning; knowledge benchmarks decline modestly.
- Why it matters / caveats: Provides scalable training pipeline without human annotation but relies on deep rewriting quality; teacher-dependent performance suggests sensitivity to instruction details.
Skill2Env: Capability-Oriented Environment Synthesis from Skills for General Agents →
Agents learn best from realistic, runnable practice tasks, but building those environments by hand does not scale. The authors start from written skill guides and use a catalogue of difficulty patterns to generate tasks together with their working environments and grading criteria, then progressively harden tasks that turn out too easy. Training on successful attempts improved an agent across a wide range of tests.
Technical breakdown
- Problem: Constructing executable tasks and their environments for agent post-training is difficult to scale beyond manual case-by-case development.
- Method: Skill2Env uses agent capability demands (environment understanding, planning, skill usage, long-horizon consistency, error recovery) to guide task and environment synthesis. It represents demands through reusable difficulty patterns and instantiates them into task blueprints specifying objectives, challenges, environment facts, information boundaries, and acceptance criteria. Iterative Task Hardening uses solver execution evidence to progressively strengthen task designs by identifying and addressing insufficiently challenging aspects.
- Key results:
- Generated 2,963 executable tasks from 3K+ curated skills
- Supervised fine-tuning on 1.5K trajectories (r > 0.9) yields consistent improvements across seven benchmarks
- Terminal-Bench 2.1: +13.5 points; SkillsBench: +14.34 points; SWE-bench Multilingual: +7.7 points
- Average improvement across benchmarks: +8.4 points (36.6% → 45.0%)
- Why it matters / caveats: Skill2Env demonstrates that capability-oriented task synthesis can scale agent post-training by bridging the gap between reusable skill knowledge and concrete, challenging executable tasks. Results show supervision transfers beyond synthesized environments to diverse downstream benchmarks.
QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents →
Some agent tasks run for hours and involve enormous amounts of text, which makes reinforcement learning wasteful because many machines sit idle waiting for the slowest run. The authors built a training system that shifts hardware between running and learning without interrupting jobs, and that prunes the repetitive branches these long runs produce. Training a very large model got substantially faster and more accurate.
Technical breakdown
- Problem: Online RL for extreme-scale agentic tasks (multi-hour executions, 700K tokens per rollout) faces severe GPU underutilization from rollout variance and prohibitive training overhead from complex trajectory branching.
- Method: QwenGyre introduces an elastic scheduler that dynamically reallocates GPUs between rollout and training without interrupting live executions, allowing newly available GPUs to join ongoing batches. A trajectory processor reconstructs branching execution histories into trajectory trees with shared prefixes, scores partial progress after timeouts, and deduplicates redundant paths by selecting bounded sets per execution with role-priority ranking.
- Key results:
- Qwen 3.8 2.4T on NL2RepoBench: 52.5% → 58.5% over 48 training steps
- 1.78× end-to-end speedup over Async baseline, 1.21× over Colocate
- Mean query duration: 2.96 hours; 9.51% of queries require ≥4 hours
- Why it matters / caveats: QwenGyre solves fundamental scheduling and trajectory-processing challenges for extreme-scale agent RL without requiring workspace checkpointing, enabling efficient distributed training of frontier-scale models on complex long-horizon software engineering tasks.
SpatialSpeak: QA-Native Reconstruction with Local and Global Context for Spatial Chain-of-Thought Reasoning →
Models that answer questions about space from several photos are usually trained only on final answers, so nothing teaches them the intermediate geometry. The authors first train the model to answer questions about where points and objects sit in three dimensions, using the same text format, then teach it to state those estimates and reason from them, checking the image again when unsure. Accuracy improved clearly.
Technical breakdown
- Problem: Vision-language models equipped with 3D geometric priors still lack explicit supervision for how to express and use intermediate geometric estimates in spatial reasoning.
- Method: SpatialSpeak employs two-stage training: Stage I (QA-Native Reconstruction Pretraining) jointly learns local point-wise 3D locations and global object-center detection in shared camera coordinates through text-based QA interface on Qwen3-VL-4B. Stage II (Spatial CoT with Visual Compensation) supervises explicit question-relevant geometric estimates, task-specific answer derivations, and reliability-conditioned refinement using visual evidence when geometric estimates are unreliable.
- Key results:
- ReVSI score of 62.8 (92.9% of baseline, +8.7 points over SpatialStack-4B at 54.1)
- QA-RP increases CoT-VC gain from 2.6 to 6.9 points
- State-of-the-art: VSI-Bench 63.3, SPAR-Bench 76.0
- Pointmap reconstruction on ScanNet: 8.9cm accuracy (vs 13.7cm CUT3R)
- Why it matters / caveats: SpatialSpeak demonstrates that reconstruction pretraining through QA creates a strong foundation for spatial reasoning, with ablations confirming both local geometry and global context are necessary components.
RoboFoundry: System-as-Policy Evolution for Self-Learning Embodied Agents →
Robot agents are usually improved one piece at a time, and experience from a run rarely turns into lasting change. The authors treat the whole supporting system - its memory files and its library of skills and recovery routines - as the thing being improved, updating it from execution records and keeping only changes that hold up on held-out tasks. It improved several models and transferred across different robots.
Technical breakdown
- Problem: Existing embodied agents optimize individual components (memory, skills, interaction harness) rather than treating the entire supporting system as a unified policy, and rarely convert execution experience into persistent, validated system changes.
- Method: RoboFoundry formulates the agent system itself (context management and hierarchical skills) as Self-Evolving System-as-Policy. It diagnoses capability gaps from execution traces, performs targeted system revisions at task-specific scope, and promotes validated improvements to general system through held-out evaluation across related tasks. Uses semantic binding layer to separate embodiment-invariant decisions from embodiment-specific execution.
- Key results:
- EmbodiedBench: GPT-5.5 improved 27.8% (56.9% → 72.7%), Qwen3.7-Plus improved 10.4% (60.0% → 70.3%)
- RoboMemArena: 53.5% TSR / 72.8% CSR (outperforms PrediMem by 213.3% relative on Transfer category)
- LIBERO-PRO: 91.9% average success (outperforms CaP-Agent0 by 679.7% on spatial perturbations)
- Real-world: nesting-doll completion 70% (vs 30% baseline), zero-shot towel-folding transfer across three fabric colors
- Why it matters / caveats: RoboFoundry advances embodied AI from code-as-policy to system-as-policy evolution, demonstrating autonomous improvement loops that transfer across heterogeneous robots through semantic decoupling.
CoWindow Attention: Full Causal Coverage Is a Collective Property →
Standard attention has every attention head look back over the entire history, which duplicates a lot of work. The authors give all heads a shared nearby window plus the very start of the text, then divide the remaining distant history among the heads so that together they still cover everything. Quality held up across model sizes while long-context training and generation ran several times faster.
Technical breakdown
- Problem: Full attention repeats complete causal history to every attention head, creating substantial redundant long-range computation and memory traffic despite advances in IO-efficient kernels.
- Method: CoWindow Attention (CoWA) distributes access to causal history across KV heads using position-defined window rule: all heads share near-diagonal and prefix-sink windows, while complementary long-range windows partition remaining history among heads. Collective coverage ensures every causal position is visible to at least one head without duplication. Single rule applies consistently during training, inference, and tensor-parallel deployment via global KV-head indexing.
- Key results:
- 8K associative recall: 89.73% accuracy (vs 89.97% FullAttn with matched window widths)
- At 128K tokens training: 7.4× forward latency, 8.6× backward latency reduction vs FullAttn
- Decoding: 3.0× latency reduction, 7.6× lower peak operator memory
- 14B model: 28.5% FLOP reduction during 32K-context training; perplexity matches FullAttn; RULER 32K: 89.13 (vs 89.42 FullAttn)
- Why it matters / caveats: CoWA shows full causal coverage can be a collective property across the head ensemble rather than duplicated in every head, enabling substantial efficiency gains while preserving model quality across 14B and 32B scales.
REALM: A Coarse-to-Fine Generative Framework for Embodied Reactive Listening →
Conversational robots need to react while listening, nodding or blinking at the right moment rather than mirroring the speaker instantly. The authors built a system that blends the listener's own recent movement with the speaker's audio using a built-in expectation of a short response delay, then adds small random facial variation on top of a smooth base motion. It outscored earlier methods and was preferred on a physical robot.
Technical breakdown
- Problem: Generating responsive listener facial motion requires balancing temporal response lags (listeners respond with delays) and behavioral continuity, while capturing both smooth overall trajectories and locally variable facial dynamics.
- Method: REALM combines Reactive Gated Speaker-Listener Fusion using a delay-centered attention prior (shifted ALiBi bias) that favors speaker representations near nominal response lag τ, with adaptive gating that learns relative weighting of listener history and speaker evidence. Coarse-to-Fine Stochastic Refinement predicts base motion trajectories, then augments only expression subspace with audio-conditioned stochastic residuals while preserving coarse pose parameters. Trained with supervised KD from vanilla model.
- Key results:
- ViCo Dtest: expression FID_Δfm of 3.91 (vs 8.71 ListenFormer baseline)
- Lowest L1 errors across both datasets for expression and pose
- L2L: expression L1 9.67 (vs 10.32 ListenFormer)
- Robot user study (N=25): REALM scores 4.4 naturalness (vs 3.6 ListenFormer), 4.0 preference (vs 3.6 baseline)
- Why it matters / caveats: REALM demonstrates that coarse motion and stochastic expression refinement are complementary mechanisms, with delay-aware fusion enabling better speaker-listener coordination and enabling physical robot deployment through deterministic retargeting.
AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research →
Search systems rank documents one by one for relevance, but a hard question needs a set of documents that complement each other without repeating. Scoring a whole set gives too blunt a signal to tell useful documents from passengers. The authors train a set-picking model with detailed criteria and give each attempt coaching matched to how good it was, improving answers while making fewer searches.
Technical breakdown
- Problem: Mainstream rerankers select individually relevant documents but not complementary sets, and set-level scalar rewards prevent fine-grained document-level credit assignment distinguishing contributors from free riders.
- Method: AdaTutoRank trains with hierarchical query-specific rubrics across document (relevance, authenticity, quality), set (complementarity, redundancy, conflict), and global levels (completeness, density, reachability). Stage I performs setwise SFT on silver labels from DeepSeek-V4-Pro. Stage II applies Adaptive Tutoring Optimization, matching hint form (rubrics alone, reflection, or sibling-set) to rollout quality, combining token-level distillation advantage with group-relative outcome advantage using GRPO.
- Key results:
- Answer-level overall: 45.28 (vs 43.84 RubricRanker, 39.08 initial retrieval)
- Setwise-level overall avg: 48.97 (vs 46.80 RubricRanker)
- First on 7 of 9 benchmarks; deep research gains larger (2.05 points) than RAG (0.80 points)
- Why it matters / caveats: AdaTutoRank demonstrates rubrics can supply both evaluation criteria and dense token-level supervision, with adaptive tutoring matching guidance to sample difficulty for more effective document set composition in complex retrieval scenarios.
WideSWE: Can Coding Agents Coordinate Changes Across Repositories? →
Coding agents are usually tested on a single codebase, but many real fixes and features require coordinated edits across several linked projects. The authors gathered real-world tasks of this kind from many software ecosystems and adapted the accompanying tests to allow different valid implementations. Even the best agent setup finished well under half the tasks, often missing some required projects or leaving edits incomplete.
Technical breakdown
- Problem: Coding agent benchmarks evaluate single-repository tasks, but real software ecosystems require coordinated changes across multiple repositories with interdependent interfaces and shared functionality.
- Method: WideSWE mines 1,729,171 merged PRs across 103 software ecosystems to identify 120 real-world cross-repository tasks (60 bug fixes, 60 features) requiring substantive changes in at least 2 repositories. Evaluates seven agent configurations (Codex CLI, Claude Code paired with various LLMs). Addresses test-requirement mismatches through two rule-guided manual review rules: relaxing implementation-specific constraints and removing unrequested functionality checks.
- Key results:
- Best configuration (Codex CLI + GPT-5.6-sol): 42.50% full task success
- 83.33% solve at least one target repository; 52-70% of failures have at least one repository passing F2P tests
- Post-edit failure dominates (59.42% of unresolved); incomplete scope identification (37.68%)
- Joint vs independent: 40.45% vs 35.96% on 89 identical-prompt tasks; joint uses 94.7 API calls vs 296.6 independent
- Why it matters / caveats: WideSWE reveals cross-repository coordination as a distinct challenge beyond single-repo issue resolution, with agents struggling to identify full modification scope or making inconsistent changes across related repositories.
Just MLPs: Efficient Visual State Reconstruction for Multimodal Language Models →
Images turn into long strings of tokens that make multimodal models slow; the usual fix throws some away, losing evidence that might matter later. The authors found the visual information reaching the text is concentrated in a few directions and is easy to predict, so they replace the costly layer-by-layer processing of image tokens with small lightweight predictors. This was faster and more accurate than discarding tokens.
Technical breakdown
- Problem: Long visual token sequences create substantial computational overhead in MLLMs because visual tokens undergo full Transformer evolution at every layer, even when relevant visual information occupies a low-dimensional subspace.
- Method: δ-Vision replaces repeated Transformer evolution of visual tokens with lightweight low-rank adapters that construct layer-wise visual memories. Embedding adapter predicts from initial visual embeddings; recurrent adapter progressively updates across layers using low-rank residual corrections (d → r → d bottleneck with SiLU). Freezes visual encoder and LLM backbone; trains only adapters with supervised KD where vanilla model serves as teacher providing token-level distributional supervision.
- Key results:
- Qwen3-VL-4B: average 74.4 (92.9% of vanilla 80.1, vs 65.3 best 5%-retention pruning, matching 20%-retention best of 74.2)
- Video-MME: 1.30× total speedup, 1.50× prefill speedup at 17.90% FLOPs of uncompressed model
- Cross-backbone: 94-96% vanilla performance on LLaVA-1.5-7B, Qwen3-VL-30B-A3B, Qwen3.5-4B
- Spectral analysis: visual attention outputs require only 38-103 directions (vs 212 available) to retain 95% energy
- Why it matters / caveats: δ-Vision shows that reducing visual computation costs along the hidden-channel dimension provides a complementary efficiency axis to token pruning, enabling faster inference while preserving all visual evidence for text retrieval across diverse model scales.
Change the Product, Keep the Parameters: Associative Algebra Layers for Transformers →
Rather than speeding up matrix multiplication, the authors ask whether a network can use a cheaper multiplication-like operation altogether, based on an algebraic structure with a sparser pattern of interactions. They trained two otherwise identical small language models differing only in this layer. The new version generated text somewhat faster but scored lower on every quality measure, so they present it as a feasibility check.
Technical breakdown
- Problem: Not stated.
- Method: Constructs family of associative unital algebras that replace ordinary matrix multiplication with sparser interaction table over same weight blocks, achieving quadratic arithmetic in matrix dimension at fixed physical block size. Uses graph construction where vertices represent diagonal feature slots and edges represent off-diagonal interactions; edge-edge products vanish. Provably optimal by Alder-Strassen bound. Implements as row-typed rectangular projections compatible with causal masking and KV-cached decoding; tests recursive Q_2^⊗2 in FFN layers of ~110M-parameter decoder-only LM.
- Key results:
- Kernel speedups: 3.04× (Qwen3-1.7B), 5.04× (Qwen3-8B), 5.06× (Qwen3-32B) on 2048-token batch
- End-to-end generation throughput: 6.2-7.8% gain across four domains (Code 6.8%, Math 7.0%, QA 7.8%, WMT 6.2%)
- Downstream quality lower: GSM8K 3.18 → 2.20 (−0.98 points), IFEval 34.5 → 31.1 (−3.4 points), MBPP 6.60 → 3.80 (−2.8 points)
- Bilinear rank R(Q_2) = 6 (optimal for 4 slots, 2 vertices); R(Q_q) = 2q² − q for q-slot law
- Why it matters / caveats: Algebraic products achieve hardware efficiency (6-7% generation speedup) but incur measurable quality loss at equal model scale; scaling laws for quality-preserving constructions at larger models remain unexplored, limiting practical deployment.
In-Flight KV Cache with Clean Anchors for Faster Autoregressive Video Diffusion →
Long AI-generated videos are made piece by piece, and earlier systems ran extra passes just to remember what came before. The authors reuse the memory each generation step already produces, add a few clean reference frames prepared in advance to keep appearance and motion steady, and run stages on separate GPUs. Videos of twenty seconds or more come out faster and stay consistent.
Technical breakdown
- Problem: Few-step autoregressive video diffusion requires cache-update-only forwards to memorize generated chunks, creating a bottleneck in generation speed.
- Method: FlashForward reuses in-flight KV cache from denoising forwards rather than executing separate cache-update-only passes. It employs two memories at different temporal scales: dense stage-matched history from current denoising outputs and sparse clean anchor KV prepared in advance for long-range structural guidance. The model uses a shared backbone with planner and renderer roles specialized with role embeddings and LoRA adapters, trained via packed supervised fine-tuning and self-rollout distillation.
- Key results:
- 1.16–1.69× faster than HiAR, 1.42–2.92× faster than Self-Forcing for 16 FPS videos 20+ seconds at 480p/720p with up to 4 GPUs
- 0.838 VBench Total score for 1.3B model at 480p, outperforming Self-Forcing (0.805) and HiAR (0.821)
- Stable generation quality at 20s, 35s, and 65s durations
- Why it matters / caveats: Eliminates cache-update-only bottleneck enabling practical long-form video generation while preserving quality. Two-sided temporal conditioning design generalizes to the autoregressive diffusion setting but depends on architectural choices specific to this paradigm.
Structured Residual Connectivity Matters for Diffusion Transformers →
Transformer-based image generators merge every earlier layer into one shared stream, unlike older designs with deliberate shortcuts. The authors let each layer actively pick which earlier representations to draw on, including its mirrored counterpart. Training converges in far fewer steps and image quality improves clearly, with a negligible increase in model size.
Technical breakdown
- Problem: Diffusion Transformers use uniform residual streams that conflate all preceding layers, unlike U-Nets with structured skip connections, limiting the ability to adaptively route information across depth.
- Method: Proposes adaptive residual routing that collects encoder representations as differentiable sources and applies softmax-based selection in decoder layers. Mirror routing pairs each decoder layer with its symmetric encoder counterpart, preserving sparse structured pathways while maintaining single-scale architecture. Residual routing operators precede self-attention and MLP sublayers, using normalized representations and learned routing vectors.
- Key results:
- 1.73× fewer training iterations to reach baseline FID
- Improved REPA-XL/2 from 5.9 to 4.34 FID without guidance, reaching 1.39 FID with classifier-free guidance
- FID drop of 0.96 in 0.28M additional training steps with less than 0.1% parameter overhead
- Why it matters / caveats: Demonstrates that adaptive structured connectivity is critical for diffusion transformers, extending U-Net principles to scalable transformer backbones. Improvements concentrated on single-scale homogeneous DiT architectures; generalization to hierarchical designs unclear.
Precise Editing and Flexible Referencing for Interactable Worlds →
Video world models let users wander through generated scenes but give little control over changing what is already there. The authors extend one so users can stream in editing instructions and reference pictures while generation continues, keeping a bounded memory of past frames for long runs, and they build matching training data and a test suite. It leads on editing while keeping navigation ability.
Technical breakdown
- Problem: Existing video world models emphasize navigation and text-driven events but lack precise control for editing existing content and flexible incorporation of reference images during real-time interaction.
- Method: EditWorld extends LingBot-World-Base with Gated Causal Attention to support streaming editing instructions and reference images during autoregressive generation. Sparse Context mechanism maintains bounded historical context for long-horizon inference. Joint autoregressive and bidirectional training with annealed self-resampling. Custom data pipeline synthesizes global (weather, style, time) and local (addition, removal, replacement) edits with detailed annotations covering scene descriptions, editing instructions, and reference images.
- Key results:
- 73.8 overall score and 80.0 editing score on WBench-Editing benchmark, substantially outperforming existing world models
- Maintains competitive performance on original WBench for navigation capabilities
- Supports 240–480 frame videos with 1–3 streaming editing instructions incorporating reference images
- Why it matters / caveats: Bridges gap from exploratory world models to fully editable ones, enabling users to modify specific world content interactively. Requires extensive data synthesis pipeline; generalization beyond supported editing operations remains to be explored.
FlowTool: Controlling Tool Parameter in Image Retouching via Flow Matching →
Automatic photo retouching is usually handled by language models that spell out tool settings one token at a time, which handles numbers poorly and costs a lot. The authors instead generate the whole set of continuous settings in one shot from noise, guided by the picture and the request, then fine-tune against how good the finished edit looks. It is dramatically faster, lighter on memory, and edits better.
Technical breakdown
- Problem: MLLM-based image retouching formulates editing as autoregressive token generation, causing poor numerical understanding of continuous parameters, cascading errors from early mistakes, and excessive computational cost.
- Method: Reformulates tool-based editing as conditional flow matching over continuous parameter space. Combines VLM backbone for multimodal understanding with DiT-based tool parameter generator producing editing plans through rectified flow. Separate tool-presence head predicts discrete tool selection. Two-stage supervised flow-matching curriculum followed by reward-based post-training (DiffusionNFT) directly optimizes rendered edit quality using image rendering rewards.
- Key results:
- 50× reduction in inference latency and 2× memory reduction versus autoregressive baselines
- Outperforms MLLM editing agents on reference-based metrics across MMArt-Bench, ArtEdit-Bench, and MIT-Adobe5K
- Competitive with proprietary models under reference-free evaluation of semantic consistency and perceptual quality
- Why it matters / caveats: Demonstrates fundamental mismatch between autoregressive language generation and continuous editing parameter prediction. Direct parameter generation dramatically improves efficiency but requires well-designed reward signals for quality optimization.
SolveEdit: Benchmarking Visual Problem Solving in Generative Models →
Most image-editing tests check whether a model follows a spelled-out instruction, not whether it can work out what needs to change to reach a goal. The authors build a test set where each case states what must end up true and what must stay untouched, allowing several valid answers. Current models solve only about half; adding a step that plans the change before editing helps every generator tried.
Technical breakdown
- Problem: Existing editing benchmarks conflate instruction following, scene understanding, and rule inference; cannot distinguish information source for valid transitions or accommodate multiple valid solutions.
- Method: Introduces SolveEdit benchmark with 2,728 cases across 10 domains organized by information source (Instruction-Specified, State-Dependent, Rule-Dependent). Defines atomic transition contracts specifying required and protected conditions enabling multi-solution evaluation. SolveScore metric measures required completion with bounded deduction for collateral changes plus six diagnostic subscores. SolveEdit-Plan uses VLMs with schema-constrained parsing to instantiate missing variables and ground editing instructions in visual evidence before generation.
- Key results:
- Strongest evaluated model achieves 57.0% SolveScore on benchmark
- SolveEdit-Plan improves GPT-Image-2 from 57.0% to 71.6% SolveScore, 6.3 point advantage over Generic Vision Rewrite baseline
- Improvements transfer to open-source editor and video generator without retraining
- Why it matters / caveats: Separates visual reasoning dimensions enabling targeted diagnosis of model failures. SolveEdit-Plan improvements are benchmark-specific; universality of the two-stage planning approach across diverse editing tasks remains unclear.
SciGen-Verifier: A Multimodal Reasoner for Explainable Verification in Scientific Image Generation →
Scientific answers are often drawings, and checking whether a generated circuit or plot is correct needs subject knowledge, not just an eye for artifacts. The authors build a test set requiring a verdict, an explanation, and a correction, then train a checker using staged reinforcement learning. It rivals much larger commercial systems and can guide repeated fixes.
Technical breakdown
- Problem: Existing verifiers target natural images and compress judgements into scalar scores; scientific image generation errors require domain knowledge and structural reasoning with explainable feedback for correction.
- Method: Constructs SciGen-Verify benchmark with 1,350 samples spanning instruction following, multidisciplinary reasoning, and world knowledge, featuring three-tier hierarchical protocol: binary judgement, explanation, and editing instruction. Develops SciGen-Verifier-8B via cold-start supervised fine-tuning on 48,311 curated instances followed by curriculum-based two-stage reinforcement learning with rubric-guided process rewards and outcome rewards.
- Key results:
- Matches much larger proprietary models on SciGen-Verify benchmark
- Achieves 94% human agreement on verification judgements with detailed explanations
- Serves as online critic for iterative image rectification in educational scenarios
- Why it matters / caveats: First benchmark and verifier for scientific image generation with explainable feedback. Limited to 1,350 samples focused on five STEM disciplines; scalability to additional domains and cross-disciplinary reasoning not demonstrated.
Program-Verified Self-Evolution for Vision-Language Models →
Vision-language models that train on their own self-written questions must guess the answers, and voting produces wrong labels surprisingly often. Instead, the authors have the model convert each image into a structured record, let fixed programs write questions and compute answers from it, and ask the model only to confirm small individual facts. Labels become much more accurate and gains keep growing over training rounds.
Technical breakdown
- Problem: Self-evolving VLMs use majority voting or model judges to label self-generated questions, but 18–24% of labels are wrong; pseudo-label accuracy degrades over training rounds (e.g., 72%→61% for VisPlay, 79%→63% for R-Zero).
- Method: VQS replaces voting with deterministic program execution over structured parses. Schema-constrained parser converts images to navigable structures (scene graphs for photos, tables for charts, node-edge sets for diagrams). Fixed templates traverse structures to generate questions and compute answers deterministically. Visual checker verifies atomic claims claim-by-claim. Label-free parser training selects best parse via verified-claim fraction. Single model serves parser, solver, and checker roles trained jointly with GRPO on verified QA pairs.
- Key results:
- 94% human-rated answer correctness versus 76% for majority voting
- Up to 3.18 point improvement on Qwen3-VL at 2B, 4B, 8B scales across 10 benchmarks
- Gains persist and grow over 3 training rounds, reaching 3.84 points at 2B scale
- Why it matters / caveats: Demonstrates deterministic programs with visual verification outperform probabilistic methods for pseudo-label generation. Approach limited to images supporting structured parsing (photos, charts, diagrams); infographics and free-form scenes less well-handled.
Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It →
In reinforcement learning for language models, the system that writes sample answers and the one that computes updates disagree slightly about token probabilities. The authors describe this gap precisely and correct it with a cap that tightens as the model grows more confident, trading a small controlled bias for far steadier updates. It gave the best average across the models tested.
Technical breakdown
- Problem: RL post-training decouples rollout generation (inference engine) and gradient computation (training engine), causing different token probability assignments; exact importance sampling correction yields uncontrolled variance, especially for mixture-of-experts models.
- Method: Calibrated Importance Sampling (CIS) decomposes mismatch into token confidence and log-odds displacement (εt). Empirically measures that displacement distribution is approximately invariant to confidence across six orders of magnitude, enabling single-constant truncation of anomalous displacements. Truncates exp(εt) at 1+λ, mapping back to confidence-dependent importance-weight cap tightening with higher token confidence. Theoretical analysis bounds second-moment term of exact correction.
- Key results:
- Highest five-benchmark average on all three mixture-of-experts models tested
- Bounds second moment at cost of controlled bias, improving upon unbiased exact correction with explosive variance
- Lower overall truncation bias than fixed-cap truncated importance sampling, particularly for low-confidence tokens
- Why it matters / caveats: Provides principled mismatch correction grounded in statistical structure rather than heuristic thresholds. Evaluation limited to mathematical reasoning benchmarks; applicability to other downstream tasks and model scales requires investigation.
Imprint Reader: From Weight-Update Readout to Behavioral Intervention →
Training leaves traces in a model's weights, but models cannot say what those changes taught them. The authors train a reader that takes a frozen weight update and describes it in words, with control cases discouraging invented claims. Descriptions are still often wrong, yet the reader's gradients can steer edits that raise refusal of harmful requests and improve tool use.
Technical breakdown
- Problem: Language models cannot decode parameter-level traces of their learning into explicit accounts of acquired knowledge or behavioral changes, limiting self-improvement capabilities.
- Method: Trains Imprint Reader with Semantic Mount-and-Read Tuning (SMaRT) to describe frozen weight updates in natural language. Reader mounts update onto its own parameters and elicits description via anchor-free meta-query semantically independent of target. Includes no-change and random-perturbation control episodes. Reader shares parameter coordinates with parent model enabling differentiable intervention via MetaEdit through coordinate-aligned gradients.
- Key results:
- Pass@100: 2% for knowledge readout, 16% for behavior under judge-based evaluation on unseen updates
- Reader-guided 0.5% pruning raises harmful-prompt refusal from 57.9% to 64.1% under safety-maintenance target
- MetaEdit increases backtracking and sub-goal expressions in mathematical reasoning; raises BFCL Overall from 41.69% to 44.60% without target-task training data
- Why it matters / caveats: Enables models to introspect their own learning and use introspection for intervention without external feedback. Readout reliability remains limited; intervention results show partial correctness suffices for useful behavioral effects.
TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining →
Comparisons of video pretraining methods change many things at once, so it is unclear what actually teaches a model about motion. The authors run a controlled grid of architectures and training goals, then propose a design pairing a per-frame image path with a small temporal module that rebuilds frames from one anchor frame plus motion cues. It leads on motion-heavy tasks with much less computation; appearance-heavy tasks favour others.
Technical breakdown
- Problem: Video SSL methods confound architecture, objective, data exposure, and schedule; unclear which combinations produce motion-prioritized representations improving on frame-to-frame change while preserving appearance.
- Method: Conducts controlled 4×6=24 architecture-objective study at 170M–190M encoder scale on 1.7M clips for 8 epochs. Proposes TT-VidT combining DINOv3-initialized ViT-B/16 spatial path with compact Temporal Transfer Layer (8 motion tokens per frame, 12 layers, block-causal attention). Diff Compression objective reconstructs frames from first-frame appearance anchor and per-frame motion tokens, creating bottleneck forcing motion information into tokens.
- Key results:
- Jester: 73.25, Something-Something V2: 25.92, ARID: 37.47, Diving48: 18.63 on motion-heavy benchmarks
- Only method leading all four motion-heavy benchmarks simultaneously with 54–121% improvement over strongest non-TT baseline
- 48–55% fewer encoder FLOPs than DisMo, VideoMAE, and V-JEPA 2; motion-inversion probe shows 99% answer change versus 9–43% for baselines
- Why it matters / caveats: Isolates motion-centric pretraining through controlled study, demonstrating structured bottleneck more effective than dense routing. Boundary cases: appearance-heavy benchmarks (HMDB51, IARD, EPIC-Kitchens) favor broader baselines, limiting generality to motion-sensitive domains.
InfiniHand: Streaming World-Space Hand Motion Estimation from Egocentric Video →
Working out how hands move through the world from head-mounted video usually means chaining a hand tracker to a separate camera-tracking system, which compounds errors and runs slowly. The authors train one streaming model that estimates hand shape, hand position and camera motion together, using thousands of hours of first-person footage. It is more accurate, drifts less over time, and runs about twice as fast.
Technical breakdown
- Problem: Recovering 3D hand motion in world-space from egocentric video requires jointly estimating MANO hand parameters, camera trajectories, and hand localization while minimizing error accumulation across separate cascaded modules.
- Method: InfiniHand is a streaming feed-forward framework built on a 3D foundation model backbone that jointly predicts hand locations, MANO parameters, and camera poses in a unified architecture. Training proceeds in two stages: Stage I learns camera-space hand priors from geometric and appearance features, then Stage II jointly estimates hands and camera trajectories with streaming spatiotemporal memory via Geometric Context Attention. Sparse bundle adjustment refines camera poses, and the approach aggregates 5,000 hours of egocentric data across nine public datasets.
- Key results:
- 21.4% reduction in ARCTIC PA-p compared to ViDiHand (9.82 to 7.72 mm)
- 11.19 FPS throughput, exceeding HaWoR by 2.04× speedup
- Strong generalization to in-the-wild videos with 57.7% improvement on EgoDex
- Why it matters / caveats: Enables scalable 3D annotation of egocentric video for embodied learning and robot policy training. Limitations remain in metric scale recovery and handling rapid viewpoint changes with severe occlusions in unconstrained scenarios.
Residual Transferability in Neural Image Watermarking →
Neural image watermarks can be forged by lifting the faint watermark pattern off a released image and pasting it onto unrelated pictures. The authors define a measure of how well such patterns transfer, show that architecture rather than training choices drives it, and identify two design features that tie a watermark to its own picture. They also offer an add-on that does this for existing systems.
Technical breakdown
- Problem: Neural image watermarks can be forged by extracting watermark-bearing residuals from legitimate images and transferring them to unrelated content, but the mechanisms enabling this cross-image transfer remain poorly understood.
- Method: Formalizes residual transferability (RT) as a metric quantifying watermark decodability after transfer across images using ground-truth residuals. Analyzes five representative methods across architectures and training objectives to identify that model architecture—not training variations—governs RT. Identifies two architectural mechanisms reducing transferability: Broadcast-GAP (spatial image-message interactions) and content-adaptive attention. Proposes CoverLock, a plug-and-play wrapper using frozen DINOv2 features with orthogonal projection to bind watermarks to image content without architectural redesign.
- Key results:
- CoverLock reduces average forgery attack success rate to below 0.25% across three high-RT systems
- Controlled interventions show RT can shift from 0.99 (high) to 0.10 (low) via architectural changes
- Table 2: CoverLock achieves 0.00% ASR on CIN across datasets while maintaining 92.50% TPR
- Why it matters / caveats: Provides principled design guidance for watermarking security and practical defense for existing systems without retraining. Analysis shows cover dependence is fundamental to forgery resistance.
BaRe-Mem: Bayesian Reliability Memory for Robust and Adaptive Agent Consultation →
When AI agents ask other agents for help, bad advice can leave them worse off than working alone. The authors keep a running estimate of each advisor's reliability for the situation at hand, drawn from the asking model's own internal state, and use it to weight advice or skip consultation. It holds up better than debate or majority voting when advisors mislead, and also helps pick team members.
Technical breakdown
- Problem: In multi-agent systems, advisor capabilities vary across tasks and contexts, making it difficult to determine when to trust external advice versus relying on autonomous reasoning without performance degradation.
- Method: BaRe-Mem maintains online Bayesian reliability estimates conditioned on the central model's belief representations, updating via Kalman gain from verified correctness outcomes. Reliability estimates modulate attention weights to advisor responses, while a decision mechanism compares estimated consultation versus autonomous abilities to determine consultation necessity. The approach uses belief representations as anchors for contextual reliability without requiring explicit feature alignment.
- Key results:
- Remains above autonomous baseline across all misleading information ratios (0-100%) on capability-challenging tasks
- 54.3% task completion on MuSiQue (vs 38.9% for success-count routing and 25.5% for random)
- Learns useful reliability estimates from sparse verified feedback (under 1% of interactions)
- Why it matters / caveats: Enables adaptive multi-agent teams to robustly incorporate external information while maintaining performance by knowing when autonomous reasoning is preferable. Sparse feedback requirement enables practical deployment.
Fewer Tokens, More Self-Teaching: On-Policy Self-Distillation for Extreme Visual Token Reduction →
Cutting the number of image pieces a multimodal model reads speeds it up but wrecks accuracy when the cut is severe. The authors let the trimmed model generate its own answers and learn from a full-view copy of itself scoring those same answers, lowering the budget gradually during training. Most of the lost ability returns while memory and setup computation drop sharply.
Technical breakdown
- Problem: Extreme visual token reduction (5% retention) in multimodal LLMs causes severe performance degradation because the compressed model generates different trajectories than its full-token counterpart, creating a distribution mismatch during training and inference.
- Method: LT-OPD (Low-Token On-Policy Distillation) addresses distribution mismatch by training the low-token student on its own generated trajectories supervised by a frozen full-token teacher using Jensen-Shannon divergence. A three-stage budget-level curriculum progressively reduces token budget during training: an initial phase at higher budget (2α+1)β, a cosine decay transition phase, then final training at target budget β. This gradual shift enables stable on-policy learning under extreme compression.
- Key results:
- 82.3% average retained performance on Qwen3.5-4B at 5% tokens (vs 68.6% without training)
- 85.2% KV-cache reduction and 85.4% prefill FLOPs reduction with no inference overhead
- Outperforms RL baselines (GRPO, GSPO, DAPO) by 4+ percentage points on recovery metric
- Why it matters / caveats: Demonstrates on-policy learning can substantially recover capabilities lost to extreme compression. Extends consistently across model scales (4B to 9B) and architectures (Qwen, GLM, LLaVA).
Hard Vision, Easy Vision: What GPT-6 Astra Reveals Across Computer Vision →
General-purpose AI systems now attempt tasks once reserved for dedicated vision software, but nobody had mapped how far that reach goes. The authors test several frontier systems across dozens of vision capabilities, comparing them with specialist models and people. Tasks involving meaning, reasoning and finding objects are close to specialist level, while precise measurement, faithful reconstruction, steady frame-by-frame prediction and niche expertise still lag.
Technical breakdown
- Problem: As frontier general-purpose systems rapidly expand beyond language into computer vision, the boundary between capabilities accessible through generalist interfaces versus those requiring specialist models remains unclear.
- Method: Systematic benchmark evaluating GPT-6 Astra alongside five other frontier systems (Fable 5, Kimi K3, Gemini 3.1 Pro, Qwen 3.8-Max, Muse Spark 1.3) across 34 capabilities spanning nine vision areas using 55 benchmarks. Compares against specialist models and human performance across recognition, reasoning, localization, 3D perception, video understanding, generation, robotics, and expert domains.
- Key results:
- Astra exceeds specialist performance on 2D object detection (+10.7 points F1) and 3D visual grounding (+10.7)
- Depth estimation lags with 32.4% higher ATE on Mip-NeRF360 vs specialist (0.071 vs 0.105)
- Video segmentation remains challenging with 7.9-point gap versus specialist baseline
- Why it matters / caveats: Maps evolving landscape showing generalists excel at semantic interpretation and object-centric tasks while metric geometry, temporal consistency, and fine-grained domain expertise remain frontier challenges. Specialist tools remain valuable for reconstruction-heavy and specialized vision tasks.
Nereus: Adaptive Parallelism for LLM Post-Training →
Reinforcement learning runs for large models spread several models across a GPU cluster using a layout chosen at the start, which can turn slow or unworkable as conditions change. The authors built a runtime that judges whether switching layouts is worth the disruption, then moves each model's distributed state in a safe order. Throughput improves several times over and switching costs almost no runtime.
Technical breakdown
- Problem: RL post-training on GPU clusters faces dynamic drift from resource volatility, workload changes (sequence length growth), and shifting hardware efficiency, but systems fix execution plans at startup, leading to inefficiency or infeasibility as conditions evolve.
- Method: Nereus is a cost-aware runtime that adapts execution plans and distributed state online via three components: (1) a low-overhead control policy using event-driven triggers and online-calibrated cost models to admit profitable transitions; (2) Elastic Model Units (EMUs) representing model-stage replicas with TP/PP encapsulated internally and DP exposed for scaling; (3) a transition engine building a global DAG of Split/Merge/Extend/Destroy primitives respecting cross-model-stage dependencies to safely move state over GPU-direct links.
- Key results:
- 2.14–7.27× throughput improvement over OpenRLHF (median 3.99×), 1.10–1.47× over Verl
- 27.7% latency reduction with online TP/PP adaptation on traced workload
- Six transitions on 1,024 GPUs consume only 0.079% of total run time
- Why it matters / caveats: Enables long-running distributed training to adapt to changing conditions while minimizing transition overhead. Integrates with vLLM, DeepSpeed, and Megatron-LM using standard NCCL/RCCL collectives.
An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning →
Teaching a small model by having a larger one grade its own attempts works but needs a constant supply of fresh samples. The authors recast this as reinforcement learning, adding deliberate exploration and reuse of previously collected attempts. It scores better on maths reasoning, keeps a wider range of solutions, and matches the usual approach using a quarter of the samples.
Technical breakdown
- Problem: On-policy distillation (OPD) for LLM reasoning requires continuous fresh rollouts without sufficient exploration or off-policy data reuse, limiting sample efficiency while maintaining policy diversity.
- Method: LSPD (Least-Square Policy Distillation) reinterprets reverse-KL distillation as KL-regularized policy optimization with teacher-induced log-ratio rewards. Formulates optimistic least-squares objective matching student and teacher log-probabilities with explicit entropy regularization for exploration. LSPD-RB extends to off-policy learning by storing trajectories in a replay buffer, enabling multiple updates per rollout batch while preserving historical data.
- Key results:
- +1.59 Avg@16 improvement over OPD on six mathematical reasoning benchmarks
- LSPD-RB reaches saturated performance in 10 rollout batches versus 40+ for standard OPD
- Stronger Pass@k scaling (k=1 to 64), with LSPD achieving 89.16 Pass@64 on AMC23 versus 87.95 for best baseline
- Why it matters / caveats: Demonstrates value-based RL principles applied to policy distillation improve both sample efficiency and solution coverage. Sharp Õ(log K) regret guarantee under online exploration provides theoretical grounding.
GeoVerse: World-Consistent Novel View Synthesis in Geometric Latent Space →
Generating new camera views from a handful of photos means both reproducing what was seen and inventing what was not, without contradictions between views. The authors run the generation inside the geometric representation of a 3D foundation model while borrowing appearance knowledge from a video generator, and keep a running memory of the scene to anchor later views. Quality and geometric consistency improve, and it runs far faster.
Technical breakdown
- Problem: Novel view synthesis must balance faithful reconstruction of observed regions with plausible completion of unseen areas while maintaining world consistency across viewpoints and long-sequence generation.
- Method: GeoVerse performs diffusion generation within geometric latent space of DA3 foundation model while injecting appearance priors from Wan2.2 VACE via ControlNet-style adapter. Global spatial memory maintains colored point cloud that accumulates observed and synthesized content across rounds, reprojecting memory to provide target-aligned RGB-D guidance. Four-step reflow distillation reduces denoising from 49 to 4 steps (3 Level-1 + 1 cascade Level-0), trained on 15 diverse 3D datasets.
- Key results:
- 2.23 dB PSNR improvement over GLD on DL3DV, 2.45 dB on Mip-NeRF360
- 32.4% ATE improvement on Mip-NeRF360 (0.105 to 0.071)
- 18× faster inference: 9.24 seconds versus 169.29 for GLD
- Why it matters / caveats: Combines geometric consistency from foundation models with generative quality from video priors for high-fidelity multi-view synthesis at practical speed. Limitations in scale recovery remain due to monocular depth ambiguity.
RenderRank: Learning to Rerank Text with Compressed Visual Tokens →
Search systems rerank candidate documents by feeding their text to a model, which is costly for long documents. The authors instead render the text as a picture, which the model encodes into far fewer pieces, training it to copy a text-based grader and then sharpen its ordering. It uses fewer inputs, ranks as well or better, and runs faster.
Technical breakdown
- Problem: Document reranking requires scoring relevance for potentially long documents, but longer sequences increase computational cost quadratically; representing document content with fewer tokens while preserving relevance information is challenging.
- Method: RenderRank renders document text as images (Roboto 12pt, 1.0 line spacing) and encodes to compressed visual tokens via Qwen3-VL vision encoder. Two-stage training: Cross-Modal Relevance Distillation (MSE loss) aligns visual inputs with text-based teacher (Qwen3-Reranker-4B) scores; Query-Local Relevance Discrimination (InfoNCE) refines relative ranking within query candidate sets. Visual embeddings precomputed offline.
- Key results:
- 55.96 NDCG@10 on 11 BEIR datasets with average 290.07 input tokens (16.5-35.5% reduction)
- 88.27 NDCG@10 on long-document datasets with ~50% token reduction (4.2K vs 8.8K)
- 76.83 PPS throughput, approximately 2× faster than comparable models
- Why it matters / caveats: Demonstrates visual compression of text enables efficient document ranking while maintaining or improving quality. Ablation shows both training stages critical for performance.
Rethinking Automated Voice Similarity by Shifting from EER to Embedding Geometry →
Speech researchers use speaker-verification models to judge whether two voices sound alike, assuming better verification means better agreement with listeners. Comparing many models against human ratings, the authors find verification accuracy does not track human judgment at all; what matters is the training objective and how many directions the model spreads voices across. Forcing a more compact representation sharply improved agreement with listeners.
Technical breakdown
- Problem: Speaker verification (SV) models are assumed to better capture voice nuances with improved EER, yet whether verification accuracy (EER) correlates with human voice similarity perception remains empirically unexamined.
- Method: Measures perceptual alignment (ρ_align) as Spearman rank correlation between embedding cosine similarities and human listener ratings on VoxSim dataset (46,348 utterances, 13 listeners). Analyzes effective dimensionality (d_eff) via participation ratio of PCA eigenvalue spectrum on speaker centroids, quantifying directional spread of embeddings across five model architectures and seven training objectives.
- Key results:
- EER-perceptual alignment Spearman correlation: +0.07 (not significant, p=0.71)
- d_eff-perceptual alignment Spearman correlation: -0.95 (highly significant, p<0.001)
- Constraining embedding dimension from 192 to 3 raises ρ_align from 0.08 to 0.74 for AM-Softmax
- Why it matters / caveats: Challenges implicit community assumption that better verification accuracy improves voice similarity estimation. Low-dimensional representations naturally align with human perception of voice, providing geometric criterion for evaluating speaker embeddings in TTS/VC applications.
Relic: From Multi-Agent Collaboration to Persistent Organizational Capability →
Teams of AI agents settle coordination disputes in conversation, but the lesson vanishes when members are swapped out. The authors turn repeated failures into rules the agents propose, vote on, and bind to the system so they actually govern later work. Task delivery improves, and newcomers behave far more correctly than when handed the same rules as plain text.
Technical breakdown
- Problem: Multi-agent teams lose learned coordination practices when members are replaced, requiring repeated negotiation of working agreements.
- Method: Relic adds a governed protocol lifecycle where agents reflect on visible work failures, propose executable rules, and vote on adoption. Adopted protocols bind triggers, responsibilities, evidence requirements, and execution consequences to a runtime decision layer via weighted utility scoring and stochastic action selection.
- Key results:
- Complete-contract delivery improves from 14.06% to 19.76% (+5.71 pp) over matched structured team without protocols across 360 runs on ten workloads
- Fresh-member behavioral correctness reaches 41.2% with executable bindings vs 34.6% with text-only rules
- External benchmark: 367/477 (76.9%) on full CooperBench after excluding broken pairs; exceeds Solo (29/48 vs 26/48)
- Why it matters / caveats: Demonstrates how organizational state persists across agent turnover through executable protocols. Evaluation uses three LLM models but focuses on software engineering tasks; generalization to other domains unclear.
Measuring Collapse and Correction in Homogeneous-Panel LLM Debate →
Judging multi-agent debate by final accuracy hides two opposite effects: discussion can rescue a wrong majority or wreck a right one. The authors log every run as a transition record separating collapses from corrections, and find that a screen predicting collapse risk, used to freeze debate, prevents some collapses but loses more corrections. Most collapses begin in the very first round.
Technical breakdown
- Problem: Multi-agent debate evaluation conflates opposing mechanisms—debate can either collapse a correct majority or correct a wrong one—making headline accuracy insufficient for assessing debate quality.
- Method: Proposes auditable transition-table protocol tracking preserved/collapsed/corrected/unrepaired outcomes. Pre-debate 8-probe factorial (2 social × 4 argument strengths) measures revisability; round-level traces identify where collapses begin; signed-replay scoring weighs prevented collapses vs lost corrections with user-specified weights.
- Key results:
- 253 collapses across 6,925 MMLU-Pro debates; 58.9% begin in Round 1
- Pre-debate probe has family-level Spearman ρ=0.893 with conditional collapse (p=0.0123)
- Probe-gated freeze prevents 29 collapses but loses 108 corrections under equal weights (net utility −79)
- Why it matters / caveats: Addresses critical gap in debate evaluation methodology with reusable audit artifacts. Limited to homogeneous 3-agent closed-book MCQ debate; MCQ format may not transfer to open-ended or heterogeneous settings.
WaveFront Decoding: Parallelized Self-Speculative Decoding for Looped Language Models →
Some language models save parameters by applying the same block many times per word, which makes them slow to write text. The authors have the model draft upcoming words using shallow passes while simultaneously deepening earlier ones in the same batched computation, correcting drafts that the full-depth pass rejects. Text comes out several times faster with no retraining and no loss of accuracy.
Technical breakdown
- Problem: Looped language models require T sequential recurrent-block applications per token, causing substantial latency despite parameter efficiency from weight sharing.
- Method: Wavefront Decoding (WFD) maintains a diagonal wavefront of mixed-depth token states, drafting at shallow depth while deepening earlier positions toward full-depth verification in a single batched forward pass. Tokens rejected at intermediate depth are corrected using full-depth output; cross-recurrence KV sharing via low-rank adaptation reduces long-context traffic.
- Key results:
- 2.42× speedup on Ouro-2.6B, 3.54× on Huginn-3.5B over autoregressive decoding across Spec-Bench
- 4.81× speedup on Huginn with cross-recurrence KV sharing
- Acceptance rates of 0.92–0.94; task accuracy maintained on GSM8K and MATH-500
- Why it matters / caveats: Training-free approach applies to both full-stack and prelude-recurrent-coda looped architectures. Evaluation limited to two public model checkpoints; long-context scenarios remain bottlenecked without KV sharing.
Beyond Timestamps: Decision-Aligned On-Policy Distillation for Long-Horizon Agents →
When a well-informed teacher coaches an agent step by step, its advice is often attached to the wrong moment, since the agent's real decisions can span several turns. The authors re-score the agent's own response under matched situations drawn from sibling attempts, then spread credit across variable-length decision stretches. It beat standard training across household, shopping and search tasks.
Technical breakdown
- Problem: On-policy self-distillation misaligns privileged supervision with student decisions when corresponding decisions occur at different timesteps and when single decisions span multiple turns.
- Method: ALIGNOPSD performs decision-aligned supervision rectification by matching student and teacher decisions across sibling rollouts using thinking traces as semantic surrogates, then re-scores the same response under matched privileged contexts. Semi-Markov hierarchical credit assignment derives variable-duration decision spans from correspondence changes and allocates outcome-grounded credit using rectified evidence.
- Key results:
- Outperforms GRPO by 5.5–8.7% across eight backbone-metric comparisons on ALFWorld, WebShop, Search-QA
- Ranks first in six of eight aggregate metrics
- Ablation: rectification alone contributes 3.7 pp on WebShop Score; adaptive spans add 1.6 pp over turn-level allocation
- Why it matters / caveats: Addresses fundamental misalignment in privileged supervision for agents. Evaluation limited to three benchmarks with two Qwen model sizes; transfer to other architectures and domains unclear.
When Do Model Internals Help? Exploring the Role of Representation Engineering in LLM Safety →
Safety work on language models either changes their outputs or reads and adjusts their internal states, but the two are rarely compared fairly. The authors run matched comparisons for both blocking unsafe behaviour and detecting it. Output-level alignment controls behaviour best but erodes after later harmless fine-tuning; internal-state methods shine with little data, detect risk almost as well far more cheaply, and can restore lost safety.
Technical breakdown
- Problem: Representation engineering (steering and probes) claims to improve LLM safety but lacks side-by-side comparison with behavioral safeguards across control and monitoring tasks using matched evaluation.
- Method: Compares DPO behavioral alignment against three steering methods (CAA, probe-based, flow-based) for control, and representation probes against text monitors (Qwen3Guard, fine-tuned LLM) for monitoring. Robustness tested via jailbreak attacks and benign fine-tuning; practicality measured via capability metrics; monitor-guided interventions tested via blocking and regeneration.
- Key results:
- DPO achieves lowest ASR under both AIM and refusal suppression before and after benign fine-tuning
- Flow steering competitive in low-data regime with high-quality contrastive data
- Representation probes achieve competitive AUROC (0.982) vs Qwen3Guard (0.996) at 7.6×10⁵ fewer marginal FLOPs
- Probe-triggered blocking reduces post-update DPO ASR from 0.500 to 0.042 with minimal over-refusal increase
- Why it matters / caveats: Provides practical guidance on when representation engineering substitutes for or complements behavioral safeguards. Evaluation uses non-adaptive attacks; adaptive attackers tailored to deployed safeguards not tested.
DroneWAM: Efficient World Action Model for Drone Visual Navigation →
Drones need to anticipate what their camera will see after a candidate move, but generating predicted images is too slow. The authors predict in a compact internal representation instead, compress the visual features, and let a learned gate decide how far ahead to look for each scene. Trajectories get more accurate, delay drops, and fewer prediction steps are needed; they also release a simulated flight dataset.
Technical breakdown
- Problem: Aerial world-action models require repeated recurrent propagation during online control, with existing approaches using fixed prediction horizons regardless of scene complexity.
- Method: DroneWAM adopts JEPA-based latent prediction in representation space. Pretrained Resampler compresses dense encoder features from 512 to 128 tokens, reducing per-step computation. Preference-trained Gate uses DPO to allocate adaptive rollout depth based on current scene. Introduces DroneNav-6D: simulated 6-DoF dataset with synchronized RGB, flight trajectories, and wind disturbances across five environments.
- Key results:
- Open-loop: ATE reduced from 1.8366 (FastWAM) to 1.1062; inference latency reduced from 712.33 ms to 483.29 ms
- Closed-loop: navigation progress improves from 0.510 to 0.774
- Adaptive rollout reduces average prediction depth from 8 to 4.58 while improving trajectory accuracy
- Compression from 512 to 128 tokens yields 18.3% latency reduction with negligible accuracy loss; generalizes to LIBERO with 53.8% latency reduction
- Why it matters / caveats: Demonstrates efficiency-oriented design principles for predictive navigation. Closed-loop task success remains low (~77%), limiting real-world deployment readiness.
Draft-KV: Learning Useful Latent Communication Between Language Models →
Models can pass internal states instead of text, but the authors show receivers barely notice when the message comes from an unrelated question, so the apparent benefit was not the content. Their method instead sends the internal states a helper forms while drafting an answer, read through a small trainable attachment while both models stay frozen. Content now matters, and gains grow with the helper's size.
Technical breakdown
- Problem: Existing latent communication interfaces show double-digit system gains but pairing gains below 1 pp, indicating the receiver exploits interface adaptation rather than benefiting from correctly paired content.
- Method: Draft-KV sends key-value states from sharer's draft answer through lightweight gated attention branch (1.05M trainable parameters, 348× smaller than C2C). Progressive training stages: (1) message reconstruction, (2) answer alignment, (3) task training with one-sided guard against mismatched message harm. Linear projections place draft KV states in separate memory with gated read.
- Key results:
- Matched accuracy on MMLU-Redux: 78.04% with Qwen3-8B sharer and Qwen2.5-0.5B receiver vs 36.40% deranged, 37.45% receiver-only
- Pairing gain averages 25.65 pp across 35 public settings vs C2C's 1.31 pp
- Sharer scaling: system gain rises 32.6 pp (0.6B→8B) vs C2C's 1.4 pp
- Transfers to private-evidence HotpotQA/2WikiMultihopQA; exceeds both models alone in 7 settings
- Why it matters / caveats: Separates content-dependent communication from system improvement via pairing gain metric. Limited to multiple-choice evaluation; open-ended tasks and heterogeneous model families may differ.
TokenCast: Forecasting Token Consumption During LLM Agent Execution →
How many tokens an AI agent burns on the same task can differ more than tenfold, so budgets are hard to set. The authors learn a cost profile for each stretch of work that records both its own usage and how much it swells the context every later step must reread, updating the forecast as the run proceeds. Predictions beat alternatives, cost little, and save tokens under a budget.
Technical breakdown
- Problem: Token consumption in LLM agents varies by over 30× across runs due to stochastic behavior and context accumulation, making cost forecasting difficult and enabling budget control.
- Method: Segment-cost factorization decomposes execution into call count, net input-length change, and residual baseline. Exact composition identity propagates context growth into downstream costs. Prefix-suffix predictor combines direct and compositional LightGBM forecasts, refreshed at task start, call start, in-call update, and task update. Quantile models provide prediction intervals; correction model reduces out-of-fold error.
- Key results:
- MAE reduction vs strongest comparator: 47.9% at In-call Update, 30.4% at Task Update, 5.3% at Task Start; average 14.5% across 96 combinations
- Budget-control replay: TokenCast uses 21.3% fewer tokens at matched trace completion across seven budgets on SWE-bench Verified
- Prediction overhead: 32.8 ms mean cumulative time per run; adds <0.03% to 129 s median wall time
- 90% interval coverage reaches 82–92% across prediction points; MIS reduction of 32% vs Self-Prediction
- Why it matters / caveats: Enables principled token budget allocation without additional LLM calls. Generalization: zero-shot transfer to unseen domains yields 1.31–1.47× target MAE; adaptation with 20 tasks improves to 0.82–0.85×.
KVCMAS: Efficient KV Cache Correction for Shared Context in Multi-Agent Systems →
When several agents share one model but each gets its own role instructions, each rebuilds its own memory of the same growing conversation. The authors capture the differences between agents' memories in a compact form and pass corrections along the chain of agents, avoiding an extra setup pass. Accuracy holds or improves while response start times and peak memory fall substantially.
Technical breakdown
- Problem: Prompt-specialized multi-agent systems with shared models but agent-specific prefixes repeat prefill of growing shared context, with direct cache reuse causing accuracy degradation and existing correction methods incurring memory-intensive online state.
- Method: KVCMAS stores cross-agent cache deviations in compact low-rank form (effective rank <32). Chained correction references previous agent cache already produced by workflow, avoiding separate context-free reference prefill and using exact first-agent cache. Online anchor pool matches current cache to stored anchors; similarity-weighted correction combines low-rank factors via truncated SVD. Reduces memory from O(V NLD) to O(V Nr(L + D)).
- Key results:
- Achieves highest accuracy among KV sharing methods on 3 of 5 benchmarks; ranks second on GSM8K and HumanEval (within 0.7 pp of KVComm)
- Lowest TTFT among delta correction methods; 2.0× speedup over NonShared at 32K tokens, 8 QPS
- Peak GPU memory: 3.7× reduction vs KVComm; within 2% of FullShared
- Chained correction reduces TTFT by 34% vs non-chained at 32K tokens, 8 QPS
- Why it matters / caveats: Solves accuracy-memory-latency trade-off for prompt-specialized agents. Evaluation limited to 5 benchmarks; scalability to deeper delegation chains unclear.
PReCache: Efficient KV Cache Sharing for Multi-LoRA Agents via Low-Rank Precomputation and Neutral Reconstruction →
Agents that share a backbone model but wear different small adapters each reprocess the same growing history. The authors precompute each agent's compact adapter-specific part once when the text is first read, and rebuild the shared part from adapter-free states so one agent's style does not contaminate the next. Waiting time and throughput improve markedly with only a slight accuracy cost, and no retraining is needed.
Technical breakdown
- Problem: Multi-LoRA agents repeatedly process growing shared trajectory while constructing separate KV caches; selective recomputation incurs computation, BaseShared retains repeated prefill, and existing methods require training or architectural constraints.
- Method: PreLRShared precomputes compact agent-specific low-rank (LR) caches when each context segment is first processed using all agent down-projections, eliminating repeated full-length prefill. ReBaseShared reconstructs shared base cache from adapter-free hidden states, reducing dependency on previous agent. Lazy prefill (LP) for single-stream and double batching (DB) for concurrent serving optimize reconstruction timing.
- Key results:
- PreLRShared: 3.1× TTFT speedup, 2.3× throughput improvement over NonShared at 66.4k tokens
- ReBaseShared: average accuracy drop only 1.1 pp vs NonShared, comparable to BaseShared
- Peak memory within 2% of FullShared at longest trajectories, 17–23% less than selective recomputation
- Under concurrent serving: 1.6× throughput improvement (PreLRShared), 1.3× (ReBaseShared) over NonShared
- DB distribution reduces queueing delay vs contiguous post-turn prefill
- Why it matters / caveats: Training-free framework supporting existing LoRA adapters with different down-projections. Evaluation limited to HotpotQA and ScienceQA; scalability to many agents and extreme context lengths needs validation.
ColNanoVDR: Document-Free Query Distillation for Multi-Vector Visual Document Retrieval via Optimal Transport →
Searching collections of scanned documents by page image normally requires running a huge model on every query, which is slow. The authors trained a small text-only model to mimic the big one's handling of queries, without ever processing document pages, using a mathematical technique for matching two sets of items. The small model kept nearly all the search quality while answering queries far faster.
Technical breakdown
- Problem: Multi-vector visual document retrievers require encoding billion-parameter vision-language models on every query, creating a computational bottleneck for deployment.
- Method: ColNanoVDR distills multi-vector retrieval using Optimal Transport with Learned Weights (OTW), aligning student query token sets to teacher token sets via entropic optimal transport. A lightweight linear head predicts learned weights for each student token, enabling document-free training on cached teacher query embeddings alone. Training uses block-wise Hessian approximation with one Sinkhorn solver pass per iteration.
- Key results:
- 149M text-only student retains 93-96% of 4.5-8.8B teacher NDCG@5 across ViDoRe v1-v3
- Query encoding 26× faster on single CPU thread (87ms vs 2290ms for teacher)
- OTW matches score distillation while reading 12.6× less cached teacher data (27 GiB vs 343 GiB)
- Why it matters / caveats: Enables efficient multi-vector visual retrieval deployment by bringing document-free distillation to late-interaction scoring. Practical path for single-CPU inference without re-indexing.
EvolvingAvatar: Interactive 3D Head Generation That Adapts as Conversations Unfold →
Conversational 3D avatars must speak, listen and react, but existing generators keep their settings fixed and ignore how a particular conversation unfolds. The authors built a system that keeps adjusting itself during the conversation from the partner's video and audio, without needing labelled example motions, and assembled a large collection of conversation videos for testing. Avatar motion matched real conversational behaviour more closely and improved as conversations continued.
Technical breakdown
- Problem: Interactive 3D head generation for conversational avatars requires coordinating speech, listening responses, and adaptation to partner behavior without target motion labels at deployment.
- Method: Test-time training with dyadic context prediction (DCP) self-supervised objective on user video and dyadic audio. Persistent fast weights accumulate conversation context across intervals via cross-attention and self-attention blocks; transient jaw adaptation responds to current speech. Region-structured FLAME codec coordinates expression, neck, and jaw with separate flow-matching heads.
- Key results:
- Expression P-FD improves 20.4% with full TTT vs no adaptation
- Out-of-distribution hard split: reduces mismatch with recorded statistics by 11.1% from first interval
- 86% human preference over DualTalk on OOD-Hard; 90% vs other baselines
- Why it matters / caveats: Shows interactive avatar generation can learn from conversation patterns without motion supervision, enabling adaptive verbal and nonverbal behavior coordination.
Adaptive Consistency Graph for Long-Horizon Agents →
Language-model agents handle short tasks well but drift away from the original goal during long chains of actions and tool calls. The authors store everything an agent learns during a run in a growing graph that records where each fact came from, then assemble a focused, size-limited briefing for each decision. Agents finished noticeably more long tasks, with the biggest gains on web-browsing work.
Technical breakdown
- Problem: LLM agents drift from original task objectives during long execution sequences as evidence accumulates and execution state becomes disconnected.
- Method: Adaptive Consistency Graph (ACG) organizes execution evidence as source-grounded memory units with event provenance and occurrence records. Dual-query retrieval seeds from task and latest interaction; graph expansion via structural-entropy clustering; budget-aware context rendering with requirement-grounded incidences prioritizes current-event evidence.
- Key results:
- GPT-5.6-luna ACG average: 50.2% vs 44.5% ReAct across three benchmarks
- BrowseComp-Plus: 73.5% vs 62.4% baseline
- SWE-bench Lite: 64.7% vs 60.3% with ReAct
- Why it matters / caveats: Addresses distribution shift in long-horizon execution through persistent, requirement-centered evidence retrieval. Task success improvements uneven across benchmarks, with DeepPlanning remaining difficult (12.4%).
DISCO: Distributed Long Context Scaling with Grounding-Reasoning Disaggregation →
Models advertise huge context windows, but their reasoning falls apart as the input grows. Inspired by distributed computing, the authors split long documents across many small models that each search their own slice, while a central model, trained to plan, decides what to look for and assembles the answer. Accuracy held up at extreme lengths where normal approaches collapsed, at a fraction of the cost.
Technical breakdown
- Problem: LLMs suffer context rot as input length increases, with reasoning quality collapsing due to structural entanglement of contextual grounding and logical reasoning in monolithic architectures.
- Method: Distributed long context scaling (DISCO) partitions context across Worker LLMs for parallel local grounding via Map actions; central Driver LLM handles global reasoning via Reduce actions. Driver trained with GRPO to optimize DAG planning; iterative refinement loop detects information gaps and replans. Lightweight Qwen3-4B Workers with reasoning-dense Qwen3-14B Driver.
- Key results:
- RULER-QA 1M tokens: 78.4% accuracy vs 10.9% retrieval baseline
- LongBench v2 Long subset: 48.7% Qwen3-14B vs 38.9% full context
- Inference cost reduction 80% vs full context while matching Gemini-3-Pro performance
- Why it matters / caveats: Paradigm shift from monolithic to distributed context processing eliminates context rot. Proof that separating grounding from reasoning enables robust million-token inference.
G^2PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation →
Shrinking a trained language model to lower-precision numbers saves memory but loses accuracy, and existing methods rely on guidance computed once at the start that goes stale. The authors refresh that guidance before each block of the network, combine two kinds of error information, and limit how far any single correction can move a weight. The compressed models stayed closer to the originals than previous approaches.
Technical breakdown
- Problem: Post-training quantization methods either use stale Hessian estimates (global objectives) or lack first-order gradient information (layer-wise local objectives), limiting quantization accuracy.
- Method: G^2PTQ integrates first- and second-order information under block-wise optimization with refreshed estimates before each Transformer block. Theorems prove block-wise Hessian approximation via outer product of gradient vectors; trust-region scaling dynamically bounds gradient steps to prevent exploding updates. Efficient recursive computation maintains GPTQ complexity.
- Key results:
- LLaMA3-70B 4-bit weight-activation: 72.94% QA accuracy vs 68.38% GPTAQ
- Qwen3-0.6B 3-bit: 39.83% QA (32.93% with sequential strategy) vs 29.92% GPTQ
- Quantization time: 2.24 hours, 12.54 GB memory for 8B model on single GPU
- Why it matters / caveats: Achieves better alignment with full-precision models by combining first- and second-order information. Practical efficiency maintained while improving upon global and local method tradeoffs.
SMAT: Simple and Efficient Merge-Aware Training →
Separately fine-tuned models can be combined into one without retraining, but they were never trained with that combination in mind, so quality drops. The authors note that merging effectively rescales, deletes and perturbs each expert's changes, and simulate those three effects during training. Merged models performed better across several merging recipes and several base models, with almost no extra training time.
Technical breakdown
- Problem: Expert models trained independently merge poorly due to overlapping updates and sign conflicts that reduce merged performance relative to individual experts.
- Method: Simulates merged parameters during expert training via Scale (sample reweighting coefficients), Mask (drop coordinates stochastically), and Perturb (add uniform noise simulating other experts). Periodic scheduling applies simulated loss every t=4 steps; kernel fusion computes Scale-Mask-Perturb together; parameter storage switching avoids repeated copying. One forward and backward pass per step maintained.
- Key results:
- Llama-3.2-1B: 45.77% merged score (+1.12pp over OrthoReg with 82% less training time)
- CLIP ViT-L/14: 87.78% average (+3.8pp over FT) across five merging methods
- Training overhead: <2% vs standard fine-tuning
- Why it matters / caveats: Enables models to remain effective after merging without sacrificing independent performance. Loss-smoothing mechanism along merge-relevant directions provides robustness to parameter variation.
KernelZero: Co-Evolving Proposer and Coder for Continuously Improved GPU Kernel Generation →
Writing fast, correct low-level code for graphics processors is hard to automate, partly because good training examples are scarce. The authors pair two models that improve each other: one invents practice problems aimed at the coder's current weak spots, the other writes the code, and rewards only start favouring speed once correctness is dependable. The small resulting model outperformed much larger commercial systems.
Technical breakdown
- Problem: Generating GPU kernels (CUDA/Triton) that are both correct and efficient requires balancing conflicting objectives, with limited high-quality training data aligned to model capabilities.
- Method: Co-evolution framework with Proposer generating Torch modules from API co-occurrence distributions and Coder translating to kernels. Frontier-driven module generation extracts APIs from failed cases; Coder trained with cold-start distillation on compiler optimization principles (Tiling, Fusion, Pipeline, Reordering) then CA-GRPO with group-level correctness threshold separating correctness and performance optimization.
- Key results:
- CUDA KernelBench Level 1/2: 75.8%/69.6% pass@1, 100%/97% pass@10
- Triton Level 1/2: 77.2%/72.5% pass@1
- Outperforms Claude-4.5-Sonnet on CUDA, DeepSeek-V4-Pro on Triton
- Why it matters / caveats: Continuous improvement paradigm through frontier-driven data generation and correctness-aware optimization. Automatic curriculum learning enables co-evolution without manual task curation.
On-Policy Self-Distillation for Multi-Turn Image Editing →
Image editors that follow written instructions work well once, but quality collapses when edits are applied repeatedly, because each new edit starts from the model's own imperfect output. The authors retrain the model on its own generated intermediate images, supervised by a copy that saw clean inputs, with no need for multi-step examples. Long editing sessions held together far better while single edits stayed as good.
Technical breakdown
- Problem: Image editing models degrade rapidly under multi-turn application where each output becomes the input to the next, with errors accumulating into severe visual artifacts.
- Method: MT-OPSD performs on-policy self-distillation via identity rollouts to isolate model-induced errors. Two-branch training: identity branch uses rollout state as target preventing further drift; editing branch pairs rollout states with real instructions, matching clean-conditioned teacher velocities via sparse query-based velocity matching. Adaptive curriculum progressively increases rollout depth; gated teacher promotion replaces frozen teacher with improved checkpoints.
- Key results:
- LME-Bench 10-turn: SR@10 improves from 3-15% to 38-52% across three backbones
- Collapse rate CR@10 reduces to 0.01-0.04
- Preserves single-turn quality (ImgEdit scores 4.28-4.52 vs 4.51 baseline)
- Why it matters / caveats: Extends strong single-turn editing to multi-turn scenarios via self-generated conditioning states. No multi-turn annotations needed; works across different editing architectures.
Why Deterministic PRM Guidance Underperforms in Discrete Diffusion Reasoning →
Some text generators reveal a rough draft at every step, tempting us to score partial drafts and keep only the promising ones. Measuring fairly under equal compute, the authors found this loses to simply generating several complete answers and picking one with a final judge. Early drafts are scored almost at random, so good candidates get thrown away, and the step-scorer also judges finished answers poorly.
Technical breakdown
- Problem: Process reward models (PRMs) for discrete diffusion LLMs guidance fail to improve over independent sampling plus outcome reranking, despite carrying useful intermediate-state signal.
- Method: Matched forward-pass diagnostic protocol separating pool damage from terminal selection quality. Bidirectional PRMs with full self-attention compared against causal variants; sequential Monte Carlo sampler decouples pool restoration from final selection; readout ablations isolate design failures.
- Key results:
- GSM8K: ORM Rerank@8 reaches 75.13% vs PRM Guided K=8 reaches 65.18% (9.95pp gap)
- PRM ROC-AUC decays from 0.77 (least-masked) to 0.54 (most-masked states)
- Top-1 guidance reduces oracle ceiling from 81.05% to 67.30% (13.75pp damage)
- Why it matters / caveats: Reveals two separable failure modes: mask-ratio signal decay and poor final judgment. Mean pooling causal PRMs critical weakness; last-token pooling recovers 70% of gap. Shows guidance evaluation must account for both defense stages.
SkillDRE: Dual-Stage Red-Team Evolution of Agent Skills via Pre-Execution and Runtime Feedback →
Reusable skill packages let AI agents improve themselves, but attackers can evolve harmful ones the same way. The authors built an automated loop that rewrites a malicious package against both the scanner that inspects it beforehand and the defences that watch it run, cycling between the two. The evolved packages succeeded far more often and slipped past scanning entirely, showing that each defence must be tested alongside the other.
Technical breakdown
- Problem: Malicious skills can evolve to bypass both pre-execution scanning and runtime defenses, making skill self-improvement a security concern for LLM agents.
- Method: Dual-stage red-team evolution with fixed attack targets and judge rules. Phase 1 scanner-guided pre-execution evolution minimizes SkillScan risk score via iterative mutation; Phase 2 runtime-guided evolution addresses SkillSonar stopping and target realization. Closed loop: runtime revisions return to Phase 1 rescanning before next execution. Autonomous target and judge rule construction with validator ensemble approval.
- Key results:
- Average attack success rate: 45.28% across four victim models vs 2.53-9.00% baselines
- SkillScan detection rate: 0.00% on final skills vs 97.34-100.00% baselines
- Largely preserves benign-task accuracy (±3.58pp degradation)
- Why it matters / caveats: Demonstrates two-stage defense feedback enables adaptive red teaming. Evaluating pre-execution and runtime defenses in isolation misses emerging attack capabilities. Shows skill self-evolution mechanism creates security risks beyond static attacks.
VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis →
Generating new camera viewpoints from only a few photos forces a trade-off: reconstruction methods respect the real geometry but cannot invent hidden areas, while generative models invent freely but drift. The authors feed a geometry estimator's 3D points, each with a confidence value, into a video generation model, and keep the generated views consistent along tracked points. Results were strong for both nearby and distant viewpoints.
Technical breakdown
- Problem: Existing novel view synthesis methods face a fundamental trade-off: reconstruction-based approaches preserve observed geometry but struggle to synthesize unseen regions, while diffusion-based methods provide strong generative priors but rely on implicit source-to-query correspondence.
- Method: VGGT-Diff combines geometry foundation priors with video diffusion by routing visual geometry latents from VGGT-Ω into a pretrained video diffusion model. The confidence-aware Visual Geometry Router (VGR) transforms VGGT-Ω features and 3D locations into query-aligned conditions, while Point-Track Residual Consistency (PTRC) aligns predicted-clean residuals along reliable 3D tracks to reduce cross-view drift.
- Key results:
- Achieves 18.104 PSNR on DL3DV (vs 17.180 for FrameCrafter) with state-of-the-art or competitive performance across interpolation and extrapolation
- Reduces Chamfer distance by 26.7% relative to FrameCrafter under VGGT-Ω evaluation
- Why it matters / caveats: The method successfully bridges geometry-aware reconstruction and generative completion, achieving strong performance with only 1K training scenes compared to competitors. The approach demonstrates that geometry conditioning can guide diffusion models while remaining robust to imperfect geometry through stochastic regularization.
ControlScope: Workflow Revision and Reliability in LLM Agents →
When an AI agent reviews a workflow already in progress, how much should it be allowed to change? The authors compared leaving the plan alone, tweaking only the next step's inputs, and rewriting everything, across file tasks, household simulations and app workflows. Wider permission solved more tasks but sometimes interrupted plans that were already working, so scope and review timing both matter.
Technical breakdown
- Problem: How much of a running workflow should a language model agent be permitted to revise—continue unchanged, edit argument values, or replace the entire unfinished workflow?
- Method: ControlScope compares three nested revision scopes (KEEP, ARG, FULL) across filesystem tasks, ALFWorld, and AppWorld environments using two source programs per task and multiple reasoning-reviewer draws. The framework measures repair availability, selected operations, and execution stability from the same public state.
- Key results:
- FULL solves 15-16 of 20 filesystem tasks vs 13 for KEEP across two source programs and three reasoning draws
- On ALFWorld 87-task reasoning panel: KEEP/ARG/FULL end 85/86/87 after common recovery
- Five-call protection saves 19.4% of logged output while losing only one success
- Why it matters / caveats: The study reveals that broader revision scope enables better task completion but also shows execution disruption when later revisions interrupt viable continuations. The work connects repair availability to agent selection and execution stability, showing that edit scope, valid operation selection, and review timing jointly shape outcomes.
Playing to Par: Reinforcement Learning for Provably Optimal Quadrilateral Block Decompositions →
Dividing a flat shape into four-sided blocks has a known mathematical floor on how many awkward corner points are unavoidable. The authors trained a program by trial and error to edit meshes toward that floor, first letting it copy easy examples built backwards from known perfect answers. It reached the theoretical minimum on most test shapes, beat a standard meshing tool, and handled shapes larger than any it trained on.
Technical breakdown
- Problem: Quadrilateral block decomposition must balance completeness, element shape quality, and minimizing irregular vertices—with the discrete Gauss-Bonnet identity providing a lower bound on vertex irregularity that depends only on topology and corner angles (the "par").
- Method: An RL agent operates directly on a mesh's half-edge data structure through local edits (chord insertion, edge splitting), using a convolution network that follows mesh connectivity. Training overcomes severe reward sparsity via behavior cloning on trivially-optimal backward-walked polyominoes followed by PPO on generated domains.
- Key results:
- Produces all-quadrilateral meshes on all 96 in-distribution test domains (8-24 boundary corners) with 95.7% passing usability threshold and 90.2% reaching provably optimal par
- On 64 larger-boundary domains (25-50 corners), completes all, is usable on 62, keeps median excess over par below 1 vs Gmsh's 39 at matched element count
- Against Gmsh at matched element count, wins on quality/regularity across 90% and 100% of decided domains respectively
- Why it matters / caveats: The method demonstrates that RL can solve combinatorial geometry problems when guided by theoretical bounds and symbolic optima. The approach is computationally efficient (trains in ~4 hours on M2) and generalizes to domains twice the training size and even unseen curved boundaries.
What Masking Geometry Works Best for EEG Foundation Models? →
Models that learn from brain recordings train by hiding parts of the signal and predicting them, but nobody had tested which hiding pattern works best on its own. The authors held everything else fixed and varied how long in time and how wide across the scalp the hidden patch was. One setting won under both learning schemes, many others were equally good, and a previously unnoticed failure mode appeared.
Technical breakdown
- Problem: EEG foundation models use diverse masking strategies but lack ablation of masking geometry in isolation; practitioners lack principled guidance on mask choice.
- Method: Controlled grid sweep across 29 spatio-temporal mask configurations (6 temporal lengths, 5 spatial radii) under both MAE (input-space reconstruction) and JEPA (latent-space prediction) frameworks on a unified REVE pipeline, evaluated on OpenEEGBench (12 datasets) with linear probing on frozen features.
- Key results:
- Optimal mask geometry (L=2s, r=9cm) achieves REVE-Base-level performance using 12.7M vs 69M parameters and 14 H100-hours vs 260 A100-hours pre-training
- Both frameworks converge on same optimal mask and shared failure modes (r="one", r="all", long temporal blocks L≥8)
- 11 MAE and 9 JEPA configurations are statistically indistinguishable from best (ex-aequo cluster)
- Identified bias-inflation collapse: JEPA at r="all" exhibits systematic downstream performance degradation as training progresses, invisible to standard collapse detectors
- Why it matters / caveats: The study provides actionable default guidance (L,r)=(2s,9cm) and reveals that masking strategy is the dominant performance lever within the shared pipeline. The work demonstrates JEPA's sensitivity to multichannel masking choices and that per-objective scale normalization matters for principled design.
Safe Error Correction for Language Models: Frozen-Base Adjustment with Capability Preservation →
Can a small add-on fix a language model's mistakes without spoiling everything else it can do? The authors left the model completely untouched and trained a tiny module that nudges its output probabilities, anchored so it cannot stray far from the original. It fixed about half the errors on a domain exam with no measurable loss elsewhere, while conventional fine-tuning fixed more but badly damaged general ability.
Technical breakdown
- Problem: Model adaptation to fix domain-specific errors typically degrades general capabilities (correction-capability tradeoff); can a small correction module preserve base capabilities while fixing errors?
- Method: CRN v2 is a lightweight logit-level correction module (34M trainable parameters, 0.73% of 4.65B Gemma 4 E2B) applying low-rank adjustment with KL preservation (λ=0.1) during supervised fine-tuning followed by reference-free DPO, keeping the base model completely frozen.
- Key results:
- Corrects 53.3% of errors on 60-question domain exam (CEHRI), 43.3% on reworded variant, with zero degradation on MMLU/BoolQ/car-wash benchmarks
- LoRA baseline at matched budget corrects 83.3% but degrades 17-75 percentage points on capability tests
- Ablation: lowering KL weight from 0.1 to 0.01 degrades correction from 53.3% to 35.0%
- Deep injection variants and longer training (5000 SFT + 2000 DPO) do not exceed 53.3%
- Why it matters / caveats: The frozen-base principle with KL anchoring prevents catastrophic forgetting by construction, but achieves lower correction rates than methods that modify weights. Results are limited by small test set (N=60) and in-distribution training data; the approach trades correction coverage for guaranteed capability preservation.
NanoForecast v0.5: Competitive Time Series Forecasting Through Training Pipeline Optimization →
A small forecasting model for time-series data was underperforming, and the authors suspected the training code rather than the design. Without changing the architecture, they fixed how the loss was computed, corrected data shapes and broadened the training variations. Error dropped sharply on the same data and compute, making it competitive with far larger models and able to make predictions without a graphics card.
Technical breakdown
- Problem: Time series forecasting models often require GPU infrastructure for both training and inference, putting them out of reach for practitioners with consumer hardware.
- Method: Retraining the NanoForecast v0.3 architecture (unchanged 6.5M parameters) with three corrected training pipeline issues: loss-scope handling (not truncating predictions to forecast horizon), tensor shape alignment (quantile loss computed before truncation), and augmentation coverage (adding flip, masking, scaling, shifting, jitter).
- Key results:
- 43.8% improvement in overall MASE (3.030 → 1.704) using identical architecture, data, and compute budget on six benchmarks
- Beats TimesFM (200M parameters) on 4 of 6 datasets: ETTh1 (0.676 vs 0.705), ETTh2 (1.110 vs 1.360), ETTm1 (0.287 vs 0.545), exchange rate (4.317 vs 4.383)
- Inference latency: 19.5ms PyTorch FP32 on Apple M4 CPU, 10.7ms via ONNX Runtime
- Training: 12 hours on single NVIDIA T4 (Google Colab)
- Why it matters / caveats: Demonstrates that training pipeline correctness can drive larger performance gains than architecture changes. The work emphasizes the importance of thorough loss implementation audit and suggests many published results may underestimate achievable performance on existing architectures.
Not All Objectives Are Born Equal: Priority-Constrained Descent for Hierarchical Multi-Objective Optimization →
In machine learning some goals matter more than others: accuracy is the point, while things like compactness are constraints. Most methods treat all goals as equals, which can stall progress. The authors keep the main goal's improvement direction and bend it only as much as needed to guarantee the secondary goals also improve, controlled by a single dial, with exact solutions for two or three goals.
Technical breakdown
- Problem: Multi-objective optimization typically treats objectives symmetrically via weighted sums, but deep learning problems have inherent hierarchies (primary objective plus secondary constraints like sparsity, robustness). Symmetric methods can halt at "conflict equilibria" where gradients cancel despite no objective reaching stationarity.
- Method: Priority-Constrained Descent (PCD) solves a convex QP to find the minimum-norm perturbation of the primary gradient that maintains progress on secondary objectives. A single tolerance τ ∈ [0,1] controls deflection; gradient normalization via EMA makes τ scale-invariant across heterogeneous objectives.
- Key results:
- On CIFAR-100 ResNet-34 structured pruning: PCD maintains 70.9% accuracy at 90% sparsity vs baselines collapsing; exceeds accuracy at matched sparsity throughout compression regime
- Unstructured pruning with sparsity 92.6% and rank 24.5: maintains 71% accuracy on tri-objective ℓ1/nuclear norm task
- For K=2 objectives: closed-form solution identifies when primary descent breaks down (threshold depends on gradient angle and τ)
- Fixed points are composite multi-stationary (every objective individually zero), strictly stronger than Pareto stationarity
- Why it matters / caveats: PCD treats hierarchy explicitly in gradient geometry rather than through learned weights, and avoids conflict equilibria. Work is limited to convex secondary objectives and provides no convergence guarantees for stochastic training.
When Privacy Moves ML-Mediated Decisions On Device: Information and Incentive Misalignment in Auctions →
When ad auctions move onto people's own devices for privacy, each device decides spending using shared budget information that is inevitably out of date. In a simulation, the authors show overspending that grows sharply with the delay, persists under milder budget limits, and is only partly fixed by a local spending check. They also show that letting the scoring step change payment units gives bidders room to game auctions.
Technical breakdown
- Problem: Moving ML-mediated auction decisions onto privacy-preserving clients decentralizes both inference and budget constraints; shared budget state cannot be globally current without communication, creating information misalignment and potential overspending.
- Method: Auction-logic-faithful simulation with 50 devices, 36 campaigns across 12 verticals, 3 bidders per vertical; studies two economic failures (information and incentive misalignment) under various sync lag (Δ ∈ {0,1,2,5,10,25,50} ticks) and budget pressure settings with proportional pacing controller.
- Key results:
- Even pacing overspend: 17.77% at 1-tick lag, 1669.31% at 50 ticks under 20× budget pressure
- At 2× budget pressure, 50-tick overspend remains 106.95% (lag effect persists under milder budgets)
- Visible-budget guard eliminates zero-lag crossing but leaves 11.88% overspend at 1 tick (other devices' debits invisible)
- Score-space payment rule: 98.23% of rival auctions admit profitable deviation; critical-base-bid payment fixes per-auction truthfulness but neither proves dynamic truthfulness nor removes paced ranking
- Why it matters / caveats: Formalizes how distributed state creates economic failures; the information misalignment is fundamental to architecture, not fixable by controller tuning. Study uses dimensionless score units (not currency) and assumes truthful reports despite non-incentive-compatible payment.
How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure →
Model evaluations usually average a handful of prompts and print a ranking table. The authors audited one such study and found identical requests often gave different results, resampling left only the worst model's position stable, and an equally reasonable way of combining runs reshuffled much of the table. Several tested services were withdrawn within weeks, so they recommend publishing stability, raw outputs and measurement dates.
Technical breakdown
- Problem: Small-sample LLM evaluations typically report ranked tables, but how stable are those rankings under resampling and reasonable analysis choices?
- Method: Eight open model variants (8B-675B parameters) across five families, tasked with inferring prompt structure (nodes with ids, categories, priorities) on 22 synthesized and 23 real prompts. Metrics: node-set Jaccard (mean of pairwise agreement) and ID-stability (all-runs agreement). Rankings audited via joint cluster bootstrap, merge-rule sensitivity, and ground-truth validation against parser-verified annotations.
- Key results:
- Node-set Jaccard spans 0.39-0.96; 72% of cells not node-set-perfect (both metrics =1)
- Joint bootstrap: gpt-oss:120b holds rank 99% of replicates, middle four at 27-48%, top two at 68% each—table identifies worst reliably, not best
- Alternative merge rule (earliest-first vs higher-budget-first) changes four of eight rows and moves headline 7.1 percentage points
- Four of eight models in Table 1 retired within 10 weeks; study no longer executable
- Reproducibility ≠ validity: Jaccard and correctness against ground truth diverge (Pearson r 0.69-0.90)
- Why it matters / caveats: Demonstrates that ranked tables look more definitive than evidence supports. Study is task-specific (prompt-structure inference) and mostly measures transcription rather than recovery; endpoint deprecation underscores reproducibility risks with commercial LLM endpoints.
Training and Inference Dynamics of PLDR-LLMs: Row-Map Collapse, Renormalization, and Predictive Reduction →
This work asks what can safely be discarded when replacing the fine detail of a model's training process with a simpler summary that still predicts how the trained model behaves. The authors derive exact bookkeeping identities for how the model's internal quantities change, and test which simplified descriptions survive. Some candidate shortcuts were ruled out; others held under stated conditions, clarifying which coarse descriptions are trustworthy.
Technical breakdown
- Problem: What must be retained when replacing microscopic training dynamics with coarser descriptions sufficient for inference prediction?
- Method: Unified theoretical framework using exact path identities, finite work decomposition on row-centered learned maps, complete AdamW state dynamics, and chronological renormalization. Single-pass RefinedWeb training on pre-trained models; no new training experiments. Theory connects row geometry, row collapse, optimizer state, and conditional prediction via staged reductions.
- Key results:
- Finite work identities resolve row-energy changes into parameter sources, signed interactions, and numerical defects
- Chronological blocking preserves loss of homogeneous dependence on scalar energy (collapse is not equivalent to erasure of optimizer memory)
- Row-RG decision procedure on four trajectories (196,608 updates, 12 cells): ten cells fail nominal precision gate; moving finite fluctuation regions do not establish thermodynamic critical class
- Executed closure tests and observer-dependent operator reduction with full error budgets
- Why it matters / caveats: Provides mathematical formalism for understanding training-to-inference reduction in LLMs, combining exact identities with finite empirical evidence. Extensive appendices give statistical methods, complete outcome grids, and reproduction boundaries. Theory applies to single-pass consumption of RefinedWeb corpus with stated assumptions; scaling laws require additional work.
How Does "English (US)" Become the Default? Triangulating Structural Bias Towards American English Across the LLM Pipeline →
Interfaces routinely offer American English as the default English. Using matched pairs of American and British word choices, the authors traced this preference through training data, fine-tuning data, how text is split into chunks, how confidently models predict each form, and what they actually write. American forms were favoured at every stage, and asking for British English shifted but did not remove the default.
Technical breakdown
- Problem: American English emerges as a default linguistic norm across LLM pipelines despite global diversity of English use, raising concerns about linguistic homogenization and inequitable AI deployment.
- Method: Introduces DIALIGN, a training-free method for regional alignment analysis using lexical/grammatical evidence from Google Books n-gram frequencies. Triangulates American English preference across three stages: data exposure (six pretraining and 21 post-training datasets), representation (tokenization efficiency and per-token prediction cost), and generation (neutral vs. British English prompting on 10 model checkpoints).
- Key results:
- American English strongly favored in all six pretraining corpora (orthographic variants: 72.94%-86.81% AmE)
- All 21 post-training datasets favor AmE at 76.17% overall (instruction-tuning 75.95%, safety-feedback 86.86%)
- British English variants require 2.85%-18.72% more tokenizer fertility depending on contrast type and tokenizer
- Semantically equivalent BrE forms receive higher per-token loss across all 10 checkpoints even at equal token counts
- Under neutral English prompting, 69-79% of outputs classified as AmE across models; British-English prompting shifts but doesn't eliminate this default
- Why it matters / caveats: First rigorous pipeline-wide study of structural linguistic bias in LLMs, identifying concrete intervention points. Demonstrates how geopolitical data histories become embedded in model infrastructure. Limitations include restriction to AmE-BrE comparison.
AdaGuard: An Adaptive Guard Model with User-defined Policies →
Safety monitors for AI agents usually check against a fixed list of risks, but real deployments have their own rules. The authors built a training set where rules and behaviour are varied one at a time so the monitor learns what actually changes a verdict, plus a training method that rewards naming exactly which rules were broken. The resulting small monitors judge agent behaviour against rules supplied on the spot.
Technical breakdown
- Problem: Guard models for LLM agents need to assess trajectories under application-specific, user-defined policies that vary across deployments, requiring interpretation of both policies and recorded behavior.
- Method: Constructs AdaptiveSafety dataset (10,939 training, 1,000 test examples) with policy and behavioral counterfactuals. Proposes SafePO, a reinforcement learning algorithm combining structured verdict rewards distinguishing complete correct predictions from partial matches, with value-guided token weighting allocating fixed analysis/verdict weight budgets via independent value model predictions on causal prefixes.
- Key results:
- AdaGuard-4B: 89.30% binary accuracy on AdaptiveSafety, 71.82% on DynaBench
- AdaGuard-8B: 89.50% accuracy with 96.25% precision, 77.10% exact rule-set match
- SafePO improves 8B model: +1.10 pp accuracy and +1.12 pp rule exact-match on AdaptiveSafety over supervised-only baseline
- Outperforms specialized DynaGuard on AdaptiveSafety exact-match by 7 pp despite lower binary performance on DynaBench
- Why it matters / caveats: Enables policy-conditioned assessment of tool-mediated agent behavior with rule identification, addressing gap between conversational guards and agent trajectory evaluation. Validated with both custom and external benchmarks. Limitations: DynaBench gap remains, mixed SFT-to-RL effects at smaller sizes, explanation factual correctness not independently verified.
Routing Drift Alone Does Not Diagnose Failure in Merged MoE LLMs →
Merging two specialised models changes which internal experts handle each word, and people assume that change is damage. Through controlled swaps the authors show most of the change comes from shifted inputs rather than the routing component itself, and forcing the original routing back does not reliably help on tasks, though deliberately scrambling routing clearly does hurt. Changed routing alone is therefore not proof of a problem.
Technical breakdown
- Problem: Whether routing drift (changed expert assignments) after Mixture-of-Experts merging constitutes actual routing failure, and what evidence should justify repair attempts.
- Method: Develops routing analysis toolkit for controlled interventions: (1) decomposes route changes into representation-induced vs. parameter-induced components via component replacement on DeepSeekMoE/OLMoE/Qwen3-MoE; (2) tests whether structural metrics predict intervention gains using Jensen-Shannon divergence; (3) operationalizes routing failure as task loss recoverable under specified routing intervention with non-routing parameters fixed; (4) proposes Selective Router Repair (SRR) fitting source-derived expert-pair corrections as case study.
- Key results:
- 77.7%-96.9% of route changes attributable to input-representation shifts, not router parameters
- JS divergence poorly predicts source-route gains (AUROC 0.47-0.52, near chance)
- Source-route restoration produces confidence intervals spanning zero on multiple benchmarks
- Controlled router corruption (logits permutation) recovers 12.50-32.81 pp accuracy; natural source-route restoration shows no reliable benefit
- SRR updates show no evidence that source-likelihood advantages identify beneficial corrections; 440 matched pairs find no significant gain
- Why it matters / caveats: Distinguishes routing disagreement from routing failure, establishing that task-level consequences rather than source agreement must justify repair. Demonstrates limitations of likelihood-based corrections. No evidence that fitted updates improve average performance on tested checkpoints.
PyroAdapt: Adapting Wildfire Prediction under Spatial Heterogeneity and Temporal Shift →
Wildfires are rare and the conditions that predict them vary by place and drift over time, so models trained on history do poorly on new years. The authors adapt a model by retrieving locations with similar past conditions and retraining it to rank places by risk rather than estimate probabilities. It prioritised fire-prone locations better under a fixed daily inspection budget, caught more severe fires, and held up in later years.
Technical breakdown
- Problem: Wildfire occurrence prediction faces severe class imbalance, spatial heterogeneity (terrain/ecoregion variation in fire-covariate relationships), and temporal distribution shift, requiring adaptation to deployment-year conditions without target labels.
- Method: Pretrain-Retrieve-Rank (PYROADAPT) framework: (1) historical focal pretraining on labeled source years; (2) unlabeled target batch retrieves k=5 nearest historical neighbors in weather/spatial covariate space; (3) ranking adaptation via pairwise losses (direct/RDPO/selective) on retrieved historical labels, with same-day spatial comparisons in California to avoid day-level confounds. Incorporates terrain, learned ecoregion embeddings, and past fire rates into risk conditioning.
- Key results:
- Daily AP increases from 21.62% (continued focal) to 24.57% (selective ranking), +2.95 pp gain
- Top5% recall improves from 18.70% to 22.79%, +4.09 pp
- At fixed 34-cell daily budget, selective ranking captures 344 additional neighborhood-smoothed positive cell-days vs. continued focal
- High-DM wildfire recall gains: selective ranking +39.70 pp at top 5% severity, +28.18 pp at top 10%, +20.50 pp at top 20%
- Rolling temporal evaluation (2019-2021 Yosemite) shows persistent gains: direct ranking 74.19% AUROC vs. 73.12% pretrained
- Why it matters / caveats: Enables domain-wide wildfire prioritization under monitoring budget constraints with temporal robustness. Combines historical retrieval with ranking adaptation without target-year labels. Limitations: history features shrink sparse-data regions toward statewide prior; label is neighborhood-smoothed occurrence indicator rather than verified ignition; no operational readiness claimed.
Rolling-WAM: World Action Models with Rolling Imagination →
Robot controllers that imagine future video while planning actions are slow, because each cycle regenerates the whole imagined sequence from scratch. The authors keep a rolling window of partly finished predictions, fully completing only the action about to run while roughing out later ones, then sliding the window forward with each new camera image. Success stayed competitive in simulation and on a real humanoid, with several times faster replanning.
Technical breakdown
- Problem: World Action Models coupling action generation with future video prediction incur high inference latency due to full-sequence joint denoising at each replanning cycle, limiting closed-loop responsiveness in robotic manipulation.
- Method: Rolling-WAM maintains sliding window of W video-action chunks at staggered noise levels. At each replanning cycle, rolling denoising schedule (Eq. 2-3) fully denoises only imminent chunk for execution while partially refining farther-future chunks; after execution, window advances with new observation and appends fresh chunk at Gaussian noise. Trains joint video-action Diffusion Transformer on matching noise profiles with flow-matching loss (Eq. 8). Action tokens attend to entire visual prediction window; direct action-to-action attention restricted to same chunk.
- Key results:
- LIBERO average: 98.1% success rate (4 suites, 40 tasks)
- RoboTwin 2.0 bimanual: 93.5% clean / 93.0% randomized (50 tasks)
- Real-world Unitree G1 humanoid: 85.0% average (Doll Placement 85%, Plate Stacking 100%, Bead Pouring 70%)
- 4.5× steady-state inference speedup (215 ms vs. 978 ms Joint-WAM) with only 2 denoising steps per replanning cycle
- Window-size ablation peaks at W=5: 78.2% on selected RoboTwin tasks
- Why it matters / caveats: Distributes joint video-action denoising across control cycles while maintaining visual context and action generation quality, achieving faster replanning than both joint WAMs and VLA baselines while predicting longer horizons (80-action window). Retained predictions may lag behind rapid scene changes with long windows; design choices like window/chunk sizes warrant exploration across different task dynamics.