AI papers — 2026-09-28
Jump to one of 24 papers
- FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders
- RayOrch: Programming and Executing Lineage-Controlled Multi-Grain Dataflows for Foundation-Model Data Preparation
- Block Sparse Attention with Log-Linear Complexity
- InternW0-Δ: A World Action Model Bridging Predictive Dynamics and Actions with 20K+ Hours of Open Data
- Tactile-JEPA: Topology-Aware Self-Supervised Representation Learning for Distributed Tactile Sensors
- Enhancing Photogrammetric Digital Surface Models with Pretrained Diffusion Models and Multimodal Conditioning
- Jev in the Wild: A Data-Driven Analysis of the Jev Model's Functionality, Applications and Ecosystem
- FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
- TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations
- AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs
- SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL
- Game Arena: Strategic LLM Evaluation in Competitive Environments
- Do Implicit Personalization and Explicit Styles Conflict? PsPLUG: A Lightweight Plug-in for Balancing Personalization and Style in Customized LLMs
- CARD: Cluster-level Adaptation with Reward-guided Decoding for Personalized Text Generation
- TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding
- IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking
- ZooWork-ShopRanker: An Open, Preference-Aligned E-Commerce Reranker
- VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models
- Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations
- MOPD-Router: Rethinking Teacher Routing in Multi-Teacher On-Policy Distillation
- Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
- Paragraph Boundaries Are Not White Space: Compression Depth as the Signature of Hierarchical Structure
- Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching
- BoundInk: Boundary-Aware Online Handwriting Generation
FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders →
Systems that turn images into compact codes for generation must pick which layers of a pretrained vision model to blend, and shallower layers favor faithful reconstruction while deeper ones favor better generated images. The authors instead train on randomly chosen subsets of layers, which they show penalizes sensitivity to disagreement between layers. One decoder then works across many layer choices, reconstructing more faithfully and improving generated image quality without touching the pretrained encoder.
Technical breakdown
Problem: Representation autoencoders (RAEs) must fuse a frozen visual encoder's layers into a single latent space, but a fixed heuristic fusion (e.g., summing the last k layers) forces the pixel decoder and the diffusion generator to share one latent even though shallow layers favor reconstruction quality while deeper layers favor generation quality, creating a reconstruction–generation gap.
Method: FuseReg replaces the fixed heuristic fusion with training over randomly sampled, normalized subsets of a frozen encoder's layers (DINOv3-L, K=23 candidate layers), using separate Bernoulli drop rates for the ViT pixel decoder (p_dec) and the DiT generator (p_dit). The decoder is trained to reconstruct images from a normalized mean of a random nonempty layer subset (mean-preserving in expectation), while the DiT is trained with a flow-matching/x-prediction loss to regress the full-layer target from noisy subset-based latents. The paper also proves theoretically (via a linear squared-loss surrogate) that this normalized subset sampling preserves the full-layer mean while adding an explicit penalty on cross-layer disagreement, and shows this cannot be matched by any deterministic global fusion (including learned softmax gates).
Key results:
- A single FuseReg decoder (p_dec=0.95) reconstructs robustly across fusions (k=7/k=23/single layer ℓ11: PSNR 23.77/27.52/25.13 dB, rFID <0.6), versus fixed-fusion RAEv2 decoders that drop to 12.51–17.50 dB PSNR off their training fusion.
- Swapping in a FuseReg decoder alone (generator unchanged) reduces unguided gFID from 3.01 to 2.21 on the native k=23 fusion, and from 27.73 to 1.92 under a shifted k=7 fusion.
- Joint decoder+generator regularization reduces DiT-Base unguided gFID from 13.96 (baseline) to 9.93 at (p_dec=0.9, p_dit=0.7); on DiT-XL, gFID improves from 2.91 to 2.38 at (p_dec=0.95, p_dit=0).
- Overall claimed gains: 27% reduction in unguided gFID from decoder replacement alone, and 29% reduction from joint regularization on DiT-Base.
Why it matters / caveats: FuseReg improves the decoder–generator interface without modifying the frozen pretrained encoder or adding extra compute/architecture, narrowing a real trade-off in RAE-based generative pipelines. The authors note limitations: results are only on ImageNet-256 with DINOv3-L and DiT-Base/XL under matched budgets; the two stages respond to regularization differently (optimal drop rates depend on model scale, metric, and guidance), so applying FuseReg elsewhere requires separately tuning both rates; and the theoretical guarantees are exact only for a linear squared-loss surrogate, with nonlinear-model benefits established empirically rather than proven.
RayOrch: Programming and Executing Lineage-Controlled Multi-Grain Dataflows for Foundation-Model Data Preparation →
Pipelines that turn documents and videos into training data repeatedly split each item into many smaller pieces, and existing tools either hide this or force programmers to track which piece came from where. RayOrch makes these parent-child relationships a built-in, checked part of the program, batching pieces from different parents together while preserving order, completion and failure isolation. It scaled near-linearly across many GPUs and finished jobs noticeably faster than leading alternatives.
Technical breakdown
Problem: Foundation-model data-preparation pipelines (e.g., PDF-to-page-to-region, video-to-clip-to-frame) repeatedly change their unit of processing, and existing distributed dataflow systems (Ray Data, Daft, Spark, Dask) either hide this fine-grained parallelism behind coarse per-document jobs or expose it as flat records that force applications to manually track parent IDs, ordinals, completion, and reconstruction/failure logic, typically introducing a global regroup/shuffle barrier after each stage.
Method: RayOrch is a Python programming model and distributed execution engine built as a layer on top of Ray that represents pipelines as finite, acyclic hierarchical dataflows over "Domains" (logical levels like PDF/Page), with declared structural primitives F.expand (parent-to-child fan-out), F.reduce (parent-scoped ordered gather), F.filter, and F.broadcast, compiler-validated for Domain compatibility. At runtime it maintains persistent "structural lineage" state (each expansion's concrete ordered child set, each child's immediate parent and immutable ordinal, and terminal outcomes) separately from transient physical execution, using a per-Call FIFO Ready Queue that can batch "Grains" (one Call applied to one Entity) across different parents for GPU utilization, generation-fenced per-Grain commits to prevent stale/duplicate reports, and typed GroupFailure containment that suppresses undispatched sibling Grains scoped to one (Call, parent) pair without affecting unrelated parents.
Key results:
- On MinerU (3,689 PDFs, 174,744 pages) with 64 NVIDIA H20 GPUs, RayOrch reduces end-to-end time by 13.1% vs. Ray Data, 29.0% vs. Daft, and 51.6% vs. native MinerU, reaching 40.68 pages/s.
- Strong scaling on MinerU from 4 to 64 GPUs gives a 15.14x processing-time speedup (94.6% of ideal linear scaling); on the video pipeline (Qwen2.5-VL-7B), scaling 8 to 64 GPUs gives 7.82x speedup vs. 7.60x (Ray Data) and 6.24x (Daft).
- On Docling (2,000 PDFs, 4 GPUs), RayOrch reduces E2E time by 16.0% vs. Ray Data and 22.3% vs. Docling Serve.
- FIFO scheduling ablation reduces wall time from 634.1s to 579.3s (8.6% reduction) over rebatching alone.
- In a controlled failure-injection experiment (99 largest parents poisoned), RayOrch's typed failure containment prevents 6,241 of 23,514 non-trigger sibling computations (26.5% of poisoned-parent siblings) from entering the UDF, cutting non-trigger UDF calls from 45,408 to 39,167 and reducing mean wall time by 14.93%, while preserving all expected outputs for unaffected parents (Ray Data and Daft show near-zero, 0.18%/0.36%, time reduction since they cannot suppress sibling work at runtime).
Why it matters / caveats: By making parent-child lineage a first-class, compiler-checked and runtime-maintained construct rather than application-managed metadata, RayOrch removes the global shuffle/regroup barrier that limits scaling and failure isolation in flat-record systems like Ray Data and Daft, showing concrete gains for real document (MinerU, Docling) and video (Qwen2.5-VL-7B) pipelines up to 64 GPUs. The paper notes its scope is limited to finite, acyclic 1:M hierarchical dataflows with ordered gathering within a source microbatch — it explicitly does not support general joins, windows spanning microbatches, or feedback loops.
Block Sparse Attention with Log-Linear Complexity →
Long-text language models save work by attending only to selected blocks of earlier text, but choosing those blocks still requires comparing everything to everything, which grows with the square of the text length. The authors build a coarse-to-fine hierarchy of summaries and narrow candidates level by level, cutting selection cost sharply. Quality matched existing sparse methods while retrieval from long documents improved, with large speed gains on very long inputs.
Technical breakdown
Problem: In conventional block-sparse attention, the Top-K block-selection stage still scores every query against every key-block summary, costing O(N²/C) and remaining a bottleneck for long-context language models even though the actual attention computation over selected blocks is linear.
Method: The paper introduces PISA (Pyramid Sparse Attention), which builds a fine-to-coarse hierarchy of key blocks via mean pooling (leaf block size C=64, branching factor g=2 above the leaf level) and performs Top-K block selection top-down: starting from a single coarsest-level candidate, at each level it computes LogSumExp (LSE) scores over a bounded candidate set, keeps the Top-K blocks, and expands only those into their children at the next finer level, continuing to the leaf level. This gives O(log N) routing complexity per query and O(N log N) overall for a sequence of length N (O(log N) at decode time), and is implemented as hardware-aware, fused Triton kernels (a two-stage kernel for training/prefill, a single-stage kernel for decoding) that never materialize the full query-key score matrix. Two ablation variants, PISA-1 (first-order Taylor truncation of the LSE score, i.e. mean) and PISA-2 (adds a variance correction term), are also tested against the full LSE score.
Key results:
- Evaluated against Full Attention, BSA, NSA, and HiLS at 418M, 1.47B, and 2.67B decoder-only backbone scales, pretrained on 100B tokens at 4K context plus 10B tokens of continued pretraining at 16K.
- PISA achieves comparable perplexity/commonsense-reasoning accuracy to BSA/NSA/HiLS (e.g., at 2.67B: loss 2.2151 vs. BSA's 2.2184, avg multiple-choice accuracy 58.10% vs. BSA's 58.61%) while reaching the highest average containment (retrieval-style) accuracy among sparse methods at all three scales (e.g., 52.14% at 2.67B vs. 51.13% for BSA).
- On RULER long-context retrieval (2.67B model, 16K continued pretraining), PISA attains an average accuracy of 62.80% across four needle-in-a-haystack task families, versus 61.07% (NSA) and 54.99% (BSA).
- In a block-selection quality probe against a Full-Attention reference, PISA achieves the highest average Recall@8 and attention-mass ratio among selectors tested.
- Prefill block-selection latency: PISA is slower than BSA up to 16K but achieves 2.86x, 5.31x, and 9.95x speedups over BSA at 64K, 128K, and 256K tokens respectively; its key-block-reuse (Qtile=4) kernel is additionally 1.30-1.35x faster than a per-query fused implementation at those lengths.
Why it matters / caveats: PISA reduces the previously-quadratic block-selection cost of block-sparse attention to O(N log N) without sacrificing language-modeling quality and while improving long-context retrieval accuracy, becoming increasingly advantageous as sequence length grows past ~64K tokens. The authors note computational constraints limited the model sizes and pretraining data budgets tested, which could affect the magnitude of relative gains over baselines at larger scale.
InternW0-Δ: A World Action Model Bridging Predictive Dynamics and Actions with 20K+ Hours of Open Data →
Robots can learn to predict how a scene will change, but turning that prediction into useful action is hard. The authors built a single model that couples a pretrained video predictor with an action generator, guided by scene understanding and by geometry and motion knowledge supplied only during training, plus a component that captures task-relevant coming changes without having to imagine future video at run time. Trained on the largest open collection of its kind, it outperformed prior methods in simulation and on real robots.
Technical breakdown
Problem: World Action Models (WAMs) can predict how a scene will evolve visually, but this predictive knowledge does not automatically translate into effective robot control, since action generation additionally needs task-relevant change detection, object geometry/motion understanding, and instruction grounding, without requiring costly future-video rollout at inference time.
Method: InternW0-Δ is a directed World-Action Mixture-of-Transformers (MoT) with 30 MoT blocks coupling a pretrained Wan2.2-TI2V-5B video expert and an ActionDiT action expert: a frozen Wan VAE encodes a sparse visual memory of anchor/recent/current observations, T5 embeddings condition the video expert on the instruction, and a frozen VLM supplies task-conditioned scene semantics to the action expert. "Causal Imprint" tokens are trained (via training-only future supervision) to encode change-oriented predictive features from recent/current observations that feed directly into the action expert without sampling future video at inference, while a separate training-only 4D-aware distillation from a Track4World teacher injects geometric/motion priors into the video expert. The model is pretrained on a curated, canonicalized 80-dimensional state-action corpus of 23,072 hours spanning robot demonstrations (11,302h), UMI data (2,075h), egocentric human demonstrations (4,061h), and Ego2Robot-converted data (5,634h), followed by task-specific post-training on target embodiments.
Key results:
- Zero-shot on LIBERO-Plus (distribution-shift robustness), InternW0-Δ achieves the best overall success rate of 92.8%.
- On RoboTwin 2.0, it reaches 71.9% success on the Clean2Random split and 90.0% on Clean2Clean (81.0% overall for the Clean2Random evaluation table).
- On EBench (mobile bimanual manipulation), it obtains the highest overall score of 66.0 (49.2% overall success rate, 76.5 score on a subset of tasks).
- On RoboDojo (memory, precision, long-horizon tasks), it achieves the best average success rate of 23.9%, nearly double the strongest prior WAM.
- Cumulative ablations on LIBERO-Plus show each component adds gains: adding SMC raises success from 49.59% to 53.47%; adding Qwen3.5-2B raises it to 60.37%, then to 69.08% with RynnBrain1.1-2B; adding Causal Imprint raises it to 70.80%, then further substantially with representation distillation; adding 4D-aware distillation (L4D) raises overall LIBERO-Plus success from 76.45% to 78.37% (Track4World teacher outperforming an alternative teacher's 76.15%).
- Real-robot deployment: pretraining increases task success rates from 20% to 95% in a reported case; the optimized deployment runtime achieves 152.8 ms average round-trip latency on a dexterous-hand platform (NVIDIA RTX 5090), a 5.11x speedup over the standard runtime.
Why it matters / caveats: The paper positions InternW0-Δ as combining the largest reported open-source corpus (20K+ hours) with a design that makes predictive visual dynamics directly usable for action generation without inference-time future-video rollout, and demonstrates transfer of one pretrained checkpoint across gripper-based and dexterous-hand real robots. The authors note limitations: their study of how different egocentric human data sources, conversion strategies, and scaling ratios affect downstream performance remains limited, and their exploration of agent-assisted control is preliminary, covering only a limited corrective setting without systematic study of agent invocation policies or long-horizon hierarchical planning.
Tactile-JEPA: Topology-Aware Self-Supervised Representation Learning for Distributed Tactile Sensors →
Robot skin made of many small touch sensors is usually processed by methods borrowed from images, even though its sensing points are sparse and irregularly placed. The authors pretrain a model to predict hidden sensors' readings from the visible ones, choosing what to hide using the sensors' actual wiring layout at both small and whole-surface scales. It estimated force and in-hand object orientation clearly better than previous approaches across several sensor types and trained more stably and faster.
Technical breakdown
Problem: Self-supervised pretraining methods for tactile sensing focus almost exclusively on vision-based tactile sensors (which output pixel-grid images), leaving distributed electronic skins largely unaddressed even though their taxels (sensing elements) are sparse and irregularly arranged, making direct reuse of visual SSL objectives suboptimal.
Method: Tactile-JEPA is a joint-embedding predictive architecture (following I-JEPA) for distributed tactile sensors: a ViT-Tiny (12 blocks, d=192, 3 heads, MLP ratio 4) context encoder and an EMA-updated target encoder of the same architecture embed per-taxel windowed signals (each taxel's response over a short time window is projected to a d-dimensional embedding via a shared affine projection + LayerNorm), and a 4-block transformer predictor infers masked target-taxel embeddings from context-taxel embeddings plus positional embeddings, trained with an MSE loss between predicted and (stop-gradient) target-encoder embeddings. The key novelty is graph-based, dual-scale mask sampling: masks are drawn over the sensor's taxel connectivity graph (rather than a pixel grid) using local masks (connected subgraphs grown via Dijkstra expansion from a random seed, capturing compact contact regions) and global masks (taxels sampled uniformly at random across the whole graph, capturing overall contact state), with the default target set mixing 2 local and 2 global masks per context.
Key results:
- Evaluated on three datasets with different sensor types/embodiments (Sparsh-skin: magnetic Xela uSkin on an Allegro hand; Tactile socks: piezoresistive knitted fabric on human feet; DECO-50: piezoresistive Inspire hands, Assembly task), Tactile-JEPA reduces in-hand orientation (θ) RMSE by 20.8% and force estimation RMSE by 6.3% versus the strongest prior baseline (BYOL/MAE/DINO variants).
- On Tactile socks, it improves action classification accuracy by 1.28% and reduces full-body pose RMSE by 2.6% over the best baseline.
- On Sparsh-skin object classification, Tactile-JEPA reaches 83.27% accuracy (vs. 83.26% for DINO, 70.61% end-to-end, 64.30% BYOL, 63.44% MAE); on in-hand pose x-accuracy it reaches 95.74% vs. 91.93% (end-to-end) and 90.27% (DINO).
- On DECO-50 visuo-tactile policy learning, Tactile-JEPA achieves the second-best RMSE (0.4634 vs. MAE's 0.4400, both better than vision-only 0.5132), with BYOL and DINO collapsing (failing to train) on this dataset.
- Tactile-JEPA pretrains stably across all three sensor types with no per-dataset tuning, while BYOL collapses on Tactile Socks and DECO-50 and MAE collapses on Tactile Socks; it also pretrains 65-75% faster than DINO across datasets.
Why it matters / caveats: By exploiting known sensor topology (taxel connectivity graph) only during mask sampling — not required at inference — Tactile-JEPA produces topology-aware, multi-scale tactile representations that generalize across sensor types, embodiments, and downstream tasks (force/pose estimation, classification, policy learning) more robustly and stably than adapted vision-SSL baselines. The paper's own ablation shows single-scale masking (pure local or pure global) favors specific downstream tasks, so the mixed local+global design is a deliberate robustness trade-off rather than uniformly optimal for every task (e.g., 4G alone gives lower DECO-50 policy error than the default 2L+2G mix).
Enhancing Photogrammetric Digital Surface Models with Pretrained Diffusion Models and Multimodal Conditioning →
Elevation maps built cheaply from satellite image pairs are noisy and full of gaps, while accurate laser-scanned elevation data is expensive. The authors adapt an existing image-generating model, trimming its text handling and rescaling values so it can work with elevation data, and guide it with both the rough elevation map and satellite photos. Errors in dense urban areas dropped by roughly a third to a half, including in a city never seen during training.
Technical breakdown
Problem: Digital Surface Models (DSMs) produced cheaply from satellite stereo-photogrammetry are noisy and contain voids, while the LiDAR data needed to correct them to high accuracy is far more expensive and limited in coverage.
Method: The authors adapt Stable Diffusion 3 (SD3) for elevation-map refinement by pruning its text stream (cutting the transformer from ~2B to ~1B parameters) while keeping the pretrained image-stream weights, and introduce a patch-wise normalization scheme that rescales/offsets noise (scale s, offset u) to stabilize flow-matching training on locally low-variance, physically-valued elevation rasters. Two separate ControlNets condition the model on the calibrated stereo DSM (CARS) and on Pléiades-HR optical imagery, trained via a Joint Conditional Flow Matching (JCFM) objective. Training data combines LiDAR-HD from 20 French cities with Pléiades-HR imagery and CARS stereo DSMs for 9 of these cities, holding out Bordeaux as a geographically separate test city.
Key results:
- DSM + RGB conditioning reduces Dense Urban RMSE from 6.00 m (calibrated stereo DSM baseline) to 3.45 m in-context, and from 4.16 m to 2.77 m in held-out Bordeaux (a 43% and 33% reduction respectively).
- Pretrained SD3 backbone initialization strongly outperforms training from scratch: FDDINOv2 of 169.4 (pretrained) vs. 692.8 (best scratch learning-rate sweep).
- DSM-only conditioning already improves RMSE across land-cover groups (e.g., in-context Croplands RMSE falls from 6.64 m to 1.04 m), and adding Pléiades RGB further lowers RMSE in nearly every group.
- Pretrained-VAE round-trip diagnostic (encode/decode with no fine-tuning) shows a baseline meter-space bias of −0.004 m, MAE 0.272 m, RMSE 0.573 m, isolating representational distortion from the learned flow's error.
Why it matters / caveats: Shows that visual priors from natural-image diffusion pretraining transfer to a structurally different domain (metric elevation maps) when combined with a simple normalization trick, offering a path to LiDAR-quality DSMs without LiDAR. Caveats acknowledged by the authors: results depend on vertical co-registration and surface-compatibility assumptions that aren't directly verified, Bordeaux Vegetation RMSE is essentially unchanged (refinement can be neutral or slightly harmful there), and the method is only validated within one country/sensor/pipeline (Pléiades-HR + CARS) with no comparison to other enhancement baselines.
Jev in the Wild: A Data-Driven Analysis of the Jev Model's Functionality, Applications and Ecosystem →
A newly released, cheap and fast model that answers questions with choices, yes/no judgments and scores has been adopted quickly, but nobody knew how people actually use it. The authors collected and labelled thousands of public code projects using it, finding most projects use it for judging attributes and scoring, often combining several of its interfaces. They also found public popularity concentrates in a few project types and poorly reflects how widely it is used.
Technical breakdown
Problem: Jev, a newly released fast, low-cost natural-language decision model (offering Choice, Noul, and Score interfaces), has seen rapid public adoption, but it is unclear how it is actually used across applications and how public attention relates to project distribution in its emerging ecosystem.
Method: The authors perform a large-scale, data-driven analysis of 2,170 publicly available Jev projects collected from GitHub as of September 22, 2026, built through a three-stage pipeline: candidate retrieval via keyword/API/SDK-based repository and code search with deduplication by repository ID, project verification (a GPT-6 Luna Max agent confirms each repo has concrete evidence of Jev use for a task, excluding mere mentions or generic wrappers), and annotation (a GPT-6 Luna Max agent labels application domain, decision purpose, and interface usage, with a second agent independently reviewing key inclusion/domain decisions and disagreements resolved by manual inspection).
Key results:
- The ecosystem grew explosively in its first week after release (Sept 15, 2026): 1,865 new GitHub repositories created and 43,750 stars gained, with 305 existing repositories also integrating Jev.
- Across 2,170 projects and 18 subcategories, Attribute judgment is the most common decision purpose (77% of projects), followed by scoring/ranking (52%) and action selection (31%); 69.7% of projects with an identified purpose use Jev for two or more purposes.
- All three interfaces are widely adopted: Choice appears in 81.0% of projects, Noul in 72.2%, Score in 45.4%, and 36.8% of projects use all three.
- Public attention is highly concentrated: Routing & Automation (250 projects, 364 stars/project average) and Interface Agents (175 projects, 271 stars/project average) together make up only 19.6% of projects but receive 63.0% of all stars, while Content & Expert Tasks and Search & Memory are the largest categories (387 and 381 projects) but average only 31 and 37 stars per project.
Why it matters / caveats: The findings suggest Jev functions as a reusable decision component whose role shifts with the surrounding application workflow, and that public GitHub attention (stars) is a poor proxy for actual usage/deployment breadth — a consideration for designing representative evaluations of general-purpose decision models. Limitations stated by the authors: the analysis is a single snapshot (as of Sept 22, 2026) of a fast-evolving ecosystem, and it covers only public GitHub repositories, excluding private repositories and commercial applications that may show different usage patterns.
FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance →
Automatic measures of how different two images look to a person are usually trained on costly, noisy human ratings. The authors exploit the fact that image-generating models build coarse structure first and fine detail last: two images that diverge early in generation look very different, so the moment of divergence serves as a free, numerical difference label. Human studies confirmed the labels match perception, and metrics trained this way outperformed those trained on human-annotated data.
Technical breakdown
Problem: Training reference-based image quality assessment (IQA) metrics requires human-annotated data — either expensive/noisy pointwise mean opinion scores (MOS) or cheaper but only-relative two-alternative forced-choice (2AFC) pairwise labels — and no existing approach produces scalable, globally-consistent pointwise perceptual distance labels without human annotation.
Method: The authors exploit diffusion model generative dynamics: given a reference image x0, they forward-diffuse it to a sampled "forking" timestep ts via xts = (1-ts)x0 + tsε, then denoise the remaining steps to synthesize a perturbed variant, using the forking timestep itself as an automatically-generated, human-annotation-free perceptual distance label (larger ts = more noise injected before divergence = perceptually farther pair). Using FLUX.1-dev as the generator (with S=50 sampling steps, forking step s~U[0,49]), they build a dataset of 480k labeled pairs (240k ImageNet references, 240k FLUX-synthetic references) and train reference-based IQA metrics (CNN backbones LPIPS/DISTS; Transformer backbones DINOv3, CLIP, MAE, DreamSim) with a RankNet-style ranked binary cross-entropy (RankBCE) loss computed over a full B×B pairwise comparison matrix within each batch, rather than triplet-wise 2AFC loss.
Key results:
- Human studies validate the forking-timestep-as-perceptual-distance hypothesis: single-reference ranking study gives Spearman correlation of 0.970 (vs. 0.960 inter-rater correlation) across 30 participants/1,982 responses/190 ImageNet images; cross-reference 2AFC study shows 90.2% individual-response agreement and 92.8% majority-consensus agreement with the forking-timestep label (Fleiss' κ=0.82), rising to 99.5-99.8% agreement when the timestep gap exceeds 30-40 steps.
- On four IQA benchmarks (PIPAL, TID2013, CSIQ, LIVE) across 7 backbone architectures, FoMo achieves the best average SROCC overall (e.g., PIPAL average SROCC 0.683 vs. 0.469-0.484 for BAPPS/PieAPP/NIGHTS/KADID-10k), despite using zero human annotation, and wins by the largest margins on Transformer backbones (e.g., DINOv3 on LIVE: 0.896 vs. best baseline 0.765).
- Ablations show RankBCE clearly outperforms 2AFC-style and L1-regression objectives on the same generated labels (e.g., LPIPS-Alex: 0.733 RankBCE vs. 0.592 2AFC vs. 0.440 L1); the method generalizes across generators (SD-1.5, SD-XL, SD3, FLUX.1) with FLUX.1 performing best, and even at a reduced 10k-pair training budget every generator surpasses the best human-annotated dataset for the CNN backbone (0.622).
- Validation at scale against established perceptual metrics (LPIPS-Alex, LPIPS-VGG, DISTS, DreamSim) as human-judgment proxies on 20,000 pairs shows the forking label achieves SROCC of 0.904-0.932 against these metrics' pooled rankings.
Why it matters / caveats: FoMo shows diffusion generative dynamics can substitute for costly human annotation in IQA and enables training with a globally-consistent pointwise ranking objective rather than only local pairwise preferences, outperforming metrics trained on established human-labeled datasets like KADID-10k and BAPPS. The authors note a key limitation: since the perturbed images are diffusion-generated, metrics trained with FoMo may fail on out-of-distribution/unseen artificial distortions uncommon in nature, though they suggest this could be mitigated by training a domain-specific diffusion generator.
TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations →
Point trackers can either follow a few points through long videos or all points through short clips, because per-frame representations grow with video length. The authors instead maintain a persistent three-dimensional map of the scene, merging duplicate observations of the same surface and computing detailed paths only for points that actually move. This is the first tracker to follow all visible points across videos of over a thousand frames within a single graphics card's memory, clearly beating comparable dense trackers.
Technical breakdown
Problem: Existing point trackers face a hard tradeoff — sparse trackers can follow user-specified query points over long videos, while dense (all-point) trackers can only handle short clips (48-96 frames) because frame-by-frame 2D/3D representations force memory and compute to scale linearly with video length, redundantly re-encoding already-observed surfaces.
Method: TrackEverything represents video as a persistent, de-duplicated 3D scene point cloud in world coordinates rather than per-frame 2D grids. It processes video in sliding windows of L=16 frames, encoding each frame with a frozen DINOv3-small backbone plus a trainable ViT Adapter, unprojecting features into a world-coordinate feature cloud using camera poses/depth (e.g., from VGGT-Ω), and voxelizing tokens to remove redundancy. Tracking is decomposed into an endpoint refiner (a transformer that iteratively predicts each point's 3D position at the window boundary plus a static/dynamic classification, using 3D WAFT — a 3D extension of Warp-Aligned Feature Transform that replaces expensive 4D correlation volumes with bilinear feature sampling at projected source/target locations) and a lightweight trajectory refiner that decodes dense within-window trajectories only for points classified as dynamic. At each window boundary, a voxelization-based de-duplication mechanism merges co-located tracks across origin frames, keeping the representation's size tied to unique scene geometry rather than video duration. The model has 41M trainable parameters (61M total including the frozen 20M DINOv3-small backbone) and is trained on Kubric, PointOdyssey, and Dynamic Replica using 8 L40S-46GB GPUs.
Key results:
- TrackEverything is, to the authors' knowledge, the first 3D tracker able to track all visible points across videos exceeding 1000 frames within 40 GB of GPU memory; prior all-frame dense trackers exhaust GPU memory beyond roughly 96 frames.
- On TAPVid-3D, it outperforms the open-source all-frame dense tracker VDPM by over 20% average APD-P (34.7 vs. 10.9 on 48-frame clips), while remaining competitive with state-of-the-art sparse trackers (e.g., TAPIP-3D: 34.1 avg APD-P) despite tracking far more points, and is competitive with the first-frame dense tracker DeltaV2 (34.7 vs. 34.4 avg APD-P on 48-frame clips; 31.2 vs. 29.9 avg APD-P on full-length videos).
- Ablations: removing the voxelization de-duplication causes the model to run out of memory; removing iterative trajectory refinement drops APD-P from 31.2 to 26.4 (-4.8), and removing 3D WAFT feature sampling drops APD-P to 29.3 (-1.9); a coarser voxel size (0.02) retains ~31 APD-P while roughly halving latency and memory versus finer voxelization (0.005: 0.23 s/frame, 24.0 GB).
- On complexity analysis (PointOdyssey), competing dense baselines (Any4D, SpatialTracker-v2, DeltaV2-dense) run out of memory between ~96-600 frames, while TrackEverything scales to 900+ frames with modest latency and under 30GB peak GPU memory.
Why it matters / caveats: By decoupling representation complexity from video duration and coupling this with all-point dense tracking, the work opens a path toward extending 3D-aware foundation models (VLMs, robotic policies, video generation) to dynamic, long-horizon real-world scenes rather than static/quasi-static ones. Stated limitations: memory still grows during perpetual open-world exploration (no cache eviction yet for hour-long trajectories), performance is coupled to the fidelity of externally-provided 3D geometry/pointmaps, and the irreversible mean-pooling used in voxel-based de-duplication can permanently and erroneously merge distinct physical surfaces that track close together.
AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs →
Tests of teams of AI agents mostly use short or competitive settings, so genuine long-term cooperation goes unmeasured. The authors built a game-world benchmark of human-written tasks needing many rounds and several agents with different roles, plus a measure that traces which actions actually contributed to success. Even the strongest model solved only about half the tasks, and most agent actions were wasted, with failures in communication, role clarity and keeping shared plans.
Technical breakdown
Problem: Existing multi-agent LLM benchmarks mostly test competitive settings, short-horizon interactions under 20 steps, or simply aggregate individual agent performance, so they fail to isolate and measure genuine long-horizon collaboration capability.
Method: The authors build AgentWorld, an MMORPG sandbox supporting up to 1,000 concurrent agents interacting under a "blackbox" protocol (no agent can observe another's internal state), with 380+ items, 144 mob types, 70+ NPCs, and 13 high-level API tools that abstract away low-level control (combat, crafting, navigation) so benchmark scores reflect collaboration rather than motor-control skill. On top of this they define 100 human-annotated tasks (plus 100 LLM-augmented variants) across 8 categories (combat, crafting, gathering, trading, exploration, survival, construction, coordination) spanning 25-55 rounds and requiring 3-20 agents with asymmetric roles. Alongside conventional binary task success rate (SR, via a Python verifier) and partial success rate (PSR, fraction of checkpoints met), they introduce Causal Collaboration Effectiveness (CCE): a graph-based metric that builds a Causal Action Graph by backward-tracing from success actions, using an LLM to make relatively objective binary causal judgments round-by-round, and computes CCE = |C|/|T| (contributing actions over total actions taken by all agents).
Key results:
- Across four frontier LLMs (Gemini 3 Flash, Claude Haiku 4.5, GPT-5 Mini, DeepSeek R1-70B) on the 100-task main set, the best model (Gemini 3 Flash) achieves only 52.0% task success rate, followed by Claude Haiku 4.5 (45.0%), GPT-5 Mini (36.0%), and DeepSeek R1-70B (20.0%); 27% of tasks are unsolved by any model.
- CCE scores are low even for the best model: Gemini 3 Flash CCE = 0.320 (less than a third of all agent actions causally contribute to success), Claude Haiku 4.5 CCE = 0.294, GPT-5 Mini CCE = 0.206, DeepSeek R1-70B CCE = 0.125 (nearly 88% of actions wasted).
- On the harder augmented task set, all models drop substantially (Gemini: 52%→24% SR, Claude: 45%→26%, GPT-5 Mini: 36%→21%, DeepSeek: 20%→10%).
- A controlled baseline comparison (Gemini 3 Flash) shows random actions solve only 5.7% of tasks, single-agent solves 28.6%, no-communication multi-agent solves 22.9%, vanilla blackbox multi-agent solves 52.0%, and an oracle-communication setting (one agent sees all messages) reaches 60.0% — showing communication helps but current models still fall well short of the oracle ceiling.
- Communication quantity does not predict success: GPT-5 Mini sends the most chat messages (44.1/task) yet ranks third in SR, while DeepSeek sends the fewest (7.5) and ranks last; Gemini communicates moderately (11.0 messages) yet achieves the highest SR.
Why it matters / caveats: The results indicate collaboration is a distinct and currently harder capability gap for LLM agents than single-agent reasoning, with systematic failure modes including communication breakdowns, role confusion, and inability to maintain shared plans across rounds (evidenced by the persistent gap between PSR and SR, e.g., 71.5% vs. 52.0% for Gemini). AgentWorld and its sandbox, task verifiers, and evaluation code are fully open-sourced to support further study of this gap.
SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL →
When AI assistants are trained by trial and error, their output mixes tool commands with prose summaries, yet the usual training method gives every word the same credit or blame, so noise from the prose corrupts the tool decisions. The authors split each attempt into tool and summary parts and compute separate credit for each, routing it only to the matching words, at no extra cost. Accuracy improved across benchmarks and model sizes while using fewer redundant tool calls.
Technical breakdown
Problem: In on-policy RL for tool-calling LLM agents, standard algorithms like GRPO broadcast a single trajectory-level scalar advantage to both structured tool-call tokens and user-facing summary tokens, causing "Global Signal Conflation" where reward noise from the summary leaks into (and misattributes credit to) tool-decision tokens.
Method: The paper proposes SLCA-GRPO, built around Segment-Locked Credit Assignment (SLCA), which splits each rollout into tool tokens (ytool) and summary tokens (ysum), computes separately normalized advantages for each segment within the same GRPO rollout group, and routes execution advantages exclusively to tool tokens and preference advantages exclusively to summary tokens (via masking, at no extra rollout cost). This is supported by Hierarchical Rewards (HierR) that supply segment-specific dense execution and summary feedback, and a Schema-Guided LLM Simulator (SGLS) that provides scalable, schema-driven tool-environment simulation for training without costly real API calls. Experiments are run on Qwen2.5-3B-Instruct, Qwen2.5-7B-Instruct, and Qwen3-8B-Base backbones, compared against standard GRPO, ToolPO, and RLTR baselines.
Key results:
- On the 7B backbone: +2.53 pp over matched GRPO on in-domain Toucan, +1.36 pp on the Berkeley Function-Calling Leaderboard (BFCL), and +9.15 pp on τ²-Bench (three-run means).
- Qwen2.5-7B-Instruct reaches 79.13% three-run mean on Toucan-Test success; on τ²-Bench, SLCA-GRPO reaches 70.31±0.50% Overall Accuracy on Qwen3-8B-Base vs. 66.96±0.32% for matched GRPO and 68.50±0.35% for SFT.
- Cross-scale gains: +2.35 pp (3B) / +2.05 pp (8B) on Toucan; +0.40 pp (3B) / +3.35 pp (8B) on BFCL; +1.09 pp (3B) / +10.03 pp (8B) on τ²-Bench.
- Gradient-norm volatility reduced by 25.0% on Qwen2.5-3B-Instruct; tool/summary gradient cosine similarity measured near zero ([-0.03, 0.08]) throughout training, consistent with the advantage-contamination diagnosis rather than gradient-direction conflict.
- SLCA-GRPO achieves higher accuracy with fewer average tool turns (reduced tool redundancy) versus baselines.
Why it matters / caveats: The method offers a zero-extra-rollout, structural (rather than temporal) fix for credit misattribution that composes with temporal credit-assignment methods like VinePPO/GiGPO/SPO, though this combination remains untested. Authors note limitations: it assumes a clear tool/summary segmentation boundary (inapplicable to settings like inline code generation without learned segmentation), routes one scalar advantage to all tool tokens so it cannot distinguish a correct action from a later failing action within the tool segment, and experiments are capped at 8B-scale models with a simulated (SGLS) rather than real-API tool environment.
Game Arena: Strategic LLM Evaluation in Competitive Environments →
Fixed question sets for testing language models are saturating and can leak into training data. The authors built an open platform where models play each other in games chosen to cover full information, hidden information, and many-player social settings, with objective win-loss outcomes and a uniform text interface. Rankings differed markedly across games and cost levels, suggesting games offer a renewable, objective way to measure strategic ability.
Technical breakdown
Problem: Static LLM benchmarks (e.g., MMLU, GSM8K) are saturating and vulnerable to data contamination, while dynamic human/LLM-judge evaluations (Chatbot Arena, MT-bench) are subjective and noisy, so there is a need for a ground-truth, evolving evaluation of LLMs' strategic planning and adaptation under uncertainty.
Method: The authors build Kaggle Game Arena, an open platform where LLMs play head-to-head matches in structured game environments using a uniform text-based harness (each model receives a natural-language state/history description and returns a single action in a prescribed format, with limited retries on invalid outputs). Three pilot environments are released covering distinct information/interaction structures: Chess (perfect information, 2 players, evaluated via Elo/Bradley-Terry ratings on FEN/PGN input with SAN move output and Stockfish-based centipawn analysis), Poker (imperfect information, 2 players, heads-up no-limit hold'em scored in BB/100), and Werewolf (imperfect + asymmetric information, 8 players, natural-language communication, scored via game-theoretic evaluation). Frontier models including Gemini 3 Pro/Flash Preview, o3, GPT-5.2, GPT-5 mini, Grok 4/4.1, Claude 4.5 Opus/Sonnet/Haiku, and DeepSeek V3.2 are evaluated across these environments with bootstrapped confidence intervals for statistical rigor.
Key results:
- Chess Text: Gemini 3 Pro Preview leads with Game Arena Elo 1325, followed by Gemini 3 Flash Preview (1297); o3 (1009) and GPT-5.2 (933) form a second tier; Grok 4 (773), Grok 4.1 Fast Reasoning (632), and GPT-5 mini (525) trail further; Claude 4.5 Opus (236), Sonnet (189), and Haiku (122) rank lowest.
- Poker (heads-up, BB/100): top tier is GPT-5.2 (+46.6), o3 (+29.7), and Grok 4 (+27.1), all well above break-even; the poker benchmark used a large sample scale (900,000 hands cited for the poker benchmark overall).
- Chess evaluation used 40 games per model pair (20 as white, 20 as black); Werewolf evaluation totaled approximately 31,500 games.
- Stockfish-based analysis shows model performance gaps emerge mainly in the middlegame/endgame rather than the opening; weaker models (e.g., DeepSeek V3.2, Claude 4.5 variants) show a steady decline in win probability and an increasing rate of illegal-move "rethink" retries as games progress.
- Cost-performance trade-offs differ by game: Chess and Werewolf share a broadly similar efficient frontier, while Poker reshapes the frontier and reorders which models are cost-effective.
Why it matters / caveats: Because games generate fresh, non-memorizable interactions and yield objective win/loss/chip outcomes, Game Arena is positioned as a saturation-resistant, statistically rigorous alternative to static or judge-based LLM benchmarks, and it is released as an open, extensible harness/dataset. Stated limitations include a fixed (non-adaptive) game-scheduling budget per model pair that is compute-inefficient, no protocol yet for handling model churn (new/retired models) in longitudinal comparisons, and no consolidated cross-game meta-rating since each game currently reports its own incompatible primary metric (Elo, BB/100, game-theoretic win rate).
Do Implicit Personalization and Explicit Styles Conflict? PsPLUG: A Lightweight Plug-in for Balancing Personalization and Style in Customized LLMs →
Language models tuned to an individual user can lose that personal flavour when the user also asks for a specific tone, a failure the authors call personalization collapse. They treat personalization as what remains after accounting for the requested style, learning a small add-on that supplies this residual and a dial for adjusting its strength at generation time. It preserved user traits better than existing methods while following style instructions more closely.
Technical breakdown
Problem: Existing LLM personalization methods suffer from "personalization collapse": when users issue explicit style instructions (e.g., "respond formally," "use a concise tone"), the strong style instruction dominates generation and overrides/collapses the user-specific persona traits that personalization methods are meant to preserve.
Method: The paper reframes personalization as a distributional residual — the divergence between a user's true (user-authored) response distribution and a base LLM's neutral, population-level distribution under the same input — rather than as an absolute output distribution to fit. Built on this view, PsPLUG is a lightweight soft-prompt plug-in that prepends a learned prefix (concatenating a trainable system-instruction vector, a user vector derived from a frozen-encoder-projected user profile, and a query vector) into a frozen LLM's input embedding layer, and is trained with a style-conditioned preference (Bradley-Terry/DPO-style) objective that contrasts user-authored texts against style-conditioned negatives to isolate the pure persona signal from style/instruction effects. An inference-time scaling coefficient α controls the trade-off between personalization strength and style adherence. Experiments use Qwen3-8B as the frozen backbone, with per-user profiles built via a PAG (Profile-Augmented Generation) pipeline and encoded with a frozen BGE-base-en-v1.5 sentence encoder, evaluated on the LaMP benchmark (LaMP-1, -2, -3, -4, -5, -7) against baselines Non-personalized, RAG, PAG, PPlug (Persona-Plug), and OPPU.
Key results:
- Without style constraints (RQ1), PsPLUG achieves the best or near-best scores across most LaMP tasks, e.g., LaMP-1 (Citation ID) F1 0.589 vs. 0.556 (OPPU)/0.493 (PPlug); LaMP-3 (Product Rating) RMSE 0.464 vs. 0.613 (OPPU)/0.583 (PPlug); LaMP-2M F1 0.334 vs. next-best 0.314 (OPPU).
- Under explicit style constraints (RQ2, 4 predefined styles: Warm, Critical, Concise, Elaborative), PsPLUG achieves best or second-best performance in over 80% of style-task-metric settings, outperforming prior baselines by up to 2-4 ROUGE points.
- PsPLUG achieves the highest style-adherence scores (LLM-judge-based) across all four tested styles while maintaining personalization, with "concise" being easiest to control and "critical"/"elaborative" hardest.
- The inference-time scaling coefficient α provides a controllable, continuous trade-off between personalization strength and style adherence (demonstrated via a strength-sensitivity test), and PsPLUG injects fixed-size vectors so inference overhead stays constant and independent of user history size |Pu|, unlike per-user fine-tuning approaches.
Why it matters / caveats: PsPLUG offers scalable, parameter-efficient persona control without per-user fine-tuning, directly targeting a previously under-examined failure mode (style-induced personalization collapse) relevant to production customized-LLM systems. The authors note limitations: style conditions are restricted to four predefined coarse categories rather than fine-grained, open-ended, or compositional real-world stylistic demands; experiments are limited to a specific backbone (Qwen3-8B) within the LaMP benchmark distribution, with robustness across different model architectures, multilingual settings, and long-context domains left unverified; and the work is framed as a preliminary step toward the broader question of disentangling personalization from style.
CARD: Cluster-level Adaptation with Reward-guided Decoding for Personalized Text Generation →
Tailoring a language model to each user either stays shallow or requires expensive per-person training. The authors first group users with similar writing habits and train one small adapter per group, then capture individual differences by contrasting a person's own writing against the group model's output, needing no manual labels. Personal preferences are applied only while generating text via a compact per-user vector, giving better quality than baselines at much lower storage and cost.
Technical breakdown
Problem: Existing personalized text generation methods force a trade-off between fine-grained per-user fidelity and scalable deployment: RAG-based methods are shallow and retrieval-quality-dependent since the generator stays frozen, while PEFT-based per-user fine-tuning captures deeper personalization but scales poorly (expensive per-user parameters, costly onboarding, and scarce high-quality per-user preference data).
Method: CARD is a hierarchical, coarse-to-fine personalization framework with three components: (1) cluster-level adaptation, where users are grouped by embedding-based clustering on shared stylistic patterns and each cluster gets a group-specific LoRA adapter trained on a frozen backbone; (2) an implicit preference learning mechanism that builds preference pairs by contrasting user-authored text against the cluster model's own generations (rather than requiring manual preference annotation), reducing semantic confounding to yield stable individual-style supervision; and (3) decoding-time personalization, where a compact (dimension J=128) user preference vector λu, learned via pairwise preference training, is combined with cluster-level logits through a low-rank logit correction/reward-guided decoding, while both the backbone and cluster LoRA parameters remain frozen. Experiments are conducted on the LaMP benchmark (LaMP-4, -5, -7) and the LongLaMP benchmark (LongLaMP-1/2/3: personalized abstract, topic, and review generation), compared against Non-personalized, RAG, PAG, PPlug, and OPPU baselines, with both automatic (ROUGE) and LLM-judge/human evaluation.
Key results:
- Across 6 tasks and 2 metrics (ROUGE-1/ROUGE-L), CARD ranks 1st in 10/12 settings, with the remaining two near-best (LaMP-5 ROUGE-1 0.459 vs. best 0.464; LongLaMP-2 ROUGE-1 0.252 vs. best 0.255).
- On LLM-judge personalization scores, CARD improves over the non-personalized baseline by 76.4% (LaMP-4), 95.5% (LaMP-5), and 113.5% (LaMP-7); human evaluation shows CARD even exceeds the reference (gold) answer's judged quality by 10.5% on LaMP-5.
- Ablation (Table 2): removing the user vector causes the largest degradation (e.g., LaMP-4 ROUGE-1 drops from 0.218 to 0.148), while removing the group LoRA also degrades performance (LaMP-4 ROUGE-1 to 0.207), confirming both hierarchical components contribute; Group LoRA alone improves LLM-judge scores over non-personalized by 66.7% (LaMP-4), 50.0% (LaMP-5), and 56.4% (LaMP-7).
- Efficiency: per-user personalization requires only a compact J=128-dimensional preference vector λu with the backbone and cluster LoRA frozen, enabling training-free, lightweight per-user storage (can be stored directly on-device) instead of maintaining per-user model parameters.
- User-vector ablations show moderate personalization strength, moderate vector dimensionality, and intermediate hidden-state extraction depth all yield the best performance, with excessive strength/dimensionality causing performance drops from overwhelming semantic content or overfitting/noise.
Why it matters / caveats: By decoupling shared group-level priors from ultra-lightweight individual decoding-time vectors, CARD targets massively scalable personalized deployment without per-user fine-tuning or long-context retrieval, and by editing only logits it avoids exposing raw user history in the context window (privacy benefit) and eliminates long-context latency. Authors note limitations: offline clustering relies on unsupervised K-means, which may not capture complex/latent user relationships; each user is represented by a single (non-interpretable) preference vector, limiting expressiveness for diverse or evolving preferences; the multi-stage pipeline requires coordination across components; and noisy or weakly relevant user histories can degrade the learned user vector, with history filtering left to future work.
TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding →
Tests of systems that understand video as it streams in report a single score without saying when the needed evidence appears, how past frames are kept, or what makes the system speak. The authors built a benchmark with checked evidence timing and a shared protocol that records what each system actually processes, reporting answer quality, timing, false and missed responses, workload and reliability separately. Systems with nearly identical accuracy differed enormously in these other respects.
Technical breakdown
Problem: Current streaming-video-understanding evaluations report task scores (e.g., QA accuracy) without specifying when evidence becomes valid, how a model maintains visual history, or how responses are triggered, so similar scores can mask substantially different workloads, failure modes, and operational behavior across systems.
Method: The paper introduces TRACE, a condition-aware benchmark and evaluation framework built on four components: (1) a temporally audited visual task set with reviewed per-question evidence timing and instruction-dependent proactive trigger annotations; (2) a unified causal "Core-Adapter" protocol, where an Evaluation Core delivers timestamped frames/tasks on a controlled causal timeline and model-specific Adapters connect heterogeneous streaming models/systems while recording actual submissions, history replay, and state; (3) explicit declared execution conditions (visual-state maintenance mechanism, self-initiated vs. externally-triggered response mode, and model-only vs. end-to-end system evaluation boundary); and (4) multidimensional reporting covering answer quality, timeliness (Time-to-First-Token, Response Latency, Median Response Delay), response-selection behavior (False-alarm Rate, Miss Rate), workload (submitted images, output tokens), completion, and reliability (invalid-output rate). The benchmark comprises 1,240 records from 517 videos and is used to evaluate eight publicly available streaming models/systems (AURA, MOSS-VL, MOSS-Preview, LiveCC, ThinkStream, VideoLLM-Online, MiniCPM-O native duplex, and the end-to-end system JoyAI) across eight declared execution configurations.
Key results:
- On QA (833 records), MOSS-VL achieves the highest accuracy at 75.03% (99.88% completion, 374.6 ms median Response Latency, 0.84% invalid output), while LiveCC (65.19%) and MOSS-Preview (65.07%) obtain nearly identical accuracy despite differing completion (93.88% vs. 100.00%) and output-token volume (146,064 vs. 42,206 recorded tokens).
- VideoLLM-Online and MiniCPM-O (native duplex) show near-collapse QA accuracy (2.64% and 2.40% respectively), driven by very high invalid-output rates (95.32% and 93.52%) rather than incomplete execution — for MiniCPM-O, 93.52% of QA outputs could not be recovered as a unique option under the text protocol.
- On Proactive Response, the end-to-end system JoyAI attains the highest In-window Accuracy (17.08%) but also a substantial Miss Rate (32.13%); among model+Adapter configs, LiveCC leads with 12.98% In-window Accuracy but a 65.0% False-alarm Rate, while MOSS-VL and AURA achieve nearly the same In-window Accuracy (~8%) yet differ sharply in False-alarm Rate (41.1% vs. 59.1%) and Miss Rate (47.24% vs. 47.64%).
- Annotation quality control reports a median absolute human revision of 3.0 seconds for QA evidence timing (127 field changes, 90th percentile 58.2s) and 2.24 seconds for Proactive Response triggers (150 field changes, 90th percentile 20.76s).
Why it matters / caveats: The results demonstrate empirically that nearly identical scalar accuracy scores can conceal large differences in completion, answer validity, generation workload, response timing, false alarms, and missed events, arguing that streaming-video performance should be reported and interpreted as execution-conditioned system behavior rather than a single comparable number. Stated limitations: findings are restricted to the tested models/systems on visual-only, single-instruction tasks at 1 FPS; the False-alarm Rate captures only responses emitted when no valid window currently exists (not a general false-positive rate on no-trigger videos); and several models have implementation-specific quirks (e.g., ThinkStream's trailing-block submission, JoyAI's complete-system evaluation boundary, MiniCPM-O's largely unparseable spoken-style output under the text protocol) that limit direct comparability; future work includes adding no-trigger/negative-event coverage, higher frame rates, audio and multi-turn interaction, and sustained-operation tests.
IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking →
Banking assistants must use each customer's account data and sometimes act on it, so grading only the final reply hides errors like using stale information or writing a wrong value. The authors built a test suite of Indian retail-banking cases with mock tools, graded in stages covering safety, actions, reply adequacy and advice, and ran every case three times. No model succeeded on all three attempts for even sixty percent of cases, and occasional success greatly overstated reliability.
Technical breakdown
Problem: Existing tool-calling and agentic benchmarks summarize performance with aggregate accuracy, which hides where and how banking assistants fail (e.g., relying on stale context, missing disputes, or taking incorrect financial actions) and overstates reliability when only occasional (at-least-once) success is reported.
Method: The authors build IndicBankBench, a 799-case benchmark of multi-turn, account-grounded Indian retail-banking interactions spanning five operational domains, a capability/refusal domain, and 20 primary axes, using deterministic mock tools (UPI, NEFT, IMPS, RTGS, IFSC identifiers) instead of live systems. Evaluation runs in four ordered stages — safety, action and tool use, response adequacy, and advisory quality — where tool use and most safety checks are deterministic, a narrow resolver handles ambiguous confirmation-before-write cases, and a separate LLM judge (GLM-5.2 at temperature 0) grades semantic response adequacy. Each case is run three times per model, and the headline metric is strict pass³ (success on all three trials) contrasted with pass@3 (at-least-once success).
Key results:
- Across eleven evaluated models, strict pass³ reliability ranges from 43.7% (Gemma 4 E4B IT) to 58.2% (DeepSeek-V4-Pro-0813), while at-least-once success (pass@3) ranges from 60.1% to 74.3%.
- At-least-once success exceeds strict pass³ by 10.8–21.4 percentage points for every model, over the 799-case benchmark (2,397 total trajectories).
- The wrong-information axis is the lowest-scoring task axis for 7 of 11 models, with no model passing more than 40.1% of its 142 cases on that axis.
- Among 26,291 trajectories where both response-judge and safety/action decisions could be compared, 3,390 (12.9%) passed all safety and action gates but failed the response check, while 1,497 (5.7%) passed the response check but failed a safety or action gate.
Why it matters / caveats: No model exceeds 60% strict reliability, showing aggregate or single-run success metrics substantially overstate dependable banking behavior; the authors note limitations including reliance on an LLM judge for response adequacy (validated on only 40 blinded transcripts), only three repeated trials per case, synthetic mock-tool environments, and English-only Indian retail-banking coverage that does not generalize to other languages, jurisdictions, or production fraud-monitoring systems.
ZooWork-ShopRanker: An Open, Preference-Aligned E-Commerce Reranker →
Search rerankers built for the general web judge topical relevance, but shopping also depends on preferences and hard constraints like budget or intended recipient, and real traffic offers no clean labels. The authors had panels of reasoning models from different families label comparison pairs, trained rerankers on those labels, and distilled the largest into smaller fast versions. The results clearly beat the strongest open reranker and, unlike it, held up on queries with explicit constraints.
Technical breakdown
Problem: Open rerankers trained for general web retrieval optimize topical relevance, not shopper preference, so they fail on e-commerce queries where a candidate violates an explicit hard constraint (product type, budget, intended recipient, or exclusion) but is still lexically/topically similar, and real search traffic offers no clean pairwise preference labels to fix this at scale.
Method: ZooWork-ShopRanker is a family of decoder-style rerankers (0.6B, 4B, 8B) built as LoRA adapters on Qwen3-Reranker bases, scoring query-product pairs via the yes/no logit gap in a single forward pass (no text generation). Training pairs are labeled by a two-family reasoning-LLM judge panel (Qwen3.5-122B and Gemma-4-31B) under a constraint-first, both-presentation-order protocol requiring cross-family agreement, and the ranker is optimized with a pairwise Bradley-Terry/RankNet logistic loss (equivalent to DPO's objective for discriminative scalar scorers). The 8B model is aligned directly on ~4.4k judge-labeled pairs plus retrieval replay data; the 4B and 0.6B are instead trained via distillation, fit with binary cross-entropy to the aligned 8B's soft scores over 134,601 query-document examples and then sharpened with the same pairwise objective on judged pairs. A new benchmark, ShopRank-Bench (~10,511 pairs from 2,991 queries, drawn from private Gensmo e-commerce search traffic, tiered by judge agreement, released in both structured and natural-language text formats), is introduced for contamination-limited evaluation.
Key results:
- ZooWork-ShopRanker-8B and -4B significantly outperform the strongest open reranker baseline (Jina-m0) on ShopRank-Bench in both text formats; on the natural-language gold tier even the 0.6B model beats the much larger Jina-m0 (91.2 vs. 88.8).
- On gold-tier pairs with an explicit hard constraint, all four open baselines (BM25, BGE-Reranker-large, BGE-Reranker-v2-m3, Jina-m0) pick the constraint-violating product on every one of 4 illustrative pairs, while ZooWork-ShopRanker-8B/-4B honor the constraint on all four (0.6B on two); more broadly, open baselines score 14-20 points lower on constrained vs. unconstrained queries (1,843 gold-tier pairs), while ZooWork-ShopRanker is flat.
- Every ZooWork-ShopRanker variant significantly improves over its own Qwen3-Reranker base at all three sizes in both formats (paired McNemar tests, p ≤ 1.3×10⁻⁴); on structured text the distilled 0.6B is statistically indistinguishable from the un-aligned 4B base (−0.9 points, 95% CI [−1.9, +0.2]) while running at 2.8x the throughput of the 4B on 40% of the memory (measured on one H200).
- Two zero-shot reasoning LLMs score 8-9 points above ZooWork-ShopRanker-8B on ShopRank-Bench, but at seconds rather than milliseconds per decision, so are treated as a reference ceiling rather than a deployable alternative; gains also transfer to common MTEB reranking tasks without degrading general reranking quality.
Why it matters / caveats: The work shows preference alignment via LLM-judge-labeled pairs (rather than hand-designed attribute hierarchies, which the authors find anti-correlate with judged preference) can close the constraint-following gap in e-commerce reranking cheaply at inference (LoRA merged before serving, no latency cost); limitations include that explicit budget compliance remains an unsolved, judge-ambiguous "harder and separate problem" not fixed by general preference alignment, and the training corpus itself is proprietary and not released (only the benchmark and models are).
VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models →
Improving robot control models by real-world trial and error is hampered by unreliable value estimates that push the policy off course and by the heavy compute cost of large models. The authors combine fast learning from human corrections early on with value estimates that are gradually calibrated as experience accumulates, plus a system design that reuses unchanging computation to raise training throughput by around an order of magnitude. Across precise chemistry-lab tasks on four robots it succeeded almost always, far exceeding prior methods.
Technical breakdown
Problem: Real-world online reinforcement learning (RL) for post-training large vision-language-action (VLA) models on high-precision manipulation is hampered by two bottlenecks: unreliable value estimates that induce policy drift, and large-VLA computational overhead that limits training throughput and sample efficiency.
Method: VLA-Precision introduces the Asymmetric Co-Bootstrapping (ACoB) algorithm, which couples fast-timescale intervention-guided behavioral learning (actor behavior cloning from successful/human-corrected actions) with slower-timescale progressive value calibration using K critics trained via a TD loss plus a ranking loss on local preferences; calibrated relative action advantages (rather than direct Q maximization) then drive reference-regularized policy improvement to suppress drift. To make ACoB tractable on large VLAs, the paper also proposes ACoB-Stream, a training architecture built on "invariant-state decoupling and on-demand streaming" across four stages (experience-context formation via frozen-VLM-context reuse, context persistence via deduplication and sliding-window sampling, objective-aligned context retrieval, and trainable-subspace policy dissemination for synchronization). The system is evaluated on nine high-precision chemistry manipulation tasks across four robotic platforms (UR5e-DexHand, UR5e-Gripper, Franka-Gripper, Dual UR5e-Gripper), compared against VLA baselines (π0, π0.5) and real-world RL baselines (HIL-SERL, ConRFT, Robo-Dopamine).
Key results:
- VLA-Precision achieves 98.3% mean success rate across nine tasks after 45.8 min average online training per task, with 27.6 s average episode time; it succeeds in 177/180 held-out trials, reaches 100% success on seven tasks, and maintains at least 90% success on every task (worst task 90%, vs. 10% for π0.5 and 5% for π0).
- Relative to π0.5 and π0, it improves average success by 30.5 and 38.9 percentage points and execution speed by 8.7% and 15.9%, respectively; relative to the strongest real-world RL baseline, Robo-Dopamine, it improves average success by 88.3 percentage points and execution speed by 61.2%.
- ACoB-Stream delivers up to 10.9x throughput improvement and reduces mean critic-to-actor (CTA) cycle latency by 90.9% versus a no-KV-cache ablation (M1), in system-efficiency comparisons against four actor-learner baseline variants.
- Ablations show removing actor behavioral cloning causes a 76.25-point success drop, and replacing relative-advantage policy improvement with direct Q maximization causes the largest degradation (87.50-point success drop, 24.93% intervention rate), confirming both components of ACoB's asymmetric co-bootstrapping are necessary.
Why it matters / caveats: The results suggest reliable real-world online RL of large VLAs hinges on resolving an early temporal asymmetry (a capable policy exists before online experience can calibrate the critic), and ACoB's rapid behavioral bootstrapping plus progressively calibrated relative advantages address this; the paper notes a discussion-section limitation that step-wise action deltas degrade sharply as execution horizon grows (autonomous success on insertion tasks fell from 58.3% to 6.7% when horizon increased from 3 to 6, with 29.3% higher action-prediction RMSE), motivating future chunk-wise deltas, and the current evaluation is task-specific rather than multi-task/shared-policy.
Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations →
Studies estimating the effects of climate policy rest on assumptions whose supporting evidence is tedious to check. The authors built a language-model pipeline that audits each paper against a fixed checklist of assumptions and required evidence, and declines to judge when it cannot find relevant text. It caught most deliberately planted flaws, far more than a keyword approach, but on real papers it abstained often and tended to judge more harshly than human annotators.
Technical breakdown
Problem: Difference-in-differences (DID) studies used to evaluate climate policy rest on identification assumptions (parallel trends, no anticipation, correct handling of staggered adoption, limited interference) that a large econometrics literature shows are fragile, yet there is no scalable way to check whether a given paper's reported evidence actually supports these assumptions.
Method: The authors introduce ARGUS, a structured, bounded (non-agentic) large-language-model pipeline with a fixed, deterministic control flow that represents a DID study via an 11-dimension assumption-implication-evidence rubric; for each dimension it retrieves relevant text spans (with a relevance gate), assesses whether the retrieved evidence is adequate, and abstains (emits "unknown") when no suitable evidence is retrieved, rather than guessing. The pipeline mainly uses gpt-4o at temperature 0 for retrieval/judging stages, with cross-model comparisons against Claude Opus 4.8, Gemini 2.5 Flash, and Llama 3.1 8B. Evaluation combines three sources: an 11-flaw injected-error benchmark and an extended 33-variant benchmark (planting known identification flaws into a supported study fixture as local ground truth), a run over 26 real economics papers tagged DID, and a five-paper pilot with reconciled gold labels from two human annotators, plus a pre-specified deterministic calibration rule (applied to fields the LLM already produces) tested against the pilot labels.
Key results:
- On the 11-flaw benchmark, ARGUS detects 73% of planted flaws versus 18% for a keyword-based pipeline; on the extended 33-variant benchmark, a cross-model panel shows near-ceiling detection for several LLMs as judges (gpt-4o 0.89, Claude Opus 4.8 1.00, Gemini 2.5 Flash 0.91, Llama 3.1 8B 0.95), but false-alarm/localization behavior differs sharply across models (Table 5/6).
- Across the 26 real economics papers, ARGUS abstains on about 40% of paper-dimension assessments for lack of retrievable evidence.
- In the five-paper, 55-cell pilot with reconciled human labels, ARGUS abstains on 22 of 55 cells, and among the 33 it answers, it assigns a higher risk level than the human label on 25 of 33 assessments (over-severe), giving low exact agreement (0.24).
- A deterministic calibration rule, fixed before the pilot labels arrived, demotes weak-retrieval "high" risk verdicts to medium or unknown; applying it raises exact agreement from 0.24 to 0.76 (8/33 to 25/33 correct) and cuts over-severe cells from 25 to 8, changing 17 incorrect labels to correct without flipping any correct label to incorrect, though the authors note weighted/chance-corrected agreement stays low and a constant "medium" label would out-score the calibrated system on exact agreement.
Why it matters / caveats: ARGUS flags where DID identification evidence is weak for expert review rather than adjudicating whether the underlying causal effect is real, offering a scalable triage layer for climate-policy causal evaluations; key caveats stated by the authors include a pilot-scale gold standard (only 5 papers, 55 cells, 2 non-expert annotators with only 2 "high"-risk cells), retrieval gaps that limit real-paper coverage and can still let partial matches reach the judge and over-score, an in-sample calibration rule whose generalization is untested, a scope confined to DID designs (transfer to other quasi-experimental methods untested), and no evaluation yet on a dedicated climate-policy paper corpus (current benchmarks use synthetic environmental-policy fixtures and a general-economics corpus).
MOPD-Router: Rethinking Teacher Routing in Multi-Teacher On-Policy Distillation →
When several specialist teacher models train one student, current practice sends each prompt to a single matching teacher, which needs topic labels and discards useful advice from the others. The authors instead let every teacher contribute at each word, weighted by whether its suggested correction reflects the specialty that teacher actually acquired. This worked best in all tested settings, beating both simple averaging on unlabelled data and the labelled-routing standard without using the labels.
Technical breakdown
Problem: Standard multi-teacher on-policy distillation (MOPD) hard-routes each prompt to a single domain-matched teacher for the entire rollout, which fails on real-world mixed datasets lacking reliable domain labels and ignores potentially useful complementary supervision from non-primary teachers at individual tokens.
Method: The paper introduces MOPD-Router, a plug-in framework that routes token-level on-policy-distillation (OPD) supervision over the full teacher pool at every generated token, without domain labels or a separately trained routing model. Within this interface it evaluates three routing metrics: Entropy (routes by teacher next-token confidence), Novelty (routes by accessible teacher-student distributional difference), and the proposed ExpertAlign, which scores each teacher by whether its correction direction to the student at the current token aligns (via cosine similarity, gated to positive alignment) with the specialization direction that teacher acquired during its own post-training relative to a shared pre-RL base. Experiments use Qwen3 teachers specialized in math, code, and instruction-following, tested under strong-to-weak (Qwen3-1.7B student, Qwen3-4B teachers) and same-size (Qwen3-4B student, Qwen3-4B teachers) distillation, on both an unlabeled mixed training set and a domain-labeled MOPD training set, evaluated across nine benchmarks (AIME 2024/2025, HMMT 2025 Feb/Nov, HumanEval+, MBPP+, LiveCodeBench v6, IFEval, IFBench).
Key results:
- On the unlabeled mixed dataset, ExpertAlign achieves the best overall macro-averaged score (38.58 for the Qwen3-1.7B student, 53.76 for the Qwen3-4B student), improving over Mean aggregation (34.49 / 47.88) by 5.88 points (+12.3%) at the reported comparison point, and outperforming label-based baselines Standard MOPD (36.33/48.76) and Open-MOPD (34.02/49.41) despite using no domain labels.
- On the domain-labeled dataset, ExpertAlign reaches overall scores of 40.19 (strong-to-weak) and 54.58 (same-size), exceeding Standard MOPD by 3.95 points (+7.8%) without using the available domain labels; per-domain gains over Standard MOPD range from +2.08 to +7.08 points across math, code, and IF subsets.
- Ablations show token-level routing matters: removing token-level routing (fixed per-response teacher weights) drops the same-size overall score from 54.58 to 52.61, and Entropy-based routing performs no better than naive Mean aggregation, indicating teacher confidence alone is not a reliable routing signal.
- Results hold across paired bootstrap significance tests in all four evaluated settings and are confirmed robust across three training seeds in the same-size domain-labeled setting.
Why it matters / caveats: The results show token-level, metric-based routing can exploit cross-domain complementary supervision and reduce dependence on prompt-level domain labels for multi-teacher distillation, which is directly relevant to industry MOPD pipelines used for LLM post-training capability integration. The authors note limitations: experiments use only three same-family Qwen3 teachers sharing a pre-RL base, so generalization to larger or architecturally diverse teacher pools is untested, the routing metrics are guided by empirical intuition rather than a unified design framework, and extending ExpertAlign to teachers built from different base models is left to future work.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy →
Recordings of people handling objects are a rich source of training data for robot hands, but the hands differ in shape, the motions may be physically impossible for a robot, and simulated skills often fail on real hardware. The authors match hand shapes while preserving where contact occurs, then refine the motion by trial and error in simulation using object position and contact cues, then distil it into a vision-based controller. Contact fidelity and task success improved clearly, and the controller worked on real hardware first try in most trials.
Technical breakdown
Problem: Transferring reconstructed human hand-object interaction (HOI) data to dexterous robot hands is hindered by morphology gaps in kinematic retargeting, dynamical infeasibility of naively retargeted motions, and the sim-to-real gap when training visuomotor policies from such demonstrations.
Method: The paper introduces Morphometric Imitation, a three-stage pipeline: (1) Morphometric Optimization (MMO), an unsupervised kinematic retargeting method that aligns human and robot hand morphology while explicitly recovering/preserving demonstrated hand-object contacts; (2) residual reinforcement learning (RL) in simulation that refines the MMO kinematic reference using object pose and contact information (in observations, rewards, and termination conditions) to produce dynamically feasible, collision-aware robot demonstrations; and (3) distillation of these demonstrations into a point-cloud-based visuomotor policy via a privileged teacher/student imitation learning setup for zero-shot sim-to-real deployment. The method is evaluated on three robot hands (Allegro, Dex3, Sharpa) and ten GRAB hand-object trajectories.
Key results:
- MMO improves location-aware contact F1 over the strongest of five baselines (DexPilot, AnyTeleop, Position, Contact-Aware PyRoki, OmniRetarget) by 8.3, 27.8, and 10.5 percentage points for Allegro, Dex3, and Sharpa respectively, while reducing mean contact patch distance to 9.1/7.5/9.7 mm vs. 18.7/18.9/10.7 mm for the best baselines.
- Using MMO references for downstream residual RL dynamic retargeting yields task success rates (SR) of 82.5% (Dex3), 69.9% (Allegro), and 91.8% (Sharpa), up to 35 points higher than the strongest baseline on each hand.
- Ablations show object pose and contact information are complementary: removing contact information drops SR to 51.2/47.6/76.6% (Dex3/Allegro/Sharpa), and further removing object pose drops it to 42.5/41.6/54.3%.
- The distilled visuomotor policy achieves 93.8% macro-average success in simulation (vs. 91.8% teacher) and 89.3% zero-shot success across 300 real-world hardware trials on 30 objects from 10 categories, with every category achieving at least 80% SR and no failures from hard table collisions.
Why it matters / caveats: The work shows that explicitly modeling morphology and contact throughout kinematic retargeting, dynamic retargeting, and policy distillation substantially improves zero-shot sim-to-real transfer from natural (unconstrained) human motion data. Limitations acknowledged by the authors include training separate policies per object category rather than a unified multi-task policy, evaluation on a single arm-hand system, reliance on only ten motion-capture trajectories (rather than monocular-video reconstructions or bimanual interactions), and remaining gaps in generalizing to a broader range of object poses and shapes.
Paragraph Boundaries Are Not White Space: Compression Depth as the Signature of Hierarchical Structure →
Language models usually treat position as a single reading-order number, which cannot express structure like paragraphs. The authors give paragraph, sentence and word position separate coordinates, then alter only the paragraph coordinate to see how attention across paragraphs changes. Attention is suppressed at paragraph edges, but a randomly labelled control also shows suppression; only the depth of suppression distinguishes real structure, and it varies by text type.
Technical breakdown
Problem: Standard positional encodings represent position as a one-dimensional reading-order coordinate, but reading order alone cannot distinguish hierarchical textual structure (document > paragraph > sentence > token), so it is unclear whether Transformer attention causally responds to hierarchical (e.g. paragraph) position beyond what token distance predicts, and whether any such response is a genuine signature of structure rather than an artifact of having an extra positional channel.
Method: The authors build a hierarchical rotary positional encoding (hRoPE) that assigns paragraph index (p1), sentence index (p2), and token index (p3) to independent, commuting rotary channel groups, enabling counterfactual interventions (fake-merge, fake-split) on the paragraph coordinate p1 while holding the token sequence fixed. They train small Transformers (8 layers, 8 heads, d_model=512, context length 1024, 5000 steps, 3 seeds) on three structurally distinct corpora — WikiText-2, OpenWebText, and a Python source-code corpus (Code) — comparing four positional variants (flat, sent axial, hrope axial, rand axial, plus a period axial mirror control), and measure cross-paragraph attention with a token-distance-exact estimator (Deff, computed by residualizing log-attention against a fitted exponential decay in token distance).
Key results:
- Intervening on p1 alone (tokens fixed) causally changes attention in all three corpora: fake-merge produces ∆B of −1.300 (WikiText-2), −1.302 (OpenWebText), and −3.769 (Code); fake-split produces ∆C of +0.841, +0.914, and +3.373 respectively, with fake-split's normalized response βC ranging from 1.06 to 1.52 across corpora.
- Collapsing p1 to a constant (const0) is the worst condition in every corpus, costing 0.30–0.34 nats more validation loss than the true paragraph coordinate (real), showing sensitivity to p1's content rather than merely its presence.
- All three corpora show attention compression near paragraph boundaries (U < 0) under the exact-distance estimator, with depth U* deepest for Code (≈ −0.772), intermediate for WikiText-2 (≈ −0.481), and shallowest for OpenWebText (≈ −0.305); a density-matched random-label control (rand axial) is also compressed but more shallowly, with hrope axial significantly deeper than rand axial in Code and WikiText-2 (95% CI excludes zero) but not resolvably different in OpenWebText.
- None of eight tested corpus-only quantities across three constructs (lexical persistence, paragraph length, embedding-based coherence) fully reproduces the cross-corpus ordering of compression depth, though embedding-based coherence comes closest.
Why it matters / caveats: The paper argues that compression depth—not compression location or mere presence of compression—is the reproducible, corpus-dependent signature that distinguishes genuine paragraph structure from a matched random coordinate, motivating treating textual position as more than a scalar linear coordinate. Acknowledged limitations include: only three corpora, making cross-corpus ordering descriptive rather than statistically generalizable; non-identifiable or fit-sensitive boundary locations (∆p*1) in two of three corpora; an unresolved real-vs-random compression gap in OpenWebText at the tested seed count; experiments limited to small 8-layer/512-dim models; and no identified corpus-level mechanism that explains the observed depth ordering.
Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching →
Looped language models reuse a block of layers a variable number of times, spending less computation on easy words, but words finishing at different depths cannot be processed together in standard serving systems. The authors build a scheduler that reassembles batches between loop repetitions, handles the stored-state bookkeeping, and predicts one step ahead which words will finish. It captured almost all of the theoretically available speedup, and worked best on models with little machinery outside the repeated block.
Technical breakdown
Problem: Looped language models promise depth-adaptive inference (fewer recurrent-core loops for "easy" tokens, more for "hard" ones), but tokens exiting the loop at different depths cannot share a uniform forward pass and thus cannot be served by standard batching systems like vLLM, so it is unclear whether depth-adaptive decoding can actually be made faster in practice.
Method: The authors present continuous depth batching (CDB), the first end-to-end inference implementation for depth-adaptive looped LMs, which splits the architecture into prelude, shared recurrent core, and coda, and forms new batches between loop steps by removing exiting tokens and (in "refill" mode) filling freed batch slots with new tokens to keep the recurrent batch large. CDB interleaves prefill/prelude/coda batches with loop steps, manages looped KV-caching for tokens with missing states, and uses a lookahead gate to predict one loop step in advance which tokens will exit so batches can be prepared asynchronously without stalling the GPU. They also derive a FLOP-based upper bound on decode speedup and a prefill-corrected, hardware-aware latency model of recurrent-step cost to quantify how close CDB gets to the theoretical maximum.
Key results:
- Evaluated on Ouro 1.4B (fully looped, minimal boundary stages) and Huginn 3.5B (2-4-2 prelude-core-coda split with costly transformer boundary layers) on offline throughput and online serving benchmarks (Alpaca, ShareGPT, ArXiv workloads) on an H100 GPU.
- On Ouro, refill CDB achieves 1.32–1.53× the throughput of full-depth continuous batching (CB) and realizes 96–99% of the estimated end-to-end speedup bound (S_e2e).
- On Huginn, no-refill CDB performs better than refill (due to costly boundary stages), reaching 95–99% of CB throughput with little scheduling overhead, and overall CDB realizes 83–96% of the available speedup.
- Online serving experiments show CDB sustains higher request rates before normalized latency rises sharply, with refill giving the largest serving-capacity gain for Ouro.
- The lookahead exit gate makes early-exit decisions one loop step ahead without reducing accuracy, but constrains minimum exit depth to r_min = 2.
Why it matters / caveats: The work shows depth-adaptive decoding can be served near-optimally (up to 99% of theoretical maximum speedup) rather than being purely a FLOP-savings concept, and offers concrete architecture guidance (keep non-looped boundary stages small, design KV-cache layouts that tolerate missing states) for future looped-LM design. Limitations: evaluation is restricted to two looped architectures on a single H100 GPU, and does not include production serving features such as multi-GPU execution, preemption, prefix caching, or speculative decoding, which could shift CDB's relative benefit in a full serving system.
BoundInk: Boundary-Aware Online Handwriting Generation →
Realistic handwriting generation depends not just on letter shapes but on how letters join, space and align, which previous methods learned only indirectly, producing broken joins and uneven spacing. The authors make the junction between neighbouring letters an explicit thing the model generates, combining local transition modelling with the surrounding sentence, and add measures that specifically assess joins and spacing. Stroke accuracy and all join and spacing measures improved, and blind human raters strongly preferred the results.
Technical breakdown
Problem: Existing writer-conditioned online handwriting generators rely on long-range sequence modeling that captures inter-character connectivity, spacing, and alignment only implicitly, often producing plausible individual glyphs but broken cursive joins, inconsistent spacing, or writer-inconsistent transitions at character boundaries.
Method: BoundInk is a writer-conditioned framework that treats inter-character boundaries as explicit generation targets rather than incidental sequence-decoding outputs. A Style Identifier extracts writer- and glyph-level style from reference handwriting, a Character Context Encoder (built on a shared multilingual text encoder such as CANINE or ByT5, plus a lightweight Transformer for sentence-dependent variation) separates character-identity embeddings from position-dependent sentence context, and a bigram-aware sliding-window Transformer (Bi-SWT) generates each character from a predecessor–current window with gated context fusion, combining local boundary dynamics with global sentence context. The authors also introduce Connectivity and Spacing Metrics (CSM) — including F1_Cursive, CRE, KGS, and SSS — as a dedicated boundary-aware evaluation suite alongside normalized Dynamic Time Warping (DTW).
Key results:
- Across three benchmark-matched protocols (vs. DeepWriting, DSD, and OLHWG), BoundInk improves all applicable CSM measures and reduces normalized DTW by 17.6–47.8% (e.g., 2.70→1.41 vs. DeepWriting, 1.25→0.99 vs. DSD, 0.17→0.14 vs. OLHWG).
- In character/glyph-level SDT-compatible comparisons, BoundInk achieves the lowest DTW among Drawing, DeepImitator, WriteLikeYou-v2, and SDT, reducing English DTW from 1.6048 (SDT) to 1.3649 and Chinese DTW from 0.8789 to 0.8718.
- Blind human preference study (30 participants, 20 trials each): BoundInk is preferred in 615/789 valid criterion-wise judgments vs. DSD (78.0%) and 673/815 valid judgments vs. DeepWriting (82.6%), matching the abstract's reported 78.0–82.6% range.
- Component ablation shows removing Bi-SWT drops F1_Cursive from 0.49 to 0.37 and SSS from 0.71 to 0.40; removing gated context fusion drops F1_Cursive to 0.35; removing both drops F1_Cursive to 0.22, indicating local boundary decoding and sentence context are complementary. Writer-style classification accuracy also improves substantially over baselines (e.g., Top-1 accuracy 6.0→69.0 vs. DSD, 35.8→61.7 vs. DeepWriting).
Why it matters / caveats: By making inter-character boundaries an explicit modeling and evaluation unit (via CSM), the paper argues this improves cursive continuity, spacing, and writer-style fidelity without sacrificing glyph-level trajectory fidelity or content legibility (OCR-based recognition remains competitive with baselines). The authors note future work is needed to incorporate optional sequence references for writer-specific boundary tendencies, generalize to broader scripts and unseen character combinations, improve long-sequence decoding efficiency and robustness to autoregressive error accumulation, and establish standardized benchmarks jointly assessing content, style, boundary quality, reliability, and efficiency.