AI papers — 2026-07-28
Jump to one of 25 papers
- Kimi K3: Open Frontier Intelligence
- JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents
- From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search
- Progress Reward Modeling for Robotic Learning: A Comprehensive Survey
- StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents
- Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation
- Data Pyramid for Embodied Manipulation
- Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification
- OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation
- The Physics of Multi-Turn Long-Horizon Planning
- Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On
- Chamaileon: Cross-Context Binder Design with Contextualized Modeling and Mixed Sampling
- dRAE: Representation Autoencoder with Hyper-Spherical Codes
- DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes
- ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding
- Codifying the Judge: Scalable Evaluation via Program Distillation
- A Frozen 12B Beats Frontier Models on Verified Work
- GNM Head: A Generative aNthropometric Model of the human head
- Leveraging External Knowledge for Historical Document Restoration via RAG LLMs
- Reasoning Denoiser: Denoising Reasoning Traces for Hallucination Detection
- FilmBench: A Film-Grade Benchmark for Cinematic Video Generation
- Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels
- IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages
- Characterizing Warp Divergence from Pascal to Blackwell
- DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style Identification
Kimi K3: Open Frontier Intelligence →
Technical breakdown
Problem: Open-source models have kept pace with proprietary frontier labs on test-time/RL scaling but have stalled around the 1T-parameter class on pre-training scale, so the gap to the strongest proprietary systems risks widening unless both scaling axes—pre-training size and test-time reasoning/agentic RL—are pushed together.
Method: Kimi K3 is a 2.8T-parameter (104B activated) native-multimodal MoE with a 1M-token context, built from 93 layers arranged in a 3:1 hybrid of Kimi Delta Attention (KDA, a delta-rule linear attention with channel-wise forget gates, lower-bounded sigmoid decay for Tensor-Core-friendly chunkwise computation, and a full-rank output gate) and Gated MLA (NoPE, no explicit position encoding); Attention Residuals (AttnRes) let each of 9 depth-blocks attend over prior block/embedding representations instead of a single residual stream; Stable LatentMoE routes 16 of 896 experts per token (2 shared) through a compact latent width with RMSNorm before up-projection, SiTU-GLU (softcapped SwiGLU) activations, and Quantile Balancing for auxiliary-loss-free load balancing at extreme sparsity; vision is handled by MoonViT-V2, a 27-layer/~0.4B-parameter ViT trained from scratch via next-token prediction rather than SigLIP contrastive init; optimization uses Per-Head Muon. Post-training is a three-stage SFT → domain/effort-specialized RL (general, agentic, coding × {low, high, max} effort = 9 expert policies) → Multi-Teacher On-Policy Distillation pipeline, plus MXFP4/MXFP8 quantization-aware training and EAGLE-3-style draft-model fine-tuning from the pre-trained MTP layer for serving.
Key results:
- Architecture/scaling: ~2.5x improvement in overall scaling efficiency (validation loss vs. FLOPs) over Kimi K2, driven by the combination of KDA/AttnRes/Stable LatentMoE plus refined data/training recipes (layers 61→93, total params 1.04T→2.78T, activated 32.6B→104.2B, routed experts 384→896, active experts 8→16, training context 128K→1M).
- Benchmarks: GPQA Diamond 93.5%, BrowseComp 91.2% (best of all models compared), AutomationBench 30.8%, MCPMark-Verified 94.5%, OmniDocBench 91.1%, Terminal-Bench 2.1 88.3%, ProgramBench 77.8%, SWE-Marathon 42.0 — generally trailing Claude Fable 5 and GPT-5.6 Sol but ahead of Claude Opus 4.8, GPT-5.5, and GLM-5.2 across the suite.
- Third-party: Artificial Analysis Intelligence Index v4.1 ranks #4/580 (57.1); Vals Index #2/39 (74.7); WebDev Arena #1/99 (Elo 1,678).
- Cost efficiency: on Kimi Code Bench 2.0, matches Claude Opus 4.8's max-effort score at ~1/3 the cost; on BrowseComp, best score (91.2%) at $2.03/task — half GPT-5.6 Sol's cost and an order of magnitude cheaper than Claude models at max effort.
Why it matters / caveats: The paper frames Kimi K3 as "the world's first open 3T-class model," pushing pre-training scale and 1M-context agentic RL together rather than trading one for the other, and releases full weights. The authors are explicit that it still trails Claude Fable 5 and GPT-5.6 Sol on research-level reasoning (HLE-Full, CritPt) and some agentic tasks, positioning it as best-in-class among open (and most proprietary) models rather than an outright frontier leader.
JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents →
Technical breakdown
Problem: Existing creative-AI systems (prompt-to-output tools, chat-based agents, node-based workflow tools) only partially preserve the evolving "project state" of long-horizon multimodal creation — references, drafts, alternatives, edits, failed attempts, version relations, and feedback — and closed commercial agent products offer no way to study how agents represent context, choose tools, revise artifacts, or recover from failures.
Method: JarvisHub treats an editable canvas — not chat history — as the shared project state, representing prompts, images, videos, audio, UI components, storyboards, versions, and feedback as typed, addressable nodes and typed directed edges (reference use, version lineage, generation dependency, grouping, workflow continuation) within a graph. A three-layer architecture implements the loop: the canvas state layer stores artifacts/versions/layouts/dependencies; the protocol bridge issues a per-turn capability manifest and execution grant and validates/commits every canvas mutation; and the agent runtime interprets the request against the canvas, selects a permitted action from five tool families (canvas tools, generation tools, native tools, recovery tools, and MCP-backed extension tools), and is supported by reusable skills, cross-turn memory, and subagents for parallel/branching subtasks. Every turn's request, canvas state, manifest, grant, action, observation, feedback, and repair decision is recorded as a trajectory for traceability, recovery/checkpointing, and future training data.
Key results:
- No quantitative benchmark or leaderboard is reported; the paper explicitly frames its evaluation as qualitative case-study demonstrations, not completed benchmarking.
- Three long-horizon creative tasks are demonstrated end-to-end in the same environment: narrative media generation (a short-drama turned into a character/scene/storyboard/shot sequence), interactive web development (a personal photography website with animations), and presentation deck generation (a Stanford-lecture-style PowerPoint deck on decision trees).
- Backend configuration used across all experiments: GPT-5.5 as the agent backend, GPT Image 2 for image generation, Seedance 2.0 for video generation, and Gemini 3.1 Pro as the multimodal evaluation backend.
- For each task the authors present paired qualitative artifacts — a "workspace trace" and the final generated deliverable — as evidence that canvas state remains inspectable and reusable throughout production.
Why it matters / caveats: The contribution is an open, inspectable harness/runtime rather than a stronger generative model — final artifact quality still depends entirely on the external generation/tool backends plugged in, and the protocol bridge guarantees actions are valid and recoverable but not that they are creatively/semantically correct. The authors position recorded trajectories as a future "data flywheel" but note raw trajectories need consent, anonymization, and copyright filtering before use as research data. Code is released at github.com/LYL1015/JarvisHub.
From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search →
Technical breakdown
Problem: Distilling proprietary LLMs into open-source agentic-search policies is blocked because token-level KL distillation needs inaccessible logits and mismatched tokenizers, while raw natural-language trajectory imitation instead transfers the teacher's superficial verbosity and phrasing (style drift/hallucination) rather than its actual reasoning strategy.
Method: MAPD (Multi-Agent Protocol Distillation) separates offline protocol synthesis from online policy alignment. An offline multi-agent system — an Orchestrator (query decomposition), parallel Searcher agents (retrieval over a local Wikipedia corpus), a Repair agent (traces back and re-decomposes failed searches using the ground-truth answer only offline), and a Protocolizer (compiles the exploration log into a JSON protocol with task_type, reasoning_plan, grounding_facts, partial_findings, and answer) — turns each teacher exploration trace into a style-normalized protocol, gated by four automated quality checks (schema validation, EM consistency, extractive-grounding verification, leak detection). During RL training, this protocol is fed only as privileged information to a conditioned "teacher" branch of the same student policy; Online Policy Self-Distillation (OPSD, reverse-KL between the student branch and the protocol-conditioned branch) supplies a dense per-token loss jointly optimized with sparse outcome-based GRPO.
Key results:
- Average success rate across 7 QA benchmarks (NQ, TriviaQA, PopQA, HotpotQA, 2Wiki, MuSiQue, Bamboogle): 39.4% on Qwen3-1.7B and 44.4% on Qwen3-4B, beating the strongest baseline SDAR by 4.8% and 3.3% relative, with larger gains on multi-hop tasks than single-hop.
- Ablations: raw proprietary trajectories without the structured protocol actually underperform even a no-teacher GRPO+OPSD baseline (30.1%/37.3% vs 30.5%/38.3%); structured protocol alone (single teacher, no MAS) recovers most of the gain (37.1%/42.9%); protocol + full MAS gives the best result (39.4%/44.4%).
- Pure OPSD collapses catastrophically in this agentic setting (5.9% on 1.7B, 20.9% on 4B) as self-rollouts drift into pathologically long outputs.
- MAPD transfers across teacher backbones with no retuning: Claude-Opus-4.6 (39.4%/44.4%), GPT-5.5 (39.0%/44.6%), Gemini-3.1-Pro (37.9%/44.0%) — all beat SDAR, within ~2 points of each other.
- Manual review found 99.33% of synthesized protocols were high-quality/compliant; a cost analysis reports 99.94% yield over 25,600 instances at ~6.3 teacher calls and ~12.5K tokens per instance, one-time cost ~$1,454 ($0.057/instance amortized), since protocols are cached and add zero inference-time overhead.
- λ_OPSD sweep shows a sweet spot at 0.05: too small gives weak/short-lived guidance; too large (0.1) induces a "retrieve-without-reasoning" collapse (mean response length drops from 135 to ~42 tokens), hurting multi-hop accuracy.
Why it matters / caveats: The results argue that the main bottleneck in cross-model distillation is the quality/format of the privileged information rather than the training mechanism, and the MAS teacher is used purely offline so it adds no inference-time cost or latency. Caveats: training used only NQ and HotpotQA splits (5 of 7 eval benchmarks are out-of-distribution), models tested are limited to Qwen3-1.7B/4B, and gains are reported via strict exact-match success rate rather than more nuanced answer-quality metrics.
Progress Reward Modeling for Robotic Learning: A Comprehensive Survey →
Technical breakdown
Problem: Robotic learning relies mostly on terminal success signals that say nothing about whether intermediate behavior is advancing, stagnating, or undoing task progress, and the growing literature on "progress rewards" that fill this gap uses inconsistent observations, goal specifications, output signals, supervision sources, and evaluation protocols, making methods hard to compare.
Method: The survey organizes the field into three connected layers rather than proposing a new technique. (1) Interface: progress models are analyzed along current-state representation (single observation, temporal window, relational comparison, or state-access APIs), goal specification (language, vision-goal/demonstration, or structured/programmatic predicates), and output form (scalar score, progress delta, ranking/preference, or executable reward function). (2) Methods: four construction paradigms — frozen foundation-model scoring (e.g., CLIP similarity, GVL/OpenGVL), learning from temporal/relative supervision (goal-proximity, VIP-style distance learning, preference-based reward learning), instruction-tuned progress prediction (fine-tuning VLMs as an explicit progress-following task, e.g., RoboReward, VLAC, ProgressLM), and programmatic reward construction (LLM-generated code/predicates, e.g., Text2Reward, Eureka). (3) Data and benchmarks: pipelines categorized by human involvement (human-driven, human-in-the-loop, fully automated), and benchmarks split into progress fidelity, robustness/generalization, and downstream utility.
Key results:
- Identifies 4 input state-representation types, 3 goal-specification types, 4 output types, 4 method paradigms, and a 3-tier data/benchmark organization spanning roughly 60+ cited methods (GVL, OpenGVL, VIP, LIV, R3M, RL-VLM-F, Eureka, VLAC, Robometer, ProgressLM, SARM, ELEMENTAL among them) evolving from 2017–2026.
- Distinguishes "temporal consistency" metrics from true "scalar calibration," arguing many methods that look accurate under demonstration-only, monotonic evaluation are never tested on whether they track genuine (non-time-correlated) task advancement.
- Notes downstream-utility results (e.g., improved RL success rate) do not by themselves prove a reward is a faithful progress estimator, since policy performance is confounded by exploration, architecture, and environment dynamics.
Why it matters / caveats: As a survey, it reports no new experiments; its contribution is the unifying framework and a public GitHub reading list rather than benchmark numbers. The authors flag four open problems: current models are too coarse-grained (missing subtle physical progress like grasp stability), wrongly assume progress advances at a roughly constant rate, incur high VLM inference latency blocking real-time control, and lack explicit long-horizon memory to disambiguate visually similar states representing different completed subgoals.
StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents →
Technical breakdown
Problem: Screenshots are a lossy, non-injective rendering of a desktop task's real program state (files, DOM, app backends), so perception-driven computer-use agents accumulate compounding misreads over long-horizon tasks and have no reliable way to confirm the final artifact is actually correct and saved.
Method: StateAct is a code-first multi-agent harness with three parts: (1) a main agent (Claude Opus 4.8) that acts through persistent bash, a file editor, and a plan checklist directly on program state, delegating only irreducibly-visual subgoals to a dedicated GUI subagent (1.1% of main-agent steps, 28/108 tasks) or a DOM-based web subagent; (2) an independent "finish gate" that spawns fresh, sees only the task instruction plus read-only machine access, re-derives the persisted result from the actual artifact, and can bounce the agent back for up to 3 correction rounds; and (3) context-sustaining mechanisms — fresh-context delegated subagents, auto-compaction, and an externalized plan checklist — to survive up to ~200 main-agent turns.
Key results:
- On OSWorld 2.0 (108 tasks), StateAct raises Claude Opus 4.8 from 20.6%→26.9% binary success and 54.8%→61.6% mean partial, cutting output tokens 224K→100K and cost ~$72→~$7.8/task (~9x cheaper) versus the same backbone's screenshot-driven reference harness.
- Component ablation: removing act-on-state (GUI back in main agent) causes the biggest drop, partial 61.6%→51.3% (below the 54.8% reference); removing the finish gate drops it to 57.5%; removing context management drops it to 58.7%.
- A bash-only variant (no GUI subagent, no finish gate, no subagents) reaches only 45.9% partial / 15.7% binary — worse than the screenshot-only reference baseline, showing code access alone is insufficient.
- Finish-gate failure analysis: of 79 non-perfect tasks, it correctly rejected only 8 and wrongly passed 68 (~90% miss rate on value/reasoning errors), though it correctly passed 28/29 true binary successes — it bounds structural correctness but not value correctness.
- Root-cause audit of 79 non-perfect tasks: 38 reasoning errors (dominant), 14 verifier-weak passes, 4 ambiguous instructions, 18 capability/modality bottlenecks, 5 undecomposed.
- Generalizes: on Claude Sonnet 4.6, binary success rises 8.3%→11.1%. Swapping the frontier GUI subagent for a compact 31B model barely changes results on 4/5 benchmarks but drops OSWorld 2.0 partial to 43.2% (vs 61.6%), showing a strong GUI subagent still matters on the hardest long-horizon visual subgoals.
Why it matters / caveats: State-grounding shifts the dominant bottleneck in long-horizon computer use from perception to reasoning — the harness fixes how agents observe and verify state but leaves value/reasoning correctness as the binding constraint, and the finish gate's structural-only ceiling is an acknowledged, measured limitation rather than a solved problem.
Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation →
Technical breakdown
Problem: On-policy diffusion distillation (OPD) methods that retain separate teacher/student CFG branches typically match only the CFG-composed velocity, but this loss constrains just a linear combination of the positive/negative branch errors, leaving infinitely many error pairs that satisfy it and letting the two branch errors cancel each other.
Method: The paper shows this branch ambiguity is benign under shared negative conditioning (both errors shrink jointly) but becomes harmful — termed Negative Branch Asymmetry (NBA) — when the teacher's negative branch carries privileged information (e.g., a reference image or dense control signal) unavailable to the student, causing naive matching to shrink the positive-branch error while growing the negative-branch error. The fix, Positive–Direction Matching (PDM), supervises before composition, penalizing the positive-branch error plus a mismatch in the CFG conditional direction, so zero loss forces both branch errors to zero, removing the compensation freedom. Both PDM and a foil baseline (Independent Branch Matching) are applied to dense-to-sparse video control on Wan-VACE (Wan2.1-VACE-1.3B), where the teacher sees dense per-frame pose/depth/scribble control but the student only sparse keyframes.
Key results:
- Text-rendering distillation (shared negative conditioning, no NBA regime): naive and PDM both track the teacher closely across guidance scales, confirming no NBA when conditioning is shared.
- Reference-conditioned distillation (privileged negative branch): naive matching shows antagonistic branch dynamics and visible style drift/distortion, while PDM tracks the teacher's guidance-scale behavior correctly.
- Dense-to-sparse video control at training-scale guidance: PDM gives best/near-best fidelity, e.g. pose MPJPE 4.13 (PDM) vs 4.43 (naive OPD) vs 5.92 (student baseline); depth CORR 91.24 (PDM) vs 90.58 (naive); scribble F1 76.22 (PDM) vs 75.00 (naive).
- Guidance-scale generalization: naive OPD degrades sharply off-scale (MPJPE 4.62→8.98, FVD 67.44→507.70), whereas PDM stays stable (MPJPE 4.27→4.48, FVD 58.65→60.97).
- Ablations: λ=1 gives the best control fidelity for PDM; supervising only 8 of 50 trajectory states gives near-best fidelity at ~132 s/step vs 790 s/step for full-trajectory supervision.
Why it matters / caveats: NBA is a previously unrecognized identifiability failure specific to branch-retaining OPD under privileged teacher conditioning, and its guidance-scale sensitivity signature is the diagnostic the authors propose for detecting it. The authors note PDM's empirical edge over the Independent Branch Matching foil is not yet theoretically explained, since both share the same branch-level zero-loss solution.
Data Pyramid for Embodied Manipulation →
Technical breakdown
Problem: Embodied foundation models need data that couples observations with physical states and actions, but the community's many heterogeneous data sources (real-robot, human video, simulation, web-scale multimodal, etc.) have no systematic taxonomy for comparing their scalability, robot-alignment, and downstream utility, nor a clear account of how existing models actually mix them.
Method: This is a survey/position paper (no new experiments or models). It organizes embodied data into a five-tier "pyramid" ordered by scalability and robot alignment, plus four secondary dimensions (quality, diversity, reusability, physical fidelity): (1) real-robot data (teleoperated/scripted trajectories on the target embodiment), (2) UMI-style data (handheld end-effector demonstrations without a robot in the loop), (3) egocentric/exocentric human video, (4) simulation data, and (5) general vision-language/web data. It then reviews how recent "embodied brain," vision-language-action (VLA), and world-action model (WAM) families compose these sources during pretraining, analyzed through action-space alignment and geometric alignment.
Key results:
- Compiles a comparative table of ~30 real-robot datasets from Pinto & Gupta (2015, 50K trials) through 2026 releases like RoboMIND 2.0 (310K trajectories, 6 embodiments, tactile+force+torque) and Baihu-VTouch (1000+ hours).
- Shows a clear historical shift: early embodied foundation models trained almost exclusively on real-robot trajectories, while recent systems increasingly co-train on real-robot + egocentric + simulation + general data simultaneously, alongside an architectural shift toward world-model-augmented/unified world-action models.
- Frames the pyramid ordering as a trade-off: each layer down sacrifices some robot-executable fidelity in exchange for collection scalability.
- Identifies tactile sensing as a still-missing "contact layer" — only a few datasets incorporate it, and it is not yet standardized in format or coverage.
Why it matters / caveats: As a survey, it contributes taxonomy and cross-model analysis rather than new empirical findings. The authors close with six open challenges: building large-scale tactile datasets, collecting failure/recovery trajectories (most datasets are success-biased), developing scalable/less-teleoperation-dependent collection pipelines, aligning state-action representations across embodiments, leveraging egocentric priors for dexterous-hand policy learning, and designing principled, architecture-aware data recipes.
Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification →
Technical breakdown
Problem: Training-free dynamic sparse attention for video diffusion transformers is bottlenecked by rigid/costly offline block routing (requiring the full proxy-score map to be materialized before sparse computation) and by lossy keep-or-drop sparsification that discards unselected blocks entirely, degrading accuracy at high sparsity.
Method: Sol-Attn derives a query-dependent threshold from the mean and standard deviation of each query block's proxy-score row, computable directly from pooled-key moments without ever forming the full proxy map. During online softmax, each query block streams pooled-key chunks tile-by-tile, comparing scores on-chip to the threshold; blocks that pass go to exact sparse attention, while blocks that fall below threshold are not dropped but approximated via a zeroth-order Taylor expansion around the block's pooled key (proxy-score reuse), with both exact and approximate contributions accumulated into the same online-softmax pass — fusing routing, sparse computation, and approximation correction into one kernel with no HBM materialization of proxy scores or routing indices.
Key results:
- Kernel-level speedup over FlashAttention-3 reaches 5.41× at 128K tokens / 90% sparsity, scaling consistently across 16K–128K sequence lengths.
- On-the-fly routing is 11.5× faster than top-k and 32.7× faster than top-p selection latency, while keeping attention-processor memory near dense (a comparable baseline needs ~8× more memory).
- End-to-end video generation: 2.02× (Wan2.1-14B), 2.12× (HunyuanVideo-13B), 1.9–2.4× (LTX 2.3-22B) speedups at ~85–90% sparsity, matching or beating other sparse-attention baselines on VBench quality and dense-reference PSNR/SSIM/LPIPS.
- Video-to-video: 3.04× speedup on SANA-WM refinement and 2.34× on Bernini-14B editing, both with best quality scores among sparse methods tested.
- Integrated into a full inference engine (with diffusion-step caching + kernel fusion) on NVIDIA B200: 5.08× end-to-end speedup on HunyuanVideo (866.9s → 170.6s) and 3.48× on Wan2.1-14B (563.8s → 161.8s).
Why it matters / caveats: By reusing already-computed proxy scores for both routing and Taylor-expansion-based correction inside a single online-softmax kernel, Sol-Attn removes the traditional separation between (offline, costly) routing and (lossy) sparse computation, narrowing the sparse-vs-dense accuracy gap at a given sparsity/speed point. Authors note the B200 kernel doesn't yet fully exploit Blackwell hardware, currently supports only forward inference (no backward/training), and evaluation is confined to bidirectional diffusion-based visual generation, not autoregressive video generation.
OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation →
Technical breakdown
Problem: Audio and video are typically encoded by independently trained, modality-specific VAEs whose latent spaces have no cross-modal alignment, so a downstream joint text-to-audio-video generator must learn audio-video synchronization from scratch on top of unstructured latents.
Method: OmniVAE keeps separate video (Wan2.2-VAE backbone) and audio (DAC-style continuous VAE) encoder/decoders trained with standard reconstruction losses, but adds two training-only objectives: (1) a segment-level bidirectional audio-video contrastive (InfoNCE) objective that aggregates video/audio latents into matched per-segment features per clip, contrasted against a hierarchical negative pool, aligning the two latent spaces at fine temporal granularity; and (2) a modality-specific semantic distillation objective supervising each branch against frozen Qwen3-Omni vision and audio encoder features via lightweight projectors. Training follows a 3-stage schedule (reconstruction-only pretrain → joint contrastive+distillation training → audio-decoder-only fine-tune to fix artifacts); both extra objectives are discarded at inference, adding zero runtime cost.
Key results:
- Model size: ~1.08B params retained at inference (704.7M video VAE + 371.6M audio VAE).
- Reconstruction is largely preserved: video PSNR 36.27/35.24 (UCF-101/Panda-70M) still beating external Wan2.1/Wan2.2 VAEs (34.5–34.9 PSNR); audio quality remains competitive with or better than other modality-specific audio VAEs tested.
- Audio-video sync probing on VGGSound-Sparse: contrastive learning drives the main gain — top-1 retrieval accuracy jumps from ~6-7% (reconstruction-only/+distillation) to ~18-20% with the contrastive objective, rising to 22.3% with aggregator fine-tuning.
- Downstream text-to-audio-video generation: OmniVAE gives best cross-modal alignment and best audio quality versus reconstruction-only and external-VAE baselines, with the contrastive objective mainly driving alignment/audio-quality gains and distillation mainly improving modality-specific quality.
Why it matters / caveats: Building cross-modal correspondence directly into the tokenizer (rather than injecting it as external conditioning) gives downstream generators a latent space that already encodes audio-video synchronization, improving both alignment and generation quality without added inference cost. Caveats: contrastive alignment causes a slight reconstruction degradation versus reconstruction-only training, and results are only demonstrated at 256×256 resolution with a single downstream architecture.
The Physics of Multi-Turn Long-Horizon Planning →
Technical breakdown
Problem: Because real foundation models are trained on opaque, uncontrollable internet data, it is impossible to pin down where multi-turn long-horizon planning ability actually comes from and how each training stage (pre-training, RL post-training, multi-teacher consolidation) contributes to or bounds it.
Method: The authors build a controllable "planning gym" — three synthetic domains, each a hierarchical AND/OR synthesis graph with a from-scratch model trained on 1.2B tokens — letting them dial task length, trajectory optimality, and planning-pattern/knowledge distribution independently. Stage 1 ablates pre-training data format (CoT world-model state-transition modeling vs. direct action prediction), distribution (atomic-only vs. small injections of long-horizon trajectories), and quality (optimal vs. mixed trajectories). Stage 2 compares multi-turn GRPO against on-policy agentic distillation (OPD), analyzed via mutual-information-motivated separation of "planning pattern" vs. "planning knowledge." Stage 3 studies cascaded multi-teacher OPD (MOPD) across teacher/domain configurations spanning shared-compatible, shared-conflicting, and non-shared-conflicting planning-pattern regimes.
Key results:
- World-model CoT beats direct action prediction: avg@8 gains of +9.4 pp (short), +39.3 pp (middle), +45.8 pp (long-horizon) planning.
- Atomic skills alone don't compose: training on 100% short-horizon data gives 93.1% pass@8 short but ~0% on middle/long; adding just 5% middle-horizon data jumps middle pass@8 from 0.83% to 51.88%.
- Suboptimal trajectories are catastrophic: mixing optimal with suboptimal trajectories collapses long-horizon avg@8 to near 0-1% even though short-horizon performance is largely preserved.
- OPD shows much more stable gradient-direction behavior than GRPO under long-horizon + low-quality settings; in one such regime OPD lifts avg@8 by up to +18.9 pp overall vs. GRPO's -1.6 pp.
- Planning-knowledge mismatch breaks OPD even with a perfect teacher: distilling a different-but-valid recipe drops overall pass@8 below the plain instruct baseline.
- MOPD performs "mode seeking, not mode covering": when patterns are shared and compatible, cross-domain transfer improves substantially; when patterns are entirely non-shared and conflicting, catastrophic forgetting occurs (one domain's avg@8 crashes from 90.00 to 15.62 after training on a conflicting domain).
Why it matters / caveats: The controlled gym isolates causal levers that are normally confounded in real pre-training corpora, giving concrete design rules — inject small amounts of long-horizon and clean data, prefer OPD over GRPO for long-horizon/low-quality regimes, and check pattern compatibility before cascading multi-teacher distillation. Caveats: findings come from a synthetic domain with a from-scratch small model, so scaling to frontier LLMs and real-world tool-use environments is not directly demonstrated.
Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On →
Technical breakdown
Problem: Existing virtual try-on systems are largely confined to a single garment category in controlled studio settings (or repurpose general image editors via prompting, which hallucinates texture/identity), so no open model handles arbitrary items (garments, shoes, bags, jewelry, accessories) with multiple heterogeneous references under real-world conditions.
Method: A dedicated data engine cleans/filters over 50M raw e-commerce, open-domain, and synthetic images down to 10M+, annotating items/subjects and pairing them into item–subject–result triplets via four complementary pipelines plus constrained multi-item pairing rules. Built on a foundation model (Qwen3-VL-8B MLLM + Wan-2.1 VAE + 16B-parameter dual-stream MMDiT, with multi-reference conditioning organized via MRoPE temporal-axis positions), it is trained with a three-stage recipe: continued pre-training (CPT) on a mix of general and try-on data, large-scale supervised fine-tuning (flow-matching loss), and reinforcement learning via DiffusionNFT with a hybrid reward combining an in-house try-on reward model and a rubric-guided Gemini 3.1 Pro judge.
Key results:
- DressCode/VITON-HD (paired): Oxygen-TryOn ranks first on all metrics — DressCode FID 1.932/SSIM 0.936; VITON-HD FID 3.941/SSIM 0.914 — beating the strongest re-evaluated baseline (DressCode paired FID 3.658).
- TStars-VTON single-item: Oxygen-TryOn overall 9.36 vs. next-best re-evaluated system 8.77 (Seedream5 Lite).
- In-house Oxygen-TryOn Bench (1,000 real-world samples): usability rate (shippable results) reaches 86.79% (Cloth-to-Model), vs. best proprietary GPT-Image-2 at 80.35% and best open-source FLUX.2-dev at 67.48%.
- Ablation: SFT lifts usability rate from 67.95%→85.95% (Cloth-to-Model) over the CPT base; RL further raises item consistency and Model-to-Model usability to 85.43%.
- Human evaluation (985 samples, 9 annotators/sample): Oxygen-TryOn best on subject consistency (3.7492) and overall score (3.5502), narrowly ahead of GPT-Image-2 (3.5375).
Why it matters / caveats: It is presented as the first open system delivering any-item, multi-reference try-on at proprietary-model fidelity, with the full data-engine and training recipe disclosed for reproducibility. Caveats: the pretrained base supports at most four reference images, item-fidelity on multi-item composition still trails top proprietary systems, and coordinating five-or-more references under complex layering remains an open problem.
Chamaileon: Cross-Context Binder Design with Contextualized Modeling and Mixed Sampling →
Technical breakdown
Problem: Existing generative protein binder design methods assume a single target and a single conformational state, so they cannot design one binder sequence that stays compatible across multiple conformational states of a target (multi-state) or that selectively binds multiple distinct target proteins (multi-target).
Method: In-Context Complex Co-Design (I3CD) trains a joint sequence-structure flow-matching model by concatenating clean target tokens with noisy binder tokens and using decoupled noise schedules for structure and sequence, so the model learns sequence-structure interplay under a shared target context. At inference, Mixture-of-Paths Sampling (MoPS) handles multiple conformations/targets by alternating, at each discretized timestep, between running full co-design on an "active" conformation and using the resulting sequence to drive sequence-conditioned forward folding on the other conformation(s), iteratively updating one shared sequence against divergent structural constraints; an added beam-search step keeps only the top candidate trajectories at each interval.
Key results:
- Training data: 60,692 filtered PDB chain pairs; trained on 8×A100 GPUs.
- New CROSS benchmark: 100 curated entries (70 high-divergence + 30 low-divergence conformation pairs).
- Full Chamaileon vs. "w/o MoPS" ablation: "both success" (simultaneous success on both conformations) rises from 2 to 7 of 10 — directly showing MoPS fixes the conformational imbalance of a sequential design approach.
- Beam search ablation: adding beam search to MoPS lifts "both success" from 5 to 7.
- Against constructed baselines (RFDiffusion+ProteinMPNN, BindCraft), neither achieved even a single "both success" cross-context binder, while Chamaileon achieved 7/10.
Why it matters / caveats: Chamaileon is presented as a first practical framework and benchmark for "cross-context" binder design, a capability the authors argue is needed for molecular switches, allosteric modulators, and multi-specific therapeutics. Caveat: validation relies mainly on ablations and two self-constructed baselines (not established published methods) on a small, newly curated 100-entry benchmark.
dRAE: Representation Autoencoder with Hyper-Spherical Codes →
Technical breakdown
Problem: Discretizing the high-dimensional, semantically-rich representations produced by vision foundation models causes severe codebook collapse under standard vector quantization, because Euclidean codebook objectives are misaligned with the anisotropic, thin-spherical-shell geometry of these feature spaces, capping usable vocabulary at roughly 16K codes.
Method: Hyper-Spherical Quantization (HSQ) replaces Euclidean nearest-codebook assignment with angular routing — selecting the code that maximizes cosine similarity rather than minimizing Euclidean distance. The codebook loss is likewise rewritten as a spherical (cosine) objective constraining code updates to directional structure, while the commitment loss is deliberately kept Euclidean to preserve magnitude information needed for reconstruction. This HSQ-based tokenizer, built on a frozen SigLIP2 ViT-So400M encoder with a symmetric ViT decoder, constitutes the discrete Representation Autoencoder (dRAE).
Key results:
- Reconstruction: at codebook size 16,384, dRAE reaches 0.69 rFID / 24.03 PSNR vs. a standard-VQ baseline's 1.31 rFID / 22.33 PSNR; scaling to a 131,072 codebook further improves dRAE to 0.42 rFID / 24.52 PSNR.
- Codebook utilization: dRAE maintains >90% global code utilization at K=65,536 without stochastic sampling tricks, vs. the VQ baseline activating only ~6K codes per forward pass.
- Multimodal understanding: dRAE outperforms the VQ baseline and prior unified tokenizers on most benchmarks (GQA 62.4, TextVQA 67.7, MMBench 81.5).
- T2I generation with only 12M training pairs: GenEval overall 0.63 and DPG-Bench 80.58, competitive with baselines trained on far larger data.
Why it matters / caveats: By decoupling angular semantic routing from magnitude-preserving reconstruction, HSQ removes the main scalability wall (codebook collapse) that has limited discrete representation-autoencoder tokenizers, enabling a simplified end-to-end training pipeline that scales to 131K-entry vocabularies with full utilization. The paper notes feature-space reconstruction remains harder than pixel-space due to encoder noise, and T2I evaluation used a comparatively small-scale setup.
DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes →
Technical breakdown
Problem: VLM pretraining data mixtures are built heuristically (stacking quality-filtered datasets, setting cross-domain ratios by intuition) with no principled, attributable way to decide inter-class ratios, intra-class composition, or whether a newly collected dataset should be admitted at all.
Method: DecoupleMix scores every candidate dataset along Quality (LLM-as-a-Judge composite of Accuracy, Hallucination, Correlation, Grammar) and Difficulty (a six-dimension aggregate), then decouples mixture design into two sub-problems: (1) inter-class allocation across capability categories via a single-variable coordinate-style iterative search; and (2) intra-class composition within a category as a constrained convex optimization over allocation weights maximizing Quality, Difficulty, and diversity, solved with an SOCP solver. A third module fixes the ratio/optimizer/budget so admitting a new dataset becomes a single controlled, attributable intervention.
Key results:
- Using 80B additional multimodal continue-pretraining tokens, their 4B model matches a comparable strong open-source instruct model on a 16-benchmark average, before any instruction tuning.
- Attributable admission test: admitting one OCR dataset under their protocol keeps off-target domain drift low and raises overall average, versus naive stacking or displacement, which cause larger drift or domain drops.
- Scaling: outperforms heuristic-stacking baselines at every scale tested; a ratio recipe searched only at a small-token proxy transfers unchanged to a 32B model, beating matched-budget stacking on 13/16 benchmarks.
- Ablations: inter-class-only ratios beat baseline stacking and two published recipes; adding intra-class convex allocation raises the average further, with the largest single-domain gain on OCR.
Why it matters / caveats: The decoupling turns an intractable joint mixture search into a linear-time procedure whose proxy-scale results transfer without retuning to larger data budgets and larger models. Authors note limitations: only tested up to 32B-parameter transfer (not frontier-scale budgets), validation is vision-language only, and quality/difficulty scoring is bounded by the frozen judge model.
ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding →
Technical breakdown
Problem: Medical MLLMs are held back by two vision-centric gaps — a monolithic vision encoder can't absorb the heterogeneous visual knowledge in 2D (X-ray, pathology, fundus) and native 3D (CT/MRI) medical imagery, and current evaluation protocols don't reflect how radiologists actually read and report.
Method: ClinFusion augments a foundational vision-language model's ViT with specialist 2D encoders (ConvNeXt for local/textural detail, DINOv2 for robust semantics) fused via the Cascade Spatial-Aware Locality (CaSL) Fusion operator — an asymmetric cascade of 2D local cross-attention blocks with stochastic residual dropout that progressively enriches the base representation while keeping it anchored to existing vision-language alignment. Native 3D volumes are handled by a dedicated pre-aligned 3D encoder fused via a depth-aware CaSL variant. Evaluation introduces MedIF-Bench (isolating instruction-following from medical accuracy) and an RoI-grounded report-generation protocol that conditions generation on extracted clinical context and buckets claims into matched/missed/hallucinated abnormalities.
Key results:
- Outperforms open-source medical MLLMs on 20/24 benchmarks and proprietary models (GPT-5.2, Gemini-3-Flash) on 13/16 benchmarks.
- 3D VQA: ClinFusion-8B scores 80.2 on AMOS-MCQ, beating comparable open medical baselines by double digits.
- MedIF-Bench Overall-IF: ClinFusion scores 98.1–98.9 vs. GPT-5.2 (96.0), Gemini-3-Flash (96.6), and medical baselines that degrade sharply after fine-tuning.
- Ablations: local cross-attention beats concatenation/global cross-attention/MoE fusion; removing the native 3D encoder drops 3D VQA and 3D report generation scores measurably.
- Blinded evaluation by 6 board-certified radiologists (300 cases): ClinFusion+agentic tools ranks first on Accuracy, Completeness, and Clinical Utility (p<0.001); its RoI-grounded metric had the highest correlation with expert judgment among 11 metrics tested.
Why it matters / caveats: The paper argues perception architecture and evaluation methodology must advance together — a compositional/cascaded vision system plus factualness-driven, context-conditioned metrics correlate better with radiologist judgment than standard text-matching metrics. Authors note remaining gaps versus proprietary models on knowledge-intensive text benchmarks, headroom on some 3D report-generation benchmarks, and the need for human-in-the-loop clinical validation before real deployment.
Codifying the Judge: Scalable Evaluation via Program Distillation →
Technical breakdown
Problem: LLM-as-a-judge evaluation is prohibitively expensive at scale, opaque in its decision logic, prone to systemic stylistic biases, and inflexible — any rubric change forces a full re-run of inference over the dataset.
Method: PAJAMA distills an LLM's judging logic into a committee of synthesized Python programs: an LLM is prompted once (seeded with few-shot examples and a curated evaluation rubric) to write a scoring function, with a similarity filter discarding redundant programs. Each program's outputs are normalized into a discrete vote via a per-program tuned threshold, low-accuracy programs are pruned, and the top-k programs' votes are combined into a joint verdict using a weak-supervision label model that learns per-program reliability weights. A confidence-aware fallback escalates abstained or low-confidence cases to an LLM judge.
Key results:
- Programmatic judges alone average 78.11% accuracy across 5 datasets, matching mid-size open LLM judges and coming within 8 points of a frontier judge, while running ~47-50× faster.
- Confidence-aware routing improves accuracy at 2-3× throughput over LLM-only evaluation, dominating length-based/random routing baselines across 12 LLM judges and 3 model families.
- Reward models distilled from 80 synthesized programs (one-time ~$7 synthesis cost) match or beat those trained on 20,000 GPT-4-labeled pairs (~$300+ cost) at 45-50× lower cost, with higher RewardBench averages.
- On bias robustness (position, rich-content, reference, gender, verbosity), PAJAMA achieves 0% flip rate on position bias (order-invariant by construction) and the lowest average flip rate among all tested judges.
Why it matters / caveats: Programs turn per-sample LLM inference into a one-time synthesis cost, yielding auditable, editable, bias-patchable evaluators with local, API-free execution. Caveats: a gap remains between PAJAMA's internal routing signals and an oracle router, and results depend on the quality/diversity of the curated rubric list and the synthesizing LLM.
A Frozen 12B Beats Frontier Models on Verified Work →
Technical breakdown
Problem: Retraining is the standard way to improve a model (expensive, opaque, non-deterministic), but for problems already solved and verified once, paying a fresh generation pass on every repeat query is architectural waste rather than genuine reasoning work.
Method: The system keeps model weights frozen forever and instead grows a persistent memory of parameterized verified solution methods: a candidate solution must pass a domain-appropriate independent verification gate (machine-checked proof, automated consistency checks, or layered candidate testing) that never consults the benchmark answer key, before being stored once via exact content-hash addressing. At query time the correct stored item is located by exact addressing (not approximate/vector similarity), and the frozen model re-executes it against the new instance's parameters to produce a bit-exact, deterministic answer at zero generation tokens. The paper deliberately discloses no architecture/algorithm details, reporting only measured input-output behavior.
Key results:
- 180/180 (100%) correct on fresh, previously-unseen parameter instances across 9 problem families, at 0 generation tokens per answer, replicated identically across 4 different open models/architectures.
- Negative controls: emptying the memory makes the (unchanged) model solve nothing.
- Exact addressing vs. approximate retrieval on a 4,500-item store: 0% wrong-item rate for exact content addressing vs. 94.3% wrong-item rate for a lossy similarity matcher; selection takes 1.4 microseconds.
- A movable 6,000,000-token working-context window at flat GPU memory on a single 46 GB GPU, vs. common inference engines hard-erroring or silently truncating far below that.
Why it matters / caveats: The system explicitly does not generalize to genuinely novel, never-verified problem families — cold-start/from-scratch reasoning is stated as out of scope, and on public benchmarks the paper concedes frontier models remain clearly ahead of any 12B; the advantage applies only to instances within an already-solved-and-verified family. This is a single-author industry experience report on a proprietary system with architecture/algorithms deliberately withheld and self-measured energy/latency/accuracy figures.
GNM Head: A Generative aNthropometric Model of the human head →
Technical breakdown
Problem: Existing publicly available 3D morphable head models treat the head as a hollow shell, omitting intra-oral structures (teeth, tongue) and fine ocular anatomy, and often suffer from low geometric fidelity due to low-quality training scans.
Method: GNM is a linear-blend-skinning-based 3D morphable model with composite, part-based linear bases: the identity basis is split into head, eyeball, and teeth sub-bases, and the expression basis is split into periocular, lower-face, tongue, and pupil sub-bases. The teeth sub-model is built from thousands of artist-procedurally-generated dental shapes; the tongue sub-model is a PCA over vertex-displacement samples from an artist-made tongue rig; the ocular sub-model is a two-sphere (sclera + cornea) parametric eyeball with a dedicated pupil-dilation expression component. Training data comes from a custom multi-view capture rig (~5,000 subjects, ~150,000 expression samples) plus artist-sculpted head meshes for cranium reconstruction under hair occlusion.
Key results:
- Final topology: 17,821 vertices, 4 joints; 253 identity components and 382 expression components across the sub-models.
- On held-out registered scans, scan-to-mesh distance: GNM 0.748 mm mean vs. FLAME 0.971 mm mean.
- On single-view synthetic image fitting: GNM 1.683 mm mean landmark error vs. FLAME 2.172 mm.
- Generalization/specificity tests show GNM reaches lower error with fewer components than FLAME and its identity manifold stays closer to real scans.
Why it matters / caveats: By unifying skin, eyeballs, teeth, and tongue in one statistical/rig space, GNM enables realistic mouth-interior and gaze reconstruction that FLAME-based pipelines structurally cannot produce, and the authors release the full model publicly. Caveat noted by the authors: the demographic conditioning used for sampling covers only a binary gender split and four broad ethnic categories, which they flag as a non-exhaustive, potential fairness/representation limitation.
Leveraging External Knowledge for Historical Document Restoration via RAG LLMs →
Technical breakdown
Problem: Historical Hanja documents like Korea's Annals of the Joseon Dynasty contain damaged/illegible characters, and prior restoration methods (masked-language-model encoders trained only on the corrupted document itself) fail badly on named entities — proper nouns like person names, places, and titles — because these require external historical knowledge not recoverable from local context alone.
Method: The authors build ARI (Archive Restoration Intelligence), an LLM fine-tuned for character-level Hanja restoration, using a prompt that combines task instructions, temporal metadata, and a RAG component: the top documents retrieved via BM25 from the training corpus, deduplicated by string similarity to avoid near-duplicate bias. Training data uses dynamic masking with a portion of samples preferentially masking named-entity spans to boost NE restoration without hurting random-character accuracy. A separate BERT-based baseline (extended with a Hanja vocabulary, trained from scratch) serves as a non-RAG, non-LLM comparison.
Key results:
- On the named-entity test set, ARI reaches 39.31% NE accuracy / 80.42% random-character accuracy, beating Gemini-2.5-Pro (31.82/73.43), Sonnet-4.5 (29.77/72.52), GPT-5.1 (30.09/69.72), and the BERT-based baseline (28.59/79.38).
- BM25 retrieval beat embedding-based and random retrieval for RAG documents; adding a reranker actually hurt performance.
- In blinded expert evaluation on 100 real-world damaged documents, ARI achieved Top-1 accuracy 0.383 vs 0.273 (Gemini-2.5-Pro) and 0.207 (BERT baseline), and the highest expert win ratio (46.0%).
- Temporal domain-shift testing shows ARI degrades as the gap between reference-corpus period and target-document period widens (accuracy drops from ~39% near-domain to ~21% at the most distant extreme).
Why it matters / caveats: ARI demonstrates that combining LLMs' implicit historical knowledge with explicit BM25-based RAG substantially improves restoration of context-dependent proper nouns in low-resource historical text, offering a practical assistive tool for historians as validated by blinded expert preference. Caveats: performance degrades under temporal domain shift to unseen eras/corpora, the model is text-only despite scanned images being available, and input length is capped, dropping a small fraction of longer documents.
Reasoning Denoiser: Denoising Reasoning Traces for Hallucination Detection →
Technical breakdown
Problem: Long reasoning traces generated by large reasoning models before their final answer contain "irrelevant steps" and "repetitive steps," and both types of noise measurably degrade hallucination detection performance, while existing confidence-based scores and naive embedding filtering fail to separate them from informative steps.
Method: REDE scores each reasoning step by the attention mass the final-answer token assigns to that step's tokens — an annotation-free signal shown to correlate with step informativeness. Rather than thresholding attention directly at inference (shown to be noisy/unstable), REDE uses the top- and bottom-ranked steps as informative/noisy proxy sets to train a lightweight, frozen-LRM projection network over step embeddings, optimized with a three-term contrastive loss (pulling informative steps together, keeping noisy steps from clustering with each other, and pushing the two groups apart). At inference, step embeddings are projected and filtered via k-NN cosine-distance scoring, and the cleaned trace is handed to any downstream hallucination detector.
Key results:
- On TruthfulQA with an 8B reasoning model, filtering raises one detector's AUROC from 68.63→87.32 and another's from 80.42→86.44; similar gains hold across MATH, CodeElo, and MULTIHOPQA benchmarks.
- REDE reaches state-of-the-art among unsupervised methods on TruthfulQA, beating the strongest representation-shaping baseline.
- Direct attention-based filtering (no learned projection) underperforms REDE, justifying the learned-projection design.
- Scales to 32B models with consistent gains over baselines, and transfers across datasets (trained on one benchmark, tested on another, performance holds or slightly improves).
Why it matters / caveats: REDE is a detector-agnostic, annotation-free plug-in that consistently improves multiple hallucination-detection paradigms by denoising reasoning traces rather than proposing a new detector itself. Its comparison baselines (generation-oriented step-selection methods designed for efficiency, not hallucination detection) all fail to improve or even degrade detection, underscoring that signals good for generation efficiency are not the same as signals useful for hallucination detection.
FilmBench: A Film-Grade Benchmark for Cinematic Video Generation →
Technical breakdown
Problem: Existing video-generation benchmarks source prompts from web text/LLM templates and score them with generic multimodal judges against rudimentary taxonomies, so they measure basic video plausibility rather than the professional Cinematic Language craft by which films are actually made and judged.
Method: FilmBench prompts are reverse-engineered via a director-driven pipeline (clip selection from award-winning films across 20 genres → draft prompt extraction → expert refinement by film-school faculty) into slotted, mostly multi-shot shot-list prompts (over 1,000 of ~1,169 total). Evaluation follows a three-level taxonomy co-designed with film-school faculty and a professional studio: 3 top-level axes (Instruction Following, Temporal Continuity, Aesthetic Quality) broken into 12+ components and 35+ sub-metrics. Scoring is done by an in-house automatic evaluation agent built on an open-sourced Cinematic Language operator suite (FilmOps).
Key results:
- The automatic evaluator reproduces expert/human model-level rankings at Spearman rho = 0.95 (text-to-video) and 0.96 (reference-to-video).
- Benchmarking multiple leading video generation models, no model saturates: top overall scores are under 90/100 for both task types.
- Cross-model variance is concentrated in Cinematic Language instruction-following (camera movement, focus, shot scale, viewing angle show the largest spreads), far more than in static aesthetic quality.
- Dynamic-aesthetics bottleneck: action performance, camera-work appeal, character-motion realism, and emotional performance are the lowest-scoring sub-metrics across all models tested.
- Multi-shot prompts are uniformly harder than single-shot: a substantial average score drop moving from single- to multi-shot, widening for weaker models.
- No single model dominates every facet — the overall leader wins only about half the individual sub-metrics.
Why it matters / caveats: By grounding prompts and metrics in actual film-school Cinematic Language rather than generic web-style criteria, FilmBench reveals large, structured gaps (multi-shot staging, dynamic performance/motion) that near-saturated prior benchmarks hide entirely. The authors note FilmBench targets professional cinematic content specifically and its rankings are not claimed to generalize to non-cinematic use cases like short-form or casual video generation.
Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels →
Technical breakdown
Problem: Vision-language models asked to attribute answers via bounding-box coordinates frequently cite the wrong region even when the answer itself is correct ("Attribution Hallucination"), and it is unclear whether this reflects a real grounding-capability gap or an artifact of forcing evidence through coordinate tokens.
Method: The authors keep backbone, document, question, and scoring fixed and only swap the evidence interface: instead of emitting bounding boxes, the model outputs a JSON answer plus verbatim quotes of its evidence; a layout parser segments each page into semantic blocks, a multimodal encoder embeds both quotes and blocks, and a one-to-one assignment resolves each quote to a specific page block. This same quote-and-retrieve pipeline is then reused as a region-label-free RL training scaffold: GRPO trains a model using a reward computed by a vision-language judge that scores the answer against the gold answer and scores crops of the retrieved regions for relevance/coverage, multiplying answer correctness by the evidence scores so evidence only counts when the answer is already right.
Key results:
- On a verified bilingual document-QA subset (six open models, four families), box recall under coordinates never exceeds 8.1% and hallucination rates run 82.1–97.0%; under the language interface recall rises to 25.9–46.9% and hallucination falls to 38.6–65.4%, with little change in answer quality.
- GRPO training on the language-interface pipeline raises an 8B backbone's evidence recall from 39.9% to 51.3% and Strict Attributed Accuracy from 22.4 to 33.8.
- Coordinate recall collapses as region-overlap requirements tighten, while language-interface recall stays nearly flat, showing coordinates lose region geometry, not page identity.
- Ablations show the recovery isn't just granularity or retrieval alone — snapping coordinate boxes to overlapping blocks still performs much worse, and the quote-based retriever beats question-only retrieval and lexical/BM25 resolvers.
Why it matters / caveats: The results suggest open VLMs already "know" where evidence is located and that the coordinate output format itself, not a capability gap, is the main obstacle to reliable attribution — offering a path to better grounding without expensive region-level box labels. Caveats: findings are scoped to single-document attribution, the language interface's ceiling is bounded by the layout parser's own accuracy ceiling, and purely graphical evidence without captions can't be quoted.
IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages →
Technical breakdown
Problem: High-quality multilingual, code-mixed conversational corpora for Indic languages are scarce — existing resources are largely discriminative-task datasets or small, manually-curated dialogue collections limited mostly to Hinglish, leaving no large-scale, event-grounded, multi-turn conversational resource across multiple Indic code-mixed varieties.
Method: The pipeline first converts each source news/blog article into a language-independent English semantic summary (via an LLM chosen after blind human evaluation) to remove document-length/style bias while keeping conversations grounded in identical facts. From this summary, a persona-conditioned dialogue generator autoregressively generates multi-turn conversations under a self-play framework, conditioned on one of five fixed persona pairs (e.g., Friends, Family Members, Colleagues), for both native-script and fully Romanized code-mixed variants. Generated conversations then pass automatic validation requiring script conformity, a minimum code-mixing threshold, and minimum utterance length.
Key results:
- ~1.33 million validated multi-turn conversations across 18 language varieties (9 Indic languages x native-script/Romanized), built from ~142,000 news articles/blogs; over 10.7 million total dialogue turns.
- Code-mixing metrics show strong, measurable code-mixing across all languages, with Telugu/Kannada showing the strongest mixing among native-script variants.
- LLM-as-a-Judge evaluation (9,000 conversations rated) shows most of the 18 varieties score above 4.0/5 on fluency, coherence, engagement, and code-mixing naturalness.
- Human evaluation (180 conversations) shows most language varieties score above 3.0/5 overall, with correlation between human and LLM-judge rankings.
- Corpus generation took roughly 10,000 GPU-hours on 16 GPUs.
Why it matters / caveats: IndicTalk is offered as the largest automatically-generated, event-grounded, persona-conditioned code-mixed conversational resource for Indic languages, intended to enable training/evaluation of multilingual conversational LLMs without manual annotation. The authors note it is a synthetic corpus that may not fully capture pragmatic nuance, dialectal variation, or natural disfluencies, and carries a topical bias toward formal, event-centric news/blog discourse.
Characterizing Warp Divergence from Pascal to Blackwell →
Technical breakdown
Problem: Since Volta's Independent Thread Scheduling (ITS) in 2017, GPU warp-divergence behavior has been widely assumed to be fixed and unchanged across generations, but this has never been tested across Ampere, Hopper, and Blackwell against a pre-ITS baseline.
Method: The authors run cycle-accurate microbenchmarks on a testbed spanning a pre-ITS Pascal GPU, Ampere, Hopper, and both datacenter and consumer Blackwell parts, forcing genuine divergent branches and confirming true divergence via disassembly. They cross-check timing results with hardware performance counters for warp execution efficiency, and separately compile a 38-kernel corpus for static SASS analysis, reconstructing control-flow graphs to classify reconvergence-instruction placement, further validating a newly-found Blackwell barrier-classification field via controlled bit-flip experiments on compiled binaries.
Key results:
- Divergence cost is strictly linear in path count k, with a per-path slope that varies modestly by generation but shows no measurable super-linear reconvergence penalty on any tested GPU.
- Warp execution efficiency tracks 32/k almost exactly across all four post-ITS GPUs tested, and the same law holds on pre-ITS Pascal.
- The full 32-way divergence penalty is occupancy-invariant, staying roughly constant from light to oversubscribed occupancy, because divergence is an issue-rate cost rather than a latency cost that resident warps can hide.
- Predication collapses the 2-way divergence cost to near-zero overhead on all tested GPUs, including pre-ITS Pascal.
- Deferred (later-than-immediate-post-dominator) reconvergence cases fall sharply from Ampere to Hopper to Blackwell as the compiler's reconvergence machinery evolves.
- Blackwell introduces a previously-undocumented two-tier convergence barrier and new uniform-branch/partial-mask-synchronization instructions not present on Ampere or Hopper; bit-flip experiments show one new barrier-classification field has no observable runtime effect, appearing to be a static compiler classification only.
Why it matters / caveats: The programmer-visible cost model of divergence (linear scaling, 32/k efficiency, predication as remedy, occupancy-invariance) is stable from pre-ITS Pascal through Blackwell, so existing performance models transfer across generations. However, the underlying static reconvergence machinery has changed substantially, meaning single-generation binary-analysis tools will mis-model Blackwell control flow unless updated; the authors also note their null-runtime-effect finding for the new barrier field is based on single-warp tests and doesn't rule out effects in untested multi-warp scheduling corners.
DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style Identification →
Technical breakdown
Problem: In naturalistic driving data, driver identity is confounded with vehicle, route, and traffic-condition exposure, so models can appear to recognize a driver's "style" while actually latching onto the car they drive or the roads they frequent.
Method: DriveDNA collects 4,121 drives from 465 drivers across 115 vehicle models (975 hours of driving at 10 Hz, with front video and CAN-derived kinematics), with 276,248 rule-generated, human-audited maneuver-event annotations across six classes. The benchmark defines three core tasks: (1) few-shot driver re-identification from short enrollment windows; (2) personalized behavior prediction (history-to-future acceleration/curvature, comparing personalized vs. generic models); and (3) condition-matched comparison (window pairs matched on vehicle, scenario, speed, and headway). Thirty baselines spanning classical descriptors, time-series encoders, multimodal fusion, and zero-shot foundation models are evaluated under a fixed protocol with explicit vehicle/route/condition leakage probes.
Key results:
- On unseen drivers, the best learned CAN-signal representation reaches AUROC .935 vs. .707 for classical descriptors; zero-shot foundation/LLM models do no better than classical descriptors.
- Under condition-matching (controlling for vehicle/scenario), descriptor AUROC collapses to chance level (.550), while the best learned CAN embedding retains .811 — a driver-specific signal survives controlling for vehicle/scenario confounds.
- Video-only re-identification looks comparable to CAN-based methods (AUROC .937) but exhibits severe route leakage — it predicts the route far above chance — and under condition-matching its AUROC drops to .675, the largest drop of any representation tested.
- Personalization yields small but consistent behavior-prediction gains, but the best re-identification embedding used as a conditioning signal gives no prediction gain, showing re-identification and prediction measure different underlying capabilities.
Why it matters / caveats: The paper's central finding is a three-way dissociation between driver identity, driving style, and predictive utility — high re-identification accuracy (especially from video) can reflect place/route/vehicle recognition rather than genuine behavioral style, so the authors argue any driving-style benchmark must report leakage alongside utility. Caveats: behavioral labels are weak percentile-based labels rather than ground-truth style measures, and driver/vehicle exposure remains naturally entangled since most drivers appear with only one car.