News · 2026-10-03
OneStreamer teaches a live-video AI to take timestamped notes and wait until it is sure before answering
OneStreamer is a new open AI model for live video that keeps its own timestamped written notes about what it has seen, and learns when to stay silent, when to wait for more evidence, and when to answer. Its 4-billion-parameter version beat every compared system on all eight streaming-video benchmarks its authors tested. The paper was submitted to arXiv on October 1 and ranked first among Hugging Face's papers of the day when checked.
Key facts
- Result: best among the compared methods on all eight streaming-video benchmarks, by the authors' own evaluation.
- When: submitted to arXiv October 1, 2026.
- Who: Xiangyu Zeng and colleagues, with the model and a dataset of over one million training records published by the MCG-NJU group on Hugging Face.
- Primary source: "OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction".
Imagine an AI assistant watching a live camera feed: a kitchen, a workshop, a security view. Something useful happens at minute three, but nobody asks about it until minute twenty. By then the frames from minute three are long gone, because keeping every frame would make the model's input grow without limit. Keeping only recent frames is cheap but forgets. "Streaming video LLMs must retain evidence before its relevance to future tasks is known and respond when sufficient evidence becomes available," the authors write. That is the problem OneStreamer tackles.
How it works
The core idea is that the model writes its memory itself, as text, while it watches. Its "Proactive Hierarchical Caption Memory" produces two kinds of notes: short observations about local details it has just seen, and summaries of events that have finished. Each note carries a timestamp and is added to the conversation history. The model then works from these notes plus a short window of the most recent frames. It is the way a good meeting note-taker works: you cannot replay the meeting, but you can look up what was said at 2:15 and who said it.
The same generation process decides when to talk. At every moment the model can emit one of three signals: stay silent and keep watching, stand by because relevant evidence is emerging but not yet enough, or respond. Training that well is hard because almost every moment in a video is a "keep watching" moment, which can drown out the rare moments when the model should speak. The authors' fix, called Proactive State Transition Learning, keeps every moment that triggers an output and samples a representative set of the waiting moments. They report this beat training on every label "while supervising only 27.5% of annotated state tokens."
The training data is part of the contribution. OneStreamer-1M combines synthetic streaming captions and questions with cleaned open data; questions are timestamped to the earliest moment the video actually supports an answer, so the model learns when it is allowed to know something, not just what the answer is.
What the experiments show
The model starts from Qwen3-VL-4B-Instruct. The most convincing result is the controlled one: with the same model and the same recent-frame window, keeping the written notes "improves historical QA without degrading real-time perception," in the authors' words. On a single NVIDIA H200, a 141 GB card, writing memory notes took less than the one-second interval between incoming inputs in their 360-second replay test.
The OneStreamer-4B weights total about 8.9 GB in bf16. The authors do not state a minimum GPU; by the shipped file sizes, the weights alone occupy about 8.9 GB at that precision, as a floor before the video window and note history are added on top. The dataset card is also public.
Why it matters
Memory is one of the central unsolved problems for AI agents that run continuously, and Ground Truth has covered several takes on it, from organizing an agent's memory to world models that forget. OneStreamer's answer is simple and inspectable: memory is written in plain language with timestamps, so a person can read what the model believes happened. A companion paper the same week, Loci, takes a different route for camera-controlled world models, mixing a compressed running state with a bank of selected old views. Our lessons on agent memory and context windows give the background.
The caveat
All benchmark results are the authors' own, and some competitor numbers were copied from earlier papers rather than re-run, so the comparison table is not one fully controlled experiment. The speed test is a replay of recorded video on a top-end data-center GPU, not a field trial with a live camera and a real user. The text notes keep accumulating, and the paper does not set a hard limit on their total size. A wrong note becomes a durable wrong memory, and the paper does not offer a way to audit or correct notes once written. Practitioner Shawn Wen, interviewed on Machine Learning Street Talk, said of real-time conversation that "it is still hard"; OneStreamer decides when to answer in text, which is a related but simpler problem than spoken turn-taking.
Key questions
How does OneStreamer remember things that left its view?
How big is the OneStreamer model download?
Is OneStreamer a voice assistant?
Cite this
APA
Ground Truth. (2026, October 3). OneStreamer teaches a live-video AI to take timestamped notes and wait until it is sure before answering. Ground Truth. https://groundtruth.day/news/onestreamer-teaches-a-video-model-to-take-its-own-notes.html
BibTeX
@misc{groundtruth:onestreamer-teaches-a-video-model-to-take-its-own-notes,
title = {OneStreamer teaches a live-video AI to take timestamped notes and wait until it is sure before answering},
author = {{Ground Truth}},
year = {2026},
month = {oct},
url = {https://groundtruth.day/news/onestreamer-teaches-a-video-model-to-take-its-own-notes.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.