Ground Truth.
AI, checked against the source.

News · 2026-10-03

OneStreamer teaches a live-video AI to take timestamped notes and wait until it is sure before answering

OneStreamer is a new open AI model for live video that keeps its own timestamped written notes about what it has seen, and learns when to stay silent, when to wait for more evidence, and when to answer. Its 4-billion-parameter version beat every compared system on all eight streaming-video benchmarks its authors tested. The paper was submitted to arXiv on October 1 and ranked first among Hugging Face's papers of the day when checked.

Key facts

Imagine an AI assistant watching a live camera feed: a kitchen, a workshop, a security view. Something useful happens at minute three, but nobody asks about it until minute twenty. By then the frames from minute three are long gone, because keeping every frame would make the model's input grow without limit. Keeping only recent frames is cheap but forgets. "Streaming video LLMs must retain evidence before its relevance to future tasks is known and respond when sufficient evidence becomes available," the authors write. That is the problem OneStreamer tackles.

How it works

The core idea is that the model writes its memory itself, as text, while it watches. Its "Proactive Hierarchical Caption Memory" produces two kinds of notes: short observations about local details it has just seen, and summaries of events that have finished. Each note carries a timestamp and is added to the conversation history. The model then works from these notes plus a short window of the most recent frames. It is the way a good meeting note-taker works: you cannot replay the meeting, but you can look up what was said at 2:15 and who said it.

The same generation process decides when to talk. At every moment the model can emit one of three signals: stay silent and keep watching, stand by because relevant evidence is emerging but not yet enough, or respond. Training that well is hard because almost every moment in a video is a "keep watching" moment, which can drown out the rare moments when the model should speak. The authors' fix, called Proactive State Transition Learning, keeps every moment that triggers an output and samples a representative set of the waiting moments. They report this beat training on every label "while supervising only 27.5% of annotated state tokens."

The training data is part of the contribution. OneStreamer-1M combines synthetic streaming captions and questions with cleaned open data; questions are timestamped to the earliest moment the video actually supports an answer, so the model learns when it is allowed to know something, not just what the answer is.

What the experiments show

The model starts from Qwen3-VL-4B-Instruct. The most convincing result is the controlled one: with the same model and the same recent-frame window, keeping the written notes "improves historical QA without degrading real-time perception," in the authors' words. On a single NVIDIA H200, a 141 GB card, writing memory notes took less than the one-second interval between incoming inputs in their 360-second replay test.

The OneStreamer-4B weights total about 8.9 GB in bf16. The authors do not state a minimum GPU; by the shipped file sizes, the weights alone occupy about 8.9 GB at that precision, as a floor before the video window and note history are added on top. The dataset card is also public.

Why it matters

Memory is one of the central unsolved problems for AI agents that run continuously, and Ground Truth has covered several takes on it, from organizing an agent's memory to world models that forget. OneStreamer's answer is simple and inspectable: memory is written in plain language with timestamps, so a person can read what the model believes happened. A companion paper the same week, Loci, takes a different route for camera-controlled world models, mixing a compressed running state with a bank of selected old views. Our lessons on agent memory and context windows give the background.

The caveat

All benchmark results are the authors' own, and some competitor numbers were copied from earlier papers rather than re-run, so the comparison table is not one fully controlled experiment. The speed test is a replay of recorded video on a top-end data-center GPU, not a field trial with a live camera and a real user. The text notes keep accumulating, and the paper does not set a hard limit on their total size. A wrong note becomes a durable wrong memory, and the paper does not offer a way to audit or correct notes once written. Practitioner Shawn Wen, interviewed on Machine Learning Street Talk, said of real-time conversation that "it is still hard"; OneStreamer decides when to answer in text, which is a related but simpler problem than spoken turn-taking.


Primary source, verified: read the paper → (arXiv 2610.01762)

Key questions

How does OneStreamer remember things that left its view?

It writes short timestamped text notes about what it sees and summaries of finished events, and keeps those notes alongside only the most recent video frames, so later questions can draw on old evidence without storing old images.

How big is the OneStreamer model download?

The OneStreamer-4B weights on Hugging Face total about 8.9 GB in bf16; the authors' timing tests ran on a single NVIDIA H200, a 141 GB card, and no minimum GPU requirement is stated.

Is OneStreamer a voice assistant?

No. It watches video and answers in text; the paper does not cover speech output or spoken turn-taking.
Cite this

APA

Ground Truth. (2026, October 3). OneStreamer teaches a live-video AI to take timestamped notes and wait until it is sure before answering. Ground Truth. https://groundtruth.day/news/onestreamer-teaches-a-video-model-to-take-its-own-notes.html

BibTeX

@misc{groundtruth:onestreamer-teaches-a-video-model-to-take-its-own-notes,
  title  = {OneStreamer teaches a live-video AI to take timestamped notes and wait until it is sure before answering},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {oct},
  url    = {https://groundtruth.day/news/onestreamer-teaches-a-video-model-to-take-its-own-notes.html}
}

Topics: research · video · multimodal · agent-memory · streaming · open-weights

Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.