Ground Truth.
AI, checked against the source.

← All topics

vision-language

Everything on Ground Truth tagged “vision-language” — 2 items.

VideoChat3 Halves Video-Model Latency by Compressing Space and Time First News

VideoChat3, an open 4-billion-parameter video model, compresses frames across space and time before the language model, roughly halving latency versus a comparable model.

moondream 3.1 (9B-A2B) Tool

An open-weight vision-language model with 9B total but only 2B active parameters, offering native object detection, pointing, captioning, and segmentation at roughly the speed of a 2B dense model.