Ground Truth.
AI, checked against the source.

← All topics

vision-language

Everything on Ground Truth tagged “vision-language” — 3 items.

The cheapest way to teach a model turned out to be blindfolding the student News

Researchers got most of the benefit of an expensive teacher model by degrading the student's view of the image instead of upgrading the teacher, lifting a 4-billion-parameter model past open models nearly sixty times its size with no labels, rewards or larger teacher.

VideoChat3 Halves Video-Model Latency by Compressing Space and Time First News

VideoChat3, an open 4-billion-parameter video model, compresses frames across space and time before the language model, roughly halving latency versus a comparable model.

moondream 3.1 (9B-A2B) Tool

An open-weight vision-language model with 9B total but only 2B active parameters, offering native object detection, pointing, captioning, and segmentation at roughly the speed of a 2B dense model.