vision-language
Everything on Ground Truth tagged “vision-language” — 3 items.
The cheapest way to teach a model turned out to be blindfolding the student News
Researchers got most of the benefit of an expensive teacher model by degrading the student's view of the image instead of upgrading the teacher, lifting a 4-billion-parameter model past open models nearly sixty times its size with no labels, rewards or larger teacher.
VideoChat3 Halves Video-Model Latency by Compressing Space and Time First News
VideoChat3, an open 4-billion-parameter video model, compresses frames across space and time before the language model, roughly halving latency versus a comparable model.
moondream 3.1 (9B-A2B) Tool
An open-weight vision-language model with 9B total but only 2B active parameters, offering native object detection, pointing, captioning, and segmentation at roughly the speed of a 2B dense model.