Ground Truth.
AI, checked against the source.

← All topics

multimodal

Everything on Ground Truth tagged “multimodal” — 56 items.

DeepSeek releases a 168 GB MIT-licensed multimodal V4 checkpoint News

DeepSeek's V4-Flash-Vision-Exp is an MIT-licensed 168 GB downloadable multimodal model whose strongest comparisons remain vendor results under DeepSeek's own harness.

MiniMax H3 and fal turn video generation into a live prompt loop News

MiniMax H3 produces short video with native stereo sound, while fal's H3 Max Director API is built for continuous real-time streams with live prompts—together enabling an interactive broadcast format where the next scene can be steered while viewers watch.

Meta's Muse Spark 1.3 caught GPT-5.6 on one scoreboard and still trails Claude News

Meta released Muse Spark 1.3 on September 2, 2026, and Artificial Analysis scored its public tier at 61 on its Intelligence Index, level with OpenAI's GPT-5.6 Sol, while the same measurement puts Anthropic's Fable 5.1 four points ahead of Meta's best variant.

DeepSeek gave its cheapest model eyes and did not change the price News

DeepSeek shipped an experimental vision version of its V4-Flash model that accepts images by base64, URL, or file upload, and bills it at exactly the same rate as the text-only model.

Gemini Omni 1.1 Flash can extend a scene instead of restarting it News

Google's updated video model reads up to ten seconds of a clip's prior context before continuing it, up from one second, and adds keyframe control, cheap 360p drafts and 4K upscaling through the Gemini API.

GLM-5.3-Flash was Ox Alpha, and it ran on Chinese chips News

Z.ai released GLM-5.3-Flash under an MIT licence and confirmed it is the anonymous \u201cOx Alpha\u201d model that topped OpenRouter for a week -- served, the company says, entirely on a cluster of Chinese AI accelerators at per-token cost comparable to NVIDIA hardware.

SenseNova generates images in pixels with no VAE in the middle News

SenseTime released SenseNova-U1.5-8B-MoT under Apache 2.0, an open-weight model that understands and generates images without the vision encoder and latent autoencoder nearly every other system relies on, working directly on pixels instead.

DeepSeek gave its cheap model eyes, then capped them at 384 tokens an image News

DeepSeek released deepseek-v4-flash-vision-exp, an experimental multimodal version of its cheapest model that accepts images directly in the same agent loop as text, but budgets each image to at most 384 tokens after resizing toward roughly 800 by 800 pixels.

Video models look right 92 percent of the time and do the task 38 percent News

A new benchmark scores AI video generation on two separate axes and finds the best model reaches 91.8 percent on visual reliability but only 37.8 percent on actually completing the instructed task, quantifying a gap that fidelity metrics have been hiding.

Frontier multimodal models still cannot build a 3D world, and a new benchmark says under 60 percent News

VibeWorlding tests whether multimodal agents can turn a plain request into an interactive 3D scene end to end, and finds that frontier models including GPT-5.5 and Qwen3.8-Max succeed on fewer than 60 percent of tasks.

MiniMax put its video-with-sound model on Hugging Face News

MiniMax released the weights for H3, a multimodal generation model that produces video with native stereo audio, two weeks after promising them at launch, and it has already been downloaded more than two million times.

The open video model tops out at fifteen seconds, not twenty-six News

MiniMax's H3 weights have drawn 2,900 stars in four days, but the company's own repository caps a single generation at fifteen seconds and says the hosted component it left out is critical to output quality.

Mistral Shipped an Open-Weight Safety Judge That Takes Its Policy as a Question News

Mistral released Shieldstral 1.0 3B, an Apache-2.0 multimodal moderation model that reads a plain-language yes/no policy question at inference time instead of a fixed harm taxonomy baked into its weights, and runs on a single 16GB GPU.

MiniMax Shipped H3's Weights and Kept the Best Part Hosted News

MiniMax released the weights for its H3 video-and-audio generation model, and its own model card says the input-processing stage that is critical to output quality is not included in the release and the 2K output stage is not open-sourced at all.

Vision Transformers: what happens when you feed a picture to a language architecture Lesson

A Vision Transformer chops an image into a grid of small patches, treats each patch as a word, and runs the exact same Transformer machinery that powers language models over the resulting sequence. Google Research showed in 2020 that this beats purpose-built image networks once you train it on enough data, and it is why today's image, video and robot models all share one architecture.

Thinking Machines ships Inkling-Small's open weights - all 532 gigabytes of them News

Thinking Machines has published the full weights for Inkling-Small, a 276-billion-parameter sparse model that activates only 12 billion parameters per token and accepts text, images and audio, under an Apache 2.0 licence with a separate use policy attached.

Microsoft lets the video codec pick which pixels the model sees News

Microsoft's Mage-VL reuses a video file's own compression decisions to choose which image patches a vision model processes, cutting visual tokens by over 75% and reporting up to a 3.5x speedup over uniform frame sampling.

Humans score 96% on a new visual exam. The best model gets one in ten. News

A new benchmark called ActiveVision asks models to keep re-examining a picture while reasoning through it, and GPT-5.5 solved 9 of 85 problems at its highest reasoning setting while unaided humans averaged 81.7.

TimeLens2 teaches video AI to point to the exact seconds that answer a question News

Researchers released TimeLens2, an open-weight video model fine-tuned to answer a text query by returning the exact timestamp intervals in a video that contain the evidence, using a new distance-sensitive reward and a carefully curated dataset, with small versions reported to beat much larger open models on temporal grounding.

VideoChat3 Halves Video-Model Latency by Compressing Space and Time First News

VideoChat3, an open 4-billion-parameter video model, compresses frames across space and time before the language model, roughly halving latency versus a comparable model.

Boogu-Image-0.1: a fully open image model that claims to close in on closed systems for about $400K News

Boogu-Image-0.1 is a fully open-source unified image generation and editing model family whose researchers say a base model reaching near-frontier quality cost roughly $400,000 to train, arguing the closed-open gap is closing through data and pipeline quality rather than raw compute scale.

Thinking Machines releases Inkling, now the top-ranked US open-weights model News

Thinking Machines Lab released Inkling, a 975-billion-parameter open-weights model under Apache 2.0 that Artificial Analysis ranks as the strongest open-weights model from any US lab, scoring 41 on its Intelligence Index.

A benchmark audit finds most video-understanding tests can be aced without watching the video News

Video-Oasis audited video-understanding benchmarks and found about 55% of samples are solvable with no visual input at all - models exploit linguistic priors instead of watching motion, and once the shortcuts are removed, state-of-the-art systems barely beat random guessing.

How AI Turns Speech Into Text Lesson

Automatic speech recognition (ASR) converts spoken audio into written text by breaking sound into tiny slices, encoding them into features a model understands, and decoding those into words -- and its accuracy is measured by word error rate, the fraction of words it gets wrong.

Vidu S1 generates video you can steer with your voice in real time News

A new paper introduces Vidu S1, a video model that generates interactive 540p video at up to 42 frames per second on consumer GPUs and lets users reshape the scene on the fly with voice commands, without the drift that usually breaks long AI video.

Google's Gemma 4 is a small open multimodal family that skips the image encoder News

Google released Gemma 4, an open-weight model family from 2.3 to 31 billion parameters that natively handles vision and audio, including a 12-billion-parameter variant that ingests raw image and audio patches with no separate encoder.

Why AI Vision Benchmarks Reward Getting Close Instead of Getting It Right News

A new evaluation method argues standard image benchmarks hide model failures by averaging all details equally, and it exposes an 8-point perception gap between open and proprietary models that looser scoring conceals.

What Are Vision-Language-Action Models? Lesson

A vision-language-action (VLA) model is a single neural network that takes in camera images and a plain-language instruction and outputs the actual motor commands to carry it out, letting one model both understand a scene and physically act on it.

Orca proposes a single 'world latent space' to replace next-token, next-frame, and next-action prediction News

Researchers introduced Orca, a world foundation model that learns one unified latent space from multimodal signals and predicts the next world state rather than the next token or frame, outperforming similar-sized specialists on text, image, and action tasks.

Image generators can't plan. This one bolts on a brain that can. News

Qwen-Image-Agent wraps planning, reasoning, and memory around a text-to-image model so it can break a hard request into steps - and the local-AI crowd immediately asked whether it runs on a gaming GPU.

One model that listens, sees, and talks back in real time News

Wan-Streamer collapses the usual chain of separate speech and video tools into a single model built for live, two-way conversation.

This model's job is to make better training data for other models News

DataClaw0 turns the grind of cleaning and labeling training data into a learned skill -- a small model that refines raw, messy multimodal streams into dense, purpose-built lessons.

fal H3 Max Director Tool

fal's API for continuous AI-generated video streams with live prompts and chunked playback controls, built around MiniMax H3 Max Director.

Wan Streamer v0.3 Tool

Streaming audio-visual interaction model that treats a video as a persistent world plus a time-varying event stream, running full-duplex real-time conversation at 640x368 / 25fps with roughly 550ms total interaction latency.

VideoChat3-4B Tool

A fully open 4B-parameter video multimodal model for general, long-form, and streaming video understanding, released with weights, training code, training strategy, and datasets.

Shieldstral 1.0 3B Tool

Mistral's open-weight multimodal moderation model. You supply the policy as a plain-language yes/no question at inference time rather than retraining for a fixed harm taxonomy, and it returns one calibrated safety score per forward pass. Handles prompts, responses, prompt-response pairs, images and image-plus-text across twelve languages. Apache 2.0, runs on a single 16GB GPU via vLLM, Transformers or llama.cpp; recommended operating context is 32k tokens.

Seed2.0 (ByteDance Seed) Tool

ByteDance's Seed2.0 family (Pro, Lite, Mini) of closed, API-hosted models aimed at long-tail knowledge and complex instruction-following, accessed through ByteDance's Volcano Engine (Ark) platform. Not open weights despite the academic-style model card.

Qwen3.8-Max-0902 Tool

Alibaba Cloud's 2.4T-parameter MoE flagship with native vision, long-horizon-task support, and a one-million-token context window.

Qwen3.8-Max Tool

Alibaba's new flagship multimodal model, live today as a paid API at $2 per million input tokens and $6 per million output tokens, with a one-million-token context, function calling, structured output, and prompt caching that drops repeated input to $0.25 per million. Weights are promised but not published.

Qwen-Image-2.0-Pro Tool

Alibaba's latest open image-generation model in the Qwen family, downloadable and runnable locally, part of a broad open-weight release wave that also refreshed the Qwen3.6 chat models.

Ox Alpha on OpenRouter Tool

A free anonymous stealth model with a 1,048,576-token context window, up to 131,072 output tokens, and text, image, and video input. Genuinely usable and genuinely free right now, with a caveat worth reading first: OpenRouter states it is not the developer or provider, its model page says prompts and completions are retained by the anonymous provider, and other documentation for the same model claims zero retention.

Ox Alpha Tool

An anonymous reasoning model with a 1,048,576-token context window and 131,072 max output tokens, accepting text, image, and video with tool and JSON support. Free to try in the browser, with no disclosed creator.

Muse Image Tool

Meta's agentic image model, free for everyday creation inside Meta AI, Instagram Stories (US), and WhatsApp; it can search, write code, and self-refine rather than mapping a prompt straight to pixels, and stamps outputs with an invisible Content Seal watermark.

MiniMax-M3 Tool

A natively multimodal open model trained on text, image, and video from the first step, with a million-token context and a sparse-attention design built for speed; downloadable for self-hosting and also offered through MiniMax's own API and agent platform.

Meta Muse Spark Tool

Meta's natively multimodal reasoning model, updated to version 1.3 on September 2, 2026. Reads images, charts, and text together, and offers a Contemplating mode in which multiple agents reason in parallel before answering. Hosted and proprietary at $1.25 per million input tokens and $4.25 output; Meta says an open-weights release is on the roadmap but has not given a date.

Meta Model API (Muse Spark 1.1) Tool

Meta's first paid, hosted model API, built around the Muse Spark 1.1 multimodal reasoning model -- a million-token context window with active context compaction, zero-shot tool and MCP support, and an OpenAI-compatible interface so existing code drops in with little more than an endpoint change.

Mage-VL Tool

Microsoft's codec-native multimodal model that reuses a video file's own bit allocation to pick visual tokens, reporting over 75% fewer tokens and up to 3.5x faster inference than uniform frame sampling. Works with H.264, HEVC and DCVC-RT.

Inkling-Small GGUF Tool

Quantized builds of Thinking Machines' newly released 276B/12B multimodal open-weight model, packaged for llama.cpp, LM Studio and Ollama so you do not have to download the 532GB original.

Inkling Tool

Thinking Machines Lab's 975B-parameter mixture-of-experts model, released July 15 under Apache 2.0. Only ~41B parameters activate per token, it accepts text, image and audio input, and it handles up to 1M tokens of context. Artificial Analysis ranks it the top US open-weights model. Free to download, modify and use commercially -- but you will need serious hardware to run it.

Gemma 4 26B A4B Tool

Google's compute-efficient multimodal model with 25.2 billion total parameters but only 3.8 billion active per token, aimed at running usefully on hardware that cannot host a dense model of comparable capability.

Gemma 4 Tool

Google's downloadable model family (2.3B-31B, dense and MoE) that natively handles text, vision, and audio, including a 12B encoder-free variant and a thinking mode.

GLM-5.3-Flash Tool

Z.ai's 320-billion-parameter multimodal model with 18 billion active per token and a one-million-token context, released under the MIT licence -- one of the most permissive terms any model this size has shipped under. The download is 328 GB of already fp8-quantized weights, and it runs locally through SGLang, vLLM or TokenSpeed.

Fara 1.5-27B Tool

Microsoft's MIT-licensed 27B multimodal agent that operates web browsers from screenshots alone, emitting clicks, typing, scrolling, and navigation, and trained to pause on ambiguity or unauthorized irreversible actions.

DeepSeek V4-Flash vision API Tool

Experimental image input for DeepSeek's cheap V4-Flash model, billed at the same rate as the text-only version. Accepts images as inline base64, as a URL the model fetches, or as a file uploaded through the Files API. The experimental tag is real, so treat the interface as unstable, but it makes high-volume image reading economically sensible.

DeepSeek V4 Flash Vision (experimental) Tool

An experimental multimodal version of DeepSeek's cheapest model, live on the DeepSeek API as deepseek-v4-flash-vision-exp. It takes images inline with text via base64, external URL, or the Files API, budgets each image to at most 384 tokens after resizing toward roughly 800 by 800 pixels, and bills at ordinary V4 Flash rates. Good for screenshots, charts, and document layout; not for small type or dense diagrams.

Boogu-Image 0.1 Tool

An open-source unified image understanding and generation model family (Base, Turbo, Edit, Edit-Turbo) with instruction-based editing and bilingual Chinese-English text rendering, trained for roughly $400K. Apache 2.0.