Ground Truth.
AI, checked against the source.

News · 2026-10-08

Kandinsky 6 releases synchronized video and audio, with demanding local tradeoffs

Kandinsky 6 released open audio-video models on October 6 that generate five-second clips with synchronized sound, including speech synchronization. The MIT-licensed release includes local inference code and memory-offloading configurations down to a 16 GB graphics card. Those configurations broaden access, but the published base-model timings show that fitting a model into memory can still mean waiting many minutes for a short clip.

Key facts

Kandinsky Lab’s paper describes “Foundation Models for Synchronized Video and Audio Generation.” The core idea is to produce moving images and sound together rather than attach a soundtrack after the visuals are complete. A person’s mouth, an object’s impact and the corresponding audio need coordinated timing. If two separate generators make independent decisions, each can produce a plausible result while disagreeing about when an event occurs.

The full technical paper describes two interacting streams inside a diffusion transformer. One works on video and one on audio. They exchange information in both directions and each reads the text prompt. Think of a film editor and sound editor sharing a timeline throughout production, rather than meeting after the picture has been locked. The model’s internal exchange gives each stream a chance to condition its choices on the other.

The family contains Lite and Pro versions. Their principal generation models have 3 billion and 29 billion parameters respectively, but those numbers do not describe the complete installation or memory footprint. The Lite model card also lists video and audio encoders and text components. The authors train separate streams before joining them on paired audiovisual material, followed by additional fine-tuning, reward-based training and distillation.

Data alignment is central to the mechanism. The paper reports a paired training collection of seven million segments filtered for audio quality and audiovisual match, including synchronization checks. The image-conditioned path starts from a frame and generates an audiovisual continuation. A separate super-resolution model produces Full HD output. These details distinguish a coordinated generation pipeline from the stronger claim that every aspect of motion, speech and sound is solved.

The installation question has a concrete answer for one practical variant. The verified Pro-distilled file listing corresponds to about 80.6 GB on disk. That is the repository download footprint, not a full runtime memory requirement and not a size claim for the Lite checkpoint. The dossier does not establish a Lite download size, so no value is guessed from its parameter count.

The low-memory path shifts work between the graphics card and system memory. At the documented 16 GB setting, only two generation blocks stay on the card at once, the text encoder is quantized and freed after use, and super-resolution decoding is divided spatially. The paper reports the process occupying about 15.1 GiB of a card with 15.6 GiB free, while Pro uses about 58 GiB of host-process memory. Disk paging can make the already expensive transfer loop much slower. The offloading lesson explains why a lower graphics-memory requirement does not eliminate the workload.

The most vivid performance number is roughly 51 minutes for Pro to generate a five-second standard-definition clip on the reported RTX 5060 Ti configuration. Lite takes about 22 minutes there. These are full, non-distilled base-model results that exclude weight loading and final encoding. They must not be assigned to the distilled variant merely because its file listing is discussed nearby. Distillation is a separate inference path, described with two network evaluations in the paper.

Quality evidence is also conditional. The authors’ side-by-side study collects roughly 200 or more judgments per criterion, with replay, ties and not-applicable options. Pro improves on the previous Kandinsky version and competes on selected synchronization, speech and prompt-following dimensions. Other systems lead on many visual or overall audio criteria. The paper acknowledges a gap from the strongest proprietary systems, making a blanket claim of superiority unsupported.

For readers wanting a low-friction trial, the official Pro-distilled hosted demo avoids a local installation. The MIT license and release artifacts make the project more inspectable than a closed preview. The honest caveat is that openness, synchronization and low-memory feasibility do not establish fast consumer use or best-in-class quality. The release’s real value is access to a joint audiovisual system whose architecture and deployment compromises can be examined directly.


Primary source, verified: read the paper →

Key questions

Does Kandinsky 6 add sound after making a silent video?

No: separate audio and video streams exchange information inside the generation model. A separate super-resolution stage then raises video resolution.

Can Kandinsky Pro run on a 16 GB graphics card?

A 16 GB offload configuration is documented, but it nearly fills the card and the Pro process uses about 58 GiB of host memory. That configuration establishes feasibility, not fast or effortless generation.

How large is the Pro distilled download?

The dossier verifies about 80.6 GB for the Pro-distilled model repository. That disk footprint is distinct from the memory needed during inference.
Cite this

APA

Ground Truth. (2026, October 8). Kandinsky 6 releases synchronized video and audio, with demanding local tradeoffs. Ground Truth. https://groundtruth.day/news/kandinsky-6-open-synchronized-audio-video.html

BibTeX

@misc{groundtruth:kandinsky-6-open-synchronized-audio-video,
  title  = {Kandinsky 6 releases synchronized video and audio, with demanding local tradeoffs},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {oct},
  url    = {https://groundtruth.day/news/kandinsky-6-open-synchronized-audio-video.html}
}

Topics: open-weights · video-generation · audio · multimodal · generative-ai

Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.