Ground Truth.
AI, checked against the source.

← All topics

audio

Everything on Ground Truth tagged “audio” — 16 items.

Kandinsky 6 releases synchronized video and audio, with demanding local tradeoffs News

Kandinsky 6 ships MIT-licensed audio-video models and consumer-GPU offload recipes, but the documented low-memory base-model runs remain slow.

Google ships Gemini 3.8 TTS with text-designed voices and consent-gated replication News

Google released Gemini 3.8 TTS with text-described voice design, directed multi-speaker scenes, and voice replication that requires a 10–30 second sample plus a matching consent recording.

MiniMax put its video-with-sound model on Hugging Face News

MiniMax released the weights for H3, a multimodal generation model that produces video with native stereo audio, two weeks after promising them at launch, and it has already been downloaded more than two million times.

Full-Duplex Speech Models: Listening and Talking at the Same Time Lesson

A full-duplex speech model processes incoming audio while it is generating outgoing audio, which removes the turn detector that decides when you have stopped speaking and makes interruption, backchannels, and overlap possible.

Vector Quantization: Turning Continuous Data Into a Vocabulary Lesson

Vector quantization forces a neural network's continuous internal representations to snap to a finite set of learned reference vectors, converting images, audio, or video into sequences of discrete symbols that a language model can predict just like words.

Neural text-to-speech: how a model turns writing into a voice Lesson

Neural text-to-speech converts written text into audio in three stages - working out the sounds, deciding how long each one lasts, and generating the actual waveform - and the last stage, the vocoder, is where most of the model's size and difficulty hides.

How AI Turns Speech Into Text Lesson

Automatic speech recognition (ASR) converts spoken audio into written text by breaking sound into tiny slices, encoding them into features a model understands, and decoding those into words -- and its accuracy is measured by word error rate, the fraction of words it gets wrong.

audio.cpp Tool

The September 30 v0.9.0 README reports a shared local audio runtime covering more than 100 models and 170 variants for recognition, synthesis, conversion, and separation.

YuE2-3B Tool

M-A-P's song generator, released 9 September: writes an editable melody-and-chord score in ABC notation, then renders the full song. A 7.3 GB download plus a 0.5 GB decoder; the model card asks for a 24 GB NVIDIA GPU. Weights are CC BY-NC 4.0 (non-commercial), code Apache 2.0.

MiniMax-H3 Tool

Open weights for MiniMax's omni-modal model that generates four to fifteen second video with native stereo audio. The locally deployable base runs at 768p through diffusers or SGLang; the prompt-interpretation and 2K regeneration stages stay behind MiniMax's API, and the licence excludes the US, EU, UK, and South Korea.

MiniMax Music 3.0 Tool

Production music model that takes a creative concept and optional lyrics and composes, arranges, performs and produces a complete song in a single generation, with instrumental-only support. Callable through MiniMax's platform API as model music-3.0, with open weights also published.

MiniMax Music 3 Tool

Open-weight model that generates complete five-minute songs with vocals in 32 kHz stereo from lyrics plus a structured style description. Runs via SGLang-Omni, Diffusers, or ComfyUI. Commercial use allowed with on-screen attribution; written permission required above $20M revenue.

Kandinsky 6 Pro-distilled demo Tool

Official hosted demo for synchronized five-second video and sound generation, avoiding the substantial local model download and memory setup.

Gemini Voice Replication Tool

Create a consent-verified synthetic voice from a 10–30 second sample and a matching adult consent recording.

Gemini 3.8 Flash TTS Tool

Generate directed speech, text-designed fictional voices, and two-speaker scenes through Google's Gemini API.

FLUX 3 (early access) Tool

Black Forest Labs' unified generation model, producing video up to 20 seconds with native synchronized audio from text, image, video or keyframe inputs. Video is behind an early-access request today; image access is promised in the following weeks and open weights are deferred.