text-to-speech
Everything on Ground Truth tagged “text-to-speech” — 3 items.
Neural text-to-speech: how a model turns writing into a voice Lesson
Neural text-to-speech converts written text into audio in three stages - working out the sounds, deciding how long each one lasts, and generating the actual waveform - and the last stage, the vocoder, is where most of the model's size and difficulty hides.
A complete text-to-speech system now fits in 9.4 million parameters News
Inflect-Micro-v2 packs an entire English speech synthesis stack, including the waveform decoder, into 9,356,513 parameters that run locally with no external vocoder or API.
Inflect-Micro-v2 Tool
A complete English speech synthesis stack in 9,356,513 parameters, waveform decoder included, producing 24 kHz mono audio locally with no external vocoder or API. One fixed synthetic male voice, no cloning, flatter prosody than large systems - but it runs anywhere.