Ground Truth.
AI, checked against the source.

← All topics

language-models

Everything on Ground Truth tagged “language-models” — 10 items.

Arithmetic coding: turning a model's probabilities into a compressed message Lesson

Arithmetic coding is a lossless compression method that represents an entire message as a narrowing interval, making precise why a model that assigns better probabilities can also compress better.

A gzip language-model demo shows why prediction and compression are related News

Nathan Barry's gzip experiment uses compressed length to rank continuations, making a real information-theory connection visible while also showing why local byte copying is not language understanding.

An open model that predicts ideas, not just words, matched OLMo 3's training loss on about half the data News

Researchers at Shanghai AI Lab and Shanghai Jiao Tong University released NCP-ArchPreview, an open 8.9-billion-parameter language model that learns to predict the next 'concept' as well as the next word, and reached the final training loss of a comparable open model after consuming only 51.3% of the training tokens.

N-gram language models Lesson

An n-gram language model predicts the next word by counting how often each word followed the previous one or two words in a large corpus -- the simplest working language model there is, and the direct ancestor of everything that came after.

Teacher forcing and exposure bias Lesson

Teacher forcing trains a sequence model by always feeding it the correct previous token instead of its own output, which makes training fast and stable but leaves the model unprepared for its own mistakes at generation time -- a mismatch called exposure bias.

Perplexity: the number that tells you a model still works Lesson

Perplexity measures how surprised a language model is by real text, expressed as the number of words it was effectively choosing between at each step -- lower means less confused, and a jump from single digits into the thousands means the model is broken.

A language model that doesn't write left to right News

iLLaDA is an 8-billion-parameter model that generates text by refining a blurry whole rather than one word at a time, and it's catching up to the mainstream.

What are diffusion language models? Lesson

Most AI writes one word at a time and can never go back. Diffusion language models start from noise and clarify it iteratively — and some versions can revise any word at any step. A growing alternative to the standard left-to-right approach.

An AI that could rewrite its own words — and gained nothing from it News

A different style of text AI can go back and change any word at any point as it writes. Given that power, it didn't actually produce better writing. A clean negative result.

LLaDA / iLLaDA Tool

An openly released diffusion language model (weights and code) that generates text by refining a whole passage at once rather than one word at a time, useful for experimenting with non-autoregressive generation and infilling.