language-models
Arithmetic coding: turning a model's probabilities into a compressed message Lesson
Arithmetic coding is a lossless compression method that represents an entire message as a narrowing interval, making precise why a model that assigns better probabilities can also compress better.
A gzip language-model demo shows why prediction and compression are related News
Nathan Barry's gzip experiment uses compressed length to rank continuations, making a real information-theory connection visible while also showing why local byte copying is not language understanding.
An open model that predicts ideas, not just words, matched OLMo 3's training loss on about half the data News
Researchers at Shanghai AI Lab and Shanghai Jiao Tong University released NCP-ArchPreview, an open 8.9-billion-parameter language model that learns to predict the next 'concept' as well as the next word, and reached the final training loss of a comparable open model after consuming only 51.3% of the training tokens.
N-gram language models Lesson
An n-gram language model predicts the next word by counting how often each word followed the previous one or two words in a large corpus -- the simplest working language model there is, and the direct ancestor of everything that came after.
Teacher forcing and exposure bias Lesson
Teacher forcing trains a sequence model by always feeding it the correct previous token instead of its own output, which makes training fast and stable but leaves the model unprepared for its own mistakes at generation time -- a mismatch called exposure bias.
Perplexity: the number that tells you a model still works Lesson
Perplexity measures how surprised a language model is by real text, expressed as the number of words it was effectively choosing between at each step -- lower means less confused, and a jump from single digits into the thousands means the model is broken.
A language model that doesn't write left to right News
iLLaDA is an 8-billion-parameter model that generates text by refining a blurry whole rather than one word at a time, and it's catching up to the mainstream.
What are diffusion language models? Lesson
Most AI writes one word at a time and can never go back. Diffusion language models start from noise and clarify it iteratively — and some versions can revise any word at any step. A growing alternative to the standard left-to-right approach.
An AI that could rewrite its own words — and gained nothing from it News
A different style of text AI can go back and change any word at any point as it writes. Given that power, it didn't actually produce better writing. A clean negative result.
LLaDA / iLLaDA Tool
An openly released diffusion language model (weights and code) that generates text by refining a whole passage at once rather than one word at a time, useful for experimenting with non-autoregressive generation and infilling.