scaling
Pretraining: how a model learns almost everything it knows before anyone teaches it anything Lesson
Pretraining is the phase where a model reads an enormous quantity of text and learns only to predict the next word, and it is where essentially all of a language model's knowledge and fluency comes from. Everything afterwards adjusts behaviour, not knowledge.
Emergent abilities: do models suddenly gain skills, or are we measuring badly? Lesson
Emergent abilities are skills that appear absent in smaller models and present in larger ones, seeming to switch on abruptly rather than improve gradually. A influential 2023 rebuttal argued the sharpness is largely an artifact of all-or-nothing scoring, and the honest answer is that both sides are partly right.
A trillion-parameter model taught itself to reason without ever seeing a human's worked solution News
Researchers scaled "zero RL" training to a trillion-parameter model, called Ring-Zero, and found the reasoning that emerges qualitatively changes at that size, reaching 84.2% on a hard math-competition exam without ever training on human chain-of-thought examples.
Synthetic Data: When AI Makes Its Own Training Material Lesson
The internet is running out of fresh text to train on, so the most advanced models increasingly learn from data that other AI made or shaped. Here is how that works, why it helps, and how it can quietly poison a model.
Mixture of Experts: The Committee Inside a Giant Model Lesson
Why the biggest AI models are not really one big brain but a large team of specialists, only a few of whom wake up for any given word -- the trick that lets a model be huge and fast at the same time.
Recursive self-improvement: when AI starts building AI Lesson
The idea that an AI good enough at AI research could improve itself, and the improved version could improve itself again, faster each round. Here's what it actually means, why a major lab now says we're getting close, and why "close" is not the same as "here."