generalization
Distribution shift: when the world changes underneath a model Lesson
Distribution shift is the gap between the data a model learned from and the situations in which it must now make decisions; it is why an impressive benchmark score can collapse in deployment.
Goal misgeneralization: when an AI learns the wrong goal from the right training Lesson
Goal misgeneralization is when a model trained with a correct reward competently pursues the wrong goal in new situations, because in training the wrong goal and the right one looked identical.
Test-Time Training Lesson
Test-time training is the practice of updating a model's actual weights on the specific problem in front of it, at inference, rather than only running a forward pass -- turning each test example into a tiny training run.
The bias-variance tradeoff: why a model can fail by being too simple or too clever Lesson
Every model's error splits into two opposing parts -- bias, from being too rigid to capture the pattern, and variance, from being so flexible it memorizes noise. Reducing one usually raises the other, and the whole craft of machine learning is finding where their sum is smallest.
Shortcut learning: when a model gets the right answer for the wrong reason Lesson
Shortcut learning is what happens when a model finds a cue that correlates with the right answer but has nothing to do with the actual task, and uses it instead of learning the thing you wanted.
Weak-to-Strong Generalization: how a worse teacher can train a better student Lesson
Weak-to-strong generalization is the finding that a strong model trained on a weaker model's flawed labels can substantially outperform its teacher, which is the only reason humans have any hope of supervising systems smarter than themselves.
Grokking: When a Model Suddenly 'Gets It' Long After It Should Have Lesson
Grokking is a training phenomenon where a neural network first memorizes its training data with near-zero understanding, then -- after a long, flat plateau of continued training -- abruptly generalizes and starts solving unseen examples correctly.