training-methods
Teacher forcing and exposure bias Lesson
Teacher forcing trains a sequence model by always feeding it the correct previous token instead of its own output, which makes training fast and stable but leaves the model unprepared for its own mistakes at generation time -- a mismatch called exposure bias.
Pseudo-labeling and self-training Lesson
Pseudo-labeling is training a model on labels it produced itself: run a model over unlabeled data, keep the predictions it is confident about, and treat them as ground truth for the next round of training. It works surprisingly well, and it fails in one specific way -- by confidently reinforcing its own mistakes.
Models can train each other without a single correct answer News
A method called Co-RL trains language models with no labels at all by rewarding each model for agreeing with a different model's majority vote, matching and sometimes beating the same recipe trained with ground-truth answers.
The cheapest way to teach a model turned out to be blindfolding the student News
Researchers got most of the benefit of an expensive teacher model by degrading the student's view of the image instead of upgrading the teacher, lifting a 4-billion-parameter model past open models nearly sixty times its size with no labels, rewards or larger teacher.