self-supervision
Everything on Ground Truth tagged “self-supervision” — 1 item.
Models can train each other without a single correct answer News
A method called Co-RL trains language models with no labels at all by rewarding each model for agreeing with a different model's majority vote, matching and sometimes beating the same recipe trained with ground-truth answers.