rlaif
Everything on Ground Truth tagged “rlaif” — 1 item.
Constitutional AI: training a model against written principles instead of human labels Lesson
Constitutional AI is Anthropic's training method in which a model critiques and revises its own answers against a written list of principles, then learns from AI-generated preference labels instead of human harmfulness ratings, making the values it is trained toward readable and revisable.