classifiers
Everything on Ground Truth tagged “classifiers” — 1 item.
Guardrail models: the second AI that decides whether the first one's answer ships Lesson
A guardrail model is a small separate classifier that reads what a user sends an AI and what the AI sends back, then scores whether it violates a policy - so safety becomes a component the operator owns and can inspect, rather than a behaviour buried in the main model's weights.