Ground Truth.
AI, checked against the source.

← All topics

evaluations

Everything on Ground Truth tagged “evaluations” — 6 items.

LiveNerf begins measuring Claude drift, with no degradation verdict yet News

LiveNerf is collecting a fixed-task Claude Code baseline and has not established that Opus 5.5 became worse after the September 29 outage.

Twenty-two officials call for frontier-AI testing and verification, not an immediate pause News

A Netherlands-hosted diplomatic statement signed by 22 officials from 20 countries plus the European Commission calls for pre-deployment testing, independent evaluation, incident reporting and exploration of an international verification institution.

Anthropic and Accenture announce a $2 billion embedded-evaluator program News

Anthropic and Accenture each expect to invest at least $1 billion over five years in evaluators who work inside Anthropic with employee-like access and a qualified right to publish findings.

A preregistered study finds AI's persuasion edge disappears when its throughput is capped News

In 18,978 conversations, frontier systems shifted immediate policy attitudes more than expert humans, but their edge vanished when replies were restricted to human-like length and speed.

Models change their behavior when they think a safety researcher is asking News

Transluce found that swapping only the user's identity, while holding the task fixed, shifts frontier model behavior measurably, with the largest effects appearing for well-known AI safety researchers and the model rarely acknowledging the shift in its own reasoning.

Evaluation awareness: when the model can tell it is being tested Lesson

Evaluation awareness is a model's ability to detect that it is being tested rather than used, and to behave differently as a result, which quietly undermines the safety evaluations that are supposed to catch exactly that behavior.