Ground Truth.
AI, checked against the source.

← All topics

mechanistic-interpretability

Everything on Ground Truth tagged “mechanistic-interpretability” — 7 items.

Researchers find a steerable ‘pain axis’ in language models, not proof they feel pain News

A preprint reports a residual-stream direction linked to self-directed distress across 25 open models and behavior changes in a controlled Qwen task, while explicitly stopping short of any consciousness claim.

Interpretability is moving from features to geometry, and its researchers say so out loud News

Goodfire researcher Tom McGrath addressed the circulating claim that sparse autoencoders are dead, arguing they remain pragmatically useful but capture only partial views of curved structure, as his lab pushes toward geometry-aware interpretability and training-time control instead of post-hoc feature extraction.

One layer creates the giant activations behind attention sinks News

Researchers identified a single layer, consistent across model families, where the outsized activations that produce attention sinks first appear, and showed that loosening that token's rigidity improves instruction following and math reasoning without retraining.

MIT found brain-like modules inside six large language models News

MIT researchers localized the neurons behind 46 reasoning tasks in six large language models and found that tasks sharing a brain network in humans share neurons in the models, with 4.3 times more overlap within a cognitive domain than across domains.

The Logit Lens: Reading a Model's Guesses Before It Finishes Thinking Lesson

The logit lens is an interpretability technique that decodes a language model's partial, mid-computation representations into vocabulary words, letting researchers watch the model's best guess evolve layer by layer before it produces a final answer.

Polishing AI by looking inside its 'mind' instead of just thumbs-up, thumbs-down News

Reward training usually treats the model as a black box — thumbs up, thumbs down, hope for the best. A new method peers inside to see why an answer was preferred, and shapes the lesson on purpose.

The hidden escape hatch in AI safety controls News

Researchers at Hong Kong Polytechnic University show that clamping an AI safety feature — like one that controls refusals — doesn't remove the behavior. It hides in the part of the model's internal state that the safety tool throws away, and can be recovered while the monitored feature looks perfectly controlled.