News · 2026-09-07
OpenAI says it is prioritising RSI and alignment over making models better at math research
OpenAI says it could make its models better at mathematical research but is not prioritising that work because recursive self-improvement and automated alignment research are more urgent. The admission is significant because it describes a concrete research-allocation tradeoff, while stopping short of a public, enforceable limit on capability development.
Key facts
- Jakub Pachocki made the statement in OpenAI’s September 6 essay, An Alien Mind.
- The stated priority is recursive self-improvement, often shortened to RSI, plus automated alignment research.
- OpenAI’s companion research-acceleration account says it is using 3.1 agent-workdays for every human workday.
- The essay proposes alignment, monitoring and potentially coordinated slowing, but names no binding safety threshold.
The important part is not the familiar claim that AI safety matters. It is that OpenAI publicly says it has chosen not to maximise one identifiable capability. Pachocki writes that the company could improve mathematics research with more focus but sees more urgency around RSI and automated alignment. Put plainly, the lab is saying that the best use of a marginal researcher, training run or engineering project is not necessarily the benchmark whose result is easiest to show.
That priority sits beside an acceleration story. OpenAI’s research-acceleration post says agents are already doing longer-horizon work and that the organisation counts 3.1 agent-workdays per human workday. The company says it pauses or constrains runs when safety bars are not met, but the essay does not disclose the bars. The result is an unusual combination: a lab that expects progress toward RSI, is actively using AI to speed research, and says that confidence in monitoring may become the bottleneck.
Pachocki separates goal alignment from value alignment. Goal alignment means a system tries to do the specified task; value alignment is the harder problem of generalising human principles in unclear or adversarial circumstances. He also discusses why OpenAI hid o1-preview’s chain of thought: making hidden reasoning a target for supervision can change the thing being observed. That is a direct connection to chain-of-thought faithfulness, the problem of whether an apparently sensible explanation is the real cause of an answer.
OpenAI’s alternative is “confessions.” In How confessions can keep language models honest, the lab describes a separate output trained only for honesty, rather than one whose reward is tied to producing a pleasing main answer. The associated paper reports a 4.4% average false-negative rate across its adversarial evaluations. A useful analogy is a post-flight incident report: it is not the pilot’s live narration, but a separate channel designed to make later auditing more candid.
The strongest criticism is not that OpenAI should ignore safety. It is that an intention is not a control. The Hacker News discussion repeatedly asks what would trigger a slowdown, who would verify it and why competitive pressure would not override it. Anthropic’s Responsible Scaling Policy offers a nearby model with capability thresholds and required safeguards, though it does not establish OpenAI’s claimed tradeoff.
OpenAI’s own language is careful: it says future commitments could be enforced by third-party auditors, governments or international bodies. “Could” is doing real work. There is no public auditor, schedule or definition of sufficient confidence. The honest caveat is therefore that this is evidence of strategic intent, not evidence that a robust external brake exists.
Why it matters: the next governance argument will not only be about whether a released model is safe. It will be about which kinds of capability work labs choose to accelerate, which they defer, and whether those choices can be checked from outside.
Key questions
What capability is OpenAI not prioritising?
Has OpenAI announced a hard cap on model progress?
Cite this
APA
Ground Truth. (2026, September 7). OpenAI says it is prioritising RSI and alignment over making models better at math research. Ground Truth. https://groundtruth.day/news/openai-says-rsi-and-alignment-outrank-math-research.html
BibTeX
@misc{groundtruth:openai-says-rsi-and-alignment-outrank-math-research,
title = {OpenAI says it is prioritising RSI and alignment over making models better at math research},
author = {{Ground Truth}},
year = {2026},
month = {sep},
url = {https://groundtruth.day/news/openai-says-rsi-and-alignment-outrank-math-research.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.