Ground Truth.
AI, checked against the source.

News · 2026-10-07

OpenAI’s Ironclad research improves workflow scores, without announcing a customer agent

OpenAI reports that GPT-6 Astra met an average 55% of evaluation criteria across 11 research tasks in Ironclad’s contract-management software, compared with 41.6% for GPT-5.6 Sol. The collaboration shows how business experts and hosted software environments can shape agent training around real workflow rules. It does not announce a shipping Ironclad customer feature or establish measured savings in live customer work.

Key facts

The hard part of an enterprise workflow is often the exception. An agent may correctly build an attractive procurement form yet omit Finance approval above a spending threshold. It may route ordinary terms to Legal while forgetting the special security review. The resulting interface looks finished, but the company’s controls are wrong. OpenAI and Ironclad’s collaboration makes those rules the center of training and evaluation.

The published examples include setting up nondisclosure agreements, creating procurement approval processes, and updating a reusable legal clause for a selected jurisdiction. The companies do not publish the complete list of task prompts. In the procurement example, success involves an intake form, templates, approvals, and a final-agreement record. Checking requests both above and below the spending threshold matters because the business rule must work in both directions.

Each task had between eight and 50 evaluation criteria. This makes partial progress visible: an agent can satisfy some requirements and miss others. It also makes the headline percentage easy to misread. A 55% mean does not say that 55% of contracts were successfully managed from start to finish. It says the model met that average share of criteria under the study’s scoring approach. The importance of an omitted criterion could vary enormously.

OpenAI does not publish criterion weights, task-level scores, repeat counts, or the grading protocol. Ironclad’s experts helped define success, but the announcement does not establish that they personally scored every attempt. The reported comparison also uses different reasoning settings: Astra at Max and Sol at High, the settings where each performed best. That is useful product-level information, but it is not a controlled comparison of identical reasoning budgets.

The training mechanism is practice in hosted Ironclad environments, using reinforcement learning and feedback. Imagine teaching a new operations employee in a training copy of the contract system. Rather than awarding credit for confidently describing a policy, the training environment can expose whether the employee configured the approval path and tested it. Agent harnesses and the surrounding environment become part of what the model can learn.

OpenAI says it created simulated tasks from publicly filed United States company contracts after filters intended to remove personal information. It says it excluded OpenAI customer data, OpenAI’s internal contracts, and nonpublic Ironclad customer contracts. That provenance is important for understanding the collaboration. Using a partner’s software environment is different from using the partner’s customers’ private contracts as training material.

The post leaves important implementation questions unanswered: how environments reset, what rewards are assigned, how much training occurs, and whether the 11 evaluation tasks were held out. Those gaps limit conclusions about generalization. They do not erase the reported gain; they establish how much additional evidence a buyer needs before treating it as a broadly deployable capability.

Time estimates need the same care. OpenAI reports simulated average attempt times of 19.2 minutes for Astra and 37.0 minutes for Sol. Its footnote says those figures use assumed processing and generation speeds. They are not stopwatch measurements of customer employees saving time. Replacing simulated durations with a claim of observed business return would change the evidence into something the source does not provide.

Ironclad CTO Sunita Verma emphasizes understanding the contracting lifecycle while preserving controls. That is relevant partner expertise, rather than independent benchmark validation. OpenAI says human oversight still matters. The partner-facing language says Ironclad “has an opportunity” to use more capable models; it does not name an available product, launch date, or customer pilot.

The Astra release page establishes the model’s separate product status. The Ironclad company update confirms the collaboration. Neither turns the research setup into an announced contract agent. Community claims lacking an authenticated discussion are omitted rather than presented as consensus.

A useful comparison is DeskForge, whose full paper describes 1.2 million annotated desktop observations. Both efforts construct learning signals from software, but DeskForge targets general interface grounding while Ironclad targets domain workflows. OpenAI does not say it used DeskForge.

The commercial implication is specific: better enterprise agents may come from teaching the rules of a job inside realistic environments. The reported average also explains why final review remains necessary. A partially correct approval system can be more dangerous than an obviously unfinished one.


Primary source, verified: read the paper →

Key questions

Is the Ironclad-trained agent available to Ironclad customers?

No customer-facing feature or deployment date is announced in the collaboration post. GPT-6 Astra’s general availability is separate from this research evaluation.

Were private customer contracts used for the research?

OpenAI says they were not. It describes simulated tasks based on filtered public SEC EDGAR contracts and excludes nonpublic Ironclad customer contracts.

Does a 55% rubric score mean 55% of contracts were completed?

No: it is an average share of evaluation criteria met across the 11 tasks. It is not an end-to-end contract-completion rate.
Cite this

APA

Ground Truth. (2026, October 7). OpenAI’s Ironclad research improves workflow scores, without announcing a customer agent. Ground Truth. https://groundtruth.day/news/openai-ironclad-research-improves-workflow-rubric-scores.html

BibTeX

@misc{groundtruth:openai-ironclad-research-improves-workflow-rubric-scores,
  title  = {OpenAI’s Ironclad research improves workflow scores, without announcing a customer agent},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {oct},
  url    = {https://groundtruth.day/news/openai-ironclad-research-improves-workflow-rubric-scores.html}
}

Topics: agents · computer-use · enterprise · evaluation · openai

Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.