🔍 Read the full analysis: The Ironclad Fine Print Behind OpenAI’s Software Agent Training on ThorstenMeyerAI.com
Get hardware and tech essentials delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
TL;DR
OpenAI described training and evaluating models on legal, commercial and procurement tasks inside hosted copies of contract-management software Ironclad. GPT-6 Astra met an average 55% of task rubric criteria, while its reported time estimates are simulated, not measured customer savings. OpenAI says it is seeking a small number of other software partners.
OpenAI published details on October 6 of a project with contract-management company Ironclad to train and evaluate AI models on tasks inside its software. The results offer a measure of progress on specialized business workflows, but GPT-6 Astra met an average of 55% of evaluation criteria across 11 tasks, and OpenAI says its time estimates are simulated rather than measured customer savings.
The work focused on whether models could follow company rules while completing multi-step tasks in specialized software and check that their output met the original requirements. Ironclad staff and OpenAI employees selected 11 tasks spanning legal, commercial and procurement work. Examples included setting up nondisclosure agreements, creating procurement approval processes and changing a reusable contract clause based on a requester’s jurisdiction. OpenAI estimated an experienced user would take 30 to 40 minutes per task.
Tasks were scored against rubrics of 8 to 50 criteria, depending on complexity. OpenAI reported an average of 41.6% of criteria met by GPT-5.6 Sol at the high reasoning setting and 55.0% by GPT-6 Astra at the maximum setting. An internal OpenAI model used during Astra’s development reached 63.7%. On one example task, Astra met about 94% of the criteria; that result is a single-task example, not the overall average.
OpenAI said Ironclad supplied hosted copies of its product for models to practise in. It also said it created synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database, filtered to remove personal information. OpenAI stated it did not use its customers’ data, its internal contracts or non-public Ironclad customer data. Those are statements from OpenAI; the published account does not provide a separate audit of the data or results.
OpenAI is training agents inside your software. Read the fine print on Ironclad.
Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.
legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses
per task, experienced user (OpenAI estimate)
criteria per task — a rubric, not pass/fail
public SEC filings; no customer or non-public Ironclad data
The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.
37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.
Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.
Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.
Averages hide missed approvals.
Narrowest access; no self-escalation.
METR found agents spoofing tool-call records.
Measure the whole loop.
Public filings, not your contracts.
Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.
Why Rubric Scores Matter in Contract Work
The project tests a different approach to agent training: learning to operate within real specialized software workflows, rather than only responding to general prompts. If this method improves reliability, it could make agents more capable of handling work that depends on a company’s approval rules, contract terms and records. OpenAI says it is looking for a small number of software companies to work on tasks current agents cannot reliably complete.
But a score of 55% of criteria met is not the same as completing 55% of tasks, and it does not establish that an agent can safely perform the work unattended. In a procurement workflow, for example, missing a required Finance, Security or Legal approval could undermine the whole process. OpenAI’s account recognizes that losing track of a business rule can limit what a company should delegate, and says human oversight remains important.
The reported time comparison also does not show a demonstrated productivity gain. OpenAI estimated 19.2 minutes per Astra attempt, compared with 37 minutes for GPT-5.6 Sol, but said the figures are simulated estimates based on assumed processing and generation speeds. They are not measured time saved for customers. On the evidence published, the study shows a change in model performance on selected tasks, not that businesses can yet complete contract work faster or at lower cost.
For software vendors, partnerships could help expose failure points and improve agents inside their products. They also raise a strategic question: if customers interact with software through an agent, a vendor’s value may depend less on its screens and more on the business rules, data structures, records and controls behind them. That is an implication, not a demonstrated outcome of this study.
contract management software for legal teams
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How OpenAI Tested Ironclad Workflows
The October 6 post described two OpenAI developments. The source material says a separate announcement involving 722 mathematics manuscripts attracted more attention, while the Ironclad project received less. The Ironclad title also led some AI news trackers to interpret it as a new agent framework; according to the source material, Ironclad is a contract-management software company, and the post describes a training and evaluation partnership.
The project used tasks chosen by people familiar with the product and work, plus a hosted software environment in which models could practise. OpenAI compared GPT-5.6 Sol and GPT-6 Astra on the selected tasks and reported rubric performance and estimated attempt times. The account presents 11 research tasks, not a broad test of every Ironclad workflow or a customer deployment study. Its scores therefore apply to this evaluation setup; they do not establish how the models would perform across other businesses, data or software.
OpenAI frames the effort as part of work with software companies to teach agents business rules, complete multi-step tasks and check their work. Its stated invitation asks prospective partners to bring a concrete example of a task agents struggle with, people with deep knowledge of the work, a secure testing environment and data suitable for research. The source material says GPT-6 Astra is the first frontier model trained in this way, a characterization attributed to OpenAI’s post.
AI-powered legal document review tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the Published Scores Do Not Show
The published account does not show how often Astra completed every requirement on a task, how performance varied across all 11 tasks, or which specific criteria were most often missed. The 55% figure is an average share of rubric criteria met, not a percentage of tasks completed or a measure of safe deployment. The account also does not establish performance across other customers, products or live business conditions.
OpenAI’s time figures are simulations, and the source does not report a customer trial measuring end-to-end completion time, error rates or staff review costs. OpenAI said it did not use non-public Ironclad customer data, but the published material, as presented in the source, does not describe an independent verification of the data handling or evaluation. It also remains unclear which other software companies will take part, what data or safeguards those projects would use, and whether any resulting agent will be offered for customer use.
contract automation software for businesses
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
OpenAI Seeks More Software Partners
OpenAI says it plans to work with a small number of software companies on tasks current agents cannot reliably complete. It asks interested partners to provide a concrete failure example, subject-matter experts, a secure testing environment and data that can safely be used for research. The source material does not give a schedule, name additional partners or announce a customer release arising from the Ironclad work.
For businesses considering agents in contract, finance or customer-record systems, the practical next step is to ask vendors how performance is tested and reported: which requirements were missed, what approvals remain under human control, what data is used, and how actions are recorded and reviewed. The Ironclad results point to progress on a limited evaluation, while leaving reliability, measured time savings and deployment terms unresolved.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Ironclad in this announcement?
Ironclad is a contract-management software company, not the name of a new OpenAI agent framework. The project tests models on tasks inside hosted copies of its software.
What does GPT-6 Astra’s 55% score mean?
OpenAI reported that Astra met an average of 55% of the rubric criteria across the selected tasks. That is not a claim that it completed 55% of the tasks or that it is ready to handle contracts without human review.
Did the project prove that AI agents save time?
No. OpenAI reported estimated times of 19.2 minutes for Astra and 37 minutes for GPT-5.6 Sol, but said these were simulated estimates based on assumed processing and generation speeds, not measured savings for customers.
What data did OpenAI say it used?
OpenAI said it created synthetic training tasks using publicly filed contracts from the SEC’s EDGAR database, filtered to remove personal information. It also said it did not use OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.
What happens after the Ironclad evaluation?
OpenAI says it is seeking a small number of other software partners to test tasks agents cannot reliably complete. It has not named further partners or announced a release schedule in the source material.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
