The Ironclad Fine Print Behind OpenAI’s Software Agent Training
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: The Ironclad Fine Print Behind OpenAI’s Software Agent Training on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get hardware and tech essentials delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI described training and evaluating models on legal, commercial and procurement tasks inside hosted copies of contract-management software Ironclad. GPT-6 Astra met an average 55% of task rubric criteria, while its reported time estimates are simulated, not measured customer savings. OpenAI says it is seeking a small number of other software partners.

OpenAI published details on October 6 of a project with contract-management company Ironclad to train and evaluate AI models on tasks inside its software. The results offer a measure of progress on specialized business workflows, but GPT-6 Astra met an average of 55% of evaluation criteria across 11 tasks, and OpenAI says its time estimates are simulated rather than measured customer savings.

The work focused on whether models could follow company rules while completing multi-step tasks in specialized software and check that their output met the original requirements. Ironclad staff and OpenAI employees selected 11 tasks spanning legal, commercial and procurement work. Examples included setting up nondisclosure agreements, creating procurement approval processes and changing a reusable contract clause based on a requester’s jurisdiction. OpenAI estimated an experienced user would take 30 to 40 minutes per task.

Tasks were scored against rubrics of 8 to 50 criteria, depending on complexity. OpenAI reported an average of 41.6% of criteria met by GPT-5.6 Sol at the high reasoning setting and 55.0% by GPT-6 Astra at the maximum setting. An internal OpenAI model used during Astra’s development reached 63.7%. On one example task, Astra met about 94% of the criteria; that result is a single-task example, not the overall average.

OpenAI said Ironclad supplied hosted copies of its product for models to practise in. It also said it created synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database, filtered to remove personal information. OpenAI stated it did not use its customers’ data, its internal contracts or non-public Ironclad customer data. Those are statements from OpenAI; the published account does not provide a separate audit of the data or results.

At a glance
reportWhen: Published October 6; further partner wo…
The developmentOpenAI published details of an Ironclad partnership that uses software workflows and hosted product environments to train and evaluate AI agents.
Crypto market snapshot
Fear & Greed Index
64/100 — Greed
Bitcoin BTC$81,778▼ 2.0%
Ethereum ETH$2,465▼ 4.2%
Tether USDT$0.9993▼ 0.0%
BNB BNB$730.64▼ 5.3%
XRP XRP$1.37▼ 3.2%
USDC USDC$0.9996▼ 0.0%
Solana SOL$109.13▼ 5.9%
TRON TRX$0.3328▼ 0.8%
Live data · CoinGecko · alternative.me (24h change)
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Rubric Scores Matter in Contract Work

The project tests a different approach to agent training: learning to operate within real specialized software workflows, rather than only responding to general prompts. If this method improves reliability, it could make agents more capable of handling work that depends on a company’s approval rules, contract terms and records. OpenAI says it is looking for a small number of software companies to work on tasks current agents cannot reliably complete.

But a score of 55% of criteria met is not the same as completing 55% of tasks, and it does not establish that an agent can safely perform the work unattended. In a procurement workflow, for example, missing a required Finance, Security or Legal approval could undermine the whole process. OpenAI’s account recognizes that losing track of a business rule can limit what a company should delegate, and says human oversight remains important.

The reported time comparison also does not show a demonstrated productivity gain. OpenAI estimated 19.2 minutes per Astra attempt, compared with 37 minutes for GPT-5.6 Sol, but said the figures are simulated estimates based on assumed processing and generation speeds. They are not measured time saved for customers. On the evidence published, the study shows a change in model performance on selected tasks, not that businesses can yet complete contract work faster or at lower cost.

For software vendors, partnerships could help expose failure points and improve agents inside their products. They also raise a strategic question: if customers interact with software through an agent, a vendor’s value may depend less on its screens and more on the business rules, data structures, records and controls behind them. That is an implication, not a demonstrated outcome of this study.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How OpenAI Tested Ironclad Workflows

The October 6 post described two OpenAI developments. The source material says a separate announcement involving 722 mathematics manuscripts attracted more attention, while the Ironclad project received less. The Ironclad title also led some AI news trackers to interpret it as a new agent framework; according to the source material, Ironclad is a contract-management software company, and the post describes a training and evaluation partnership.

The project used tasks chosen by people familiar with the product and work, plus a hosted software environment in which models could practise. OpenAI compared GPT-5.6 Sol and GPT-6 Astra on the selected tasks and reported rubric performance and estimated attempt times. The account presents 11 research tasks, not a broad test of every Ironclad workflow or a customer deployment study. Its scores therefore apply to this evaluation setup; they do not establish how the models would perform across other businesses, data or software.

OpenAI frames the effort as part of work with software companies to teach agents business rules, complete multi-step tasks and check their work. Its stated invitation asks prospective partners to bring a concrete example of a task agents struggle with, people with deep knowledge of the work, a secure testing environment and data suitable for research. The source material says GPT-6 Astra is the first frontier model trained in this way, a characterization attributed to OpenAI’s post.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Published Scores Do Not Show

The published account does not show how often Astra completed every requirement on a task, how performance varied across all 11 tasks, or which specific criteria were most often missed. The 55% figure is an average share of rubric criteria met, not a percentage of tasks completed or a measure of safe deployment. The account also does not establish performance across other customers, products or live business conditions.

OpenAI’s time figures are simulations, and the source does not report a customer trial measuring end-to-end completion time, error rates or staff review costs. OpenAI said it did not use non-public Ironclad customer data, but the published material, as presented in the source, does not describe an independent verification of the data handling or evaluation. It also remains unclear which other software companies will take part, what data or safeguards those projects would use, and whether any resulting agent will be offered for customer use.

Amazon

contract automation software for businesses

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

OpenAI Seeks More Software Partners

OpenAI says it plans to work with a small number of software companies on tasks current agents cannot reliably complete. It asks interested partners to provide a concrete failure example, subject-matter experts, a secure testing environment and data that can safely be used for research. The source material does not give a schedule, name additional partners or announce a customer release arising from the Ironclad work.

For businesses considering agents in contract, finance or customer-record systems, the practical next step is to ask vendors how performance is tested and reported: which requirements were missed, what approvals remain under human control, what data is used, and how actions are recorded and reviewed. The Ironclad results point to progress on a limited evaluation, while leaving reliability, measured time savings and deployment terms unresolved.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Ironclad in this announcement?

Ironclad is a contract-management software company, not the name of a new OpenAI agent framework. The project tests models on tasks inside hosted copies of its software.

What does GPT-6 Astra’s 55% score mean?

OpenAI reported that Astra met an average of 55% of the rubric criteria across the selected tasks. That is not a claim that it completed 55% of the tasks or that it is ready to handle contracts without human review.

Did the project prove that AI agents save time?

No. OpenAI reported estimated times of 19.2 minutes for Astra and 37 minutes for GPT-5.6 Sol, but said these were simulated estimates based on assumed processing and generation speeds, not measured savings for customers.

What data did OpenAI say it used?

OpenAI said it created synthetic training tasks using publicly filed contracts from the SEC’s EDGAR database, filtered to remove personal information. It also said it did not use OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.

What happens after the Ironclad evaluation?

OpenAI says it is seeking a small number of other software partners to test tasks agents cannot reliably complete. It has not named further partners or announced a release schedule in the source material.

Source: ThorstenMeyerAI.com

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.
HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Understanding Investment Costs: A DIY Approach With A Dollar Cost Calculator

A new online tool enables fee-conscious investors to calculate the lifetime dollar impact of investment fees, promoting greater fee transparency.

How Artificial Intelligence Is Elevating SaaS Competitive Strategies

Artificial intelligence is shifting SaaS market dynamics by changing switching costs, market valuation, and competitive edges, with significant implications for the industry.

SEC Approves Binance.US Sale to US Consortium, Averting Shutdown

Find out how the SEC’s approval of Binance.US’s sale to a U.S. consortium is shaping the future of the platform and what it means for users.

What Makes a Crypto Brand Trustworthy in 2026

The trustworthiness of a crypto brand in 2026 hinges on key qualities that ensure safety and transparency—discover what truly makes a brand reliable.