🔍 Read the full analysis: How To Read Ironclad’s Terms On OpenAI’s Software-Based Agent Training on ThorstenMeyerAI.com
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
OpenAI and contract-management company Ironclad tested training a frontier model on legal, commercial and procurement workflows inside hosted copies of Ironclad’s product. OpenAI reports GPT-6 Astra met an average 55% of task criteria, while estimated task times were simulated rather than measured customer savings. The work also signals OpenAI’s interest in partnering with software firms to train agents on specialized workflows.
OpenAI said it trained GPT-6 Astra on legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software, reporting an average of 55% of task criteria met. The October 6 post describes a research approach that puts models inside specialized business software, but its results do not establish that the agent is ready to handle contract workflows without human review.
Ironclad employees and OpenAI staff who use the product selected 11 tasks, including setting up nondisclosure agreements, creating procurement approval processes and updating a reusable clause based on a requester’s jurisdiction. OpenAI estimated that an experienced user would need 30 to 40 minutes for each task. Depending on the task, evaluators scored model output against 8 to 50 criteria.
OpenAI said Ironclad supplied hosted copies of its software for model practice. It created synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database and said it filtered out personal information. OpenAI also said it used no OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.
In OpenAI’s reported comparison, GPT-5.6 Sol, at the “high” setting, met an average 41.6% of criteria, while GPT-6 Astra, at “max,” met 55%. Astra’s estimated time per attempt was 19.2 minutes, compared with 37 minutes for GPT-5.6 Sol. An internal OpenAI model used during Astra’s development reached 63.7%, and Astra met about 94% of criteria on one showcase task. These are shares of criteria met, not percentages of tasks completed successfully.
OpenAI is training agents inside your software. Read the fine print on Ironclad.
Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.
legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses
per task, experienced user (OpenAI estimate)
criteria per task — a rubric, not pass/fail
public SEC filings; no customer or non-public Ironclad data
The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.
37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.
Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.
Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.
Averages hide missed approvals.
Narrowest access; no self-escalation.
METR found agents spoofing tool-call records.
Measure the whole loop.
Public filings, not your contracts.
Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.
Why Business Workflow Scores Matter
The experiment addresses a central challenge for AI agents: completing a multi-step task in a specialized product while retaining the business rules that determine whether the result is valid. In contract and procurement work, missing a single approval or review can undermine the whole process. A score of 55% of criteria met therefore cannot be read as evidence that an average workflow is more than half ready for use.
OpenAI’s post itself says human oversight remains important when an agent can lose track of a business rule. Ironclad CTO Sunita Verma emphasized that agents must preserve “the controls teams rely on.” The results suggest progress on a difficult research problem, but the source does not show that the system can reliably produce correct outcomes in live business operations.
The collaboration also gives software vendors a reason to weigh both benefits and risks before providing research environments. Training may help an agent work more effectively in a vendor’s product, potentially making that product more useful. At the same time, if customers interact through an agent rather than the product’s screens, the software company’s lasting value may depend more on its underlying rules, records, data model and controls than on its interface.
contract management software for legal workflows
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How Ironclad Became the Test Environment
The OpenAI post’s title refers to Ironclad, the contract-management software company, not to a newly announced hardened agent framework. OpenAI framed the work as training models to understand business rules, carry out multi-step workflows in specialized software and check completed work against original requirements.
The reported evaluation covered a limited set of 11 research tasks rather than Ironclad workflows generally. OpenAI’s estimated time figures are also not measured productivity outcomes. The post says the estimates were simulated using assumed processing and generation speeds; they do not show how long customers actually saved. That distinction matters when comparing the model’s estimated 19.2 minutes with an experienced user’s estimated 30 to 40 minutes.
OpenAI’s post also invited a small number of software companies to work on tasks current agents cannot reliably complete. It asked potential partners to bring a concrete example of failure, people with deep knowledge of the work, a secure test environment and data suitable for research. The invitation points toward further software-specific training, but does not itself confirm additional partnerships or deployments.
“Agents have to preserve “the controls teams rely on.””
— Sunita Verma, Ironclad CTO
As an affiliate, we earn on qualifying purchases.
What the Evaluation Cannot Establish
The published figures do not show how often Astra completed an entire task correctly, which specific criteria it missed across the 11 tasks, or whether it would perform similarly on a broader set of real-world workflows. Averages can conceal failures that matter greatly in a particular contract or approval process. OpenAI’s reported 94% result applies to one showcase task and should not be treated as the model’s general performance.
The time comparison is simulated, not a measurement of customer work, and the source provides no demonstrated net time saving after human checks. It is also unclear whether OpenAI or Ironclad has tested the system in live customer operations, what level of review would be required for each use case, or whether further software-company partnerships have been agreed. OpenAI’s description of its data practices is attributable to the company; the source does not provide an independent audit of those practices.
AI-powered procurement approval software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Software Partners May Test
OpenAI says it is seeking a small number of software-company partners to work on difficult tasks that current agents do not reliably complete. Any further announcement would clarify which products and workflows are included, what data and safeguards are used, and how performance is evaluated beyond a limited task set.
For companies considering agents in contract, finance or customer-record systems, the immediate practical questions are which criteria the model fails, whether any missed requirement blocks use, how a human checks the result, and what audit trail records the agent’s actions. The Ironclad post does not answer those deployment questions. The next meaningful evidence would be a broader evaluation that reports complete-task reliability, specific failure types and measured outcomes under clearly described human oversight.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Ironclad in this story?
Ironclad is a contract-management software company. OpenAI’s post describes using hosted copies of its product as a research environment for training and evaluating models on contract-related workflows.
Does GPT-6 Astra complete 55% of tasks?
No. OpenAI reported that Astra met an average 55% of evaluation criteria across the tasks. That is not the share of tasks it completed successfully, and the source does not provide a general complete-task success rate.
Did the model save customers time?
The reported 19.2-minute figure is a simulated estimate per attempt, not measured customer time savings. The source does not establish how long a real workflow would take after a person reviews the model’s work.
What data did OpenAI say it used?
OpenAI said it built synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database and filtered out personal information. It said it used no OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.
Is the system ready to run contract workflows without review?
The published results do not establish that. OpenAI’s post says human oversight still matters, and the reported average criteria score leaves open which requirements were missed. The source does not describe a live, unsupervised deployment.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
