AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: How To Read Ironclad’s Terms On OpenAI’s Software-Based Agent Training on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI and contract-management company Ironclad tested training a frontier model on legal, commercial and procurement workflows inside hosted copies of Ironclad’s product. OpenAI reports GPT-6 Astra met an average 55% of task criteria, while estimated task times were simulated rather than measured customer savings. The work also signals OpenAI’s interest in partnering with software firms to train agents on specialized workflows.

OpenAI said it trained GPT-6 Astra on legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software, reporting an average of 55% of task criteria met. The October 6 post describes a research approach that puts models inside specialized business software, but its results do not establish that the agent is ready to handle contract workflows without human review.

Ironclad employees and OpenAI staff who use the product selected 11 tasks, including setting up nondisclosure agreements, creating procurement approval processes and updating a reusable clause based on a requester’s jurisdiction. OpenAI estimated that an experienced user would need 30 to 40 minutes for each task. Depending on the task, evaluators scored model output against 8 to 50 criteria.

OpenAI said Ironclad supplied hosted copies of its software for model practice. It created synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database and said it filtered out personal information. OpenAI also said it used no OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.

In OpenAI’s reported comparison, GPT-5.6 Sol, at the “high” setting, met an average 41.6% of criteria, while GPT-6 Astra, at “max,” met 55%. Astra’s estimated time per attempt was 19.2 minutes, compared with 37 minutes for GPT-5.6 Sol. An internal OpenAI model used during Astra’s development reached 63.7%, and Astra met about 94% of criteria on one showcase task. These are shares of criteria met, not percentages of tasks completed successfully.

At a glance
reportWhen: Published October 6; results and partne…
The developmentOpenAI published details of an Ironclad collaboration in which models practised contract-related tasks in hosted copies of the company’s software.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Business Workflow Scores Matter

The experiment addresses a central challenge for AI agents: completing a multi-step task in a specialized product while retaining the business rules that determine whether the result is valid. In contract and procurement work, missing a single approval or review can undermine the whole process. A score of 55% of criteria met therefore cannot be read as evidence that an average workflow is more than half ready for use.

OpenAI’s post itself says human oversight remains important when an agent can lose track of a business rule. Ironclad CTO Sunita Verma emphasized that agents must preserve “the controls teams rely on.” The results suggest progress on a difficult research problem, but the source does not show that the system can reliably produce correct outcomes in live business operations.

The collaboration also gives software vendors a reason to weigh both benefits and risks before providing research environments. Training may help an agent work more effectively in a vendor’s product, potentially making that product more useful. At the same time, if customers interact through an agent rather than the product’s screens, the software company’s lasting value may depend more on its underlying rules, records, data model and controls than on its interface.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Ironclad Became the Test Environment

The OpenAI post’s title refers to Ironclad, the contract-management software company, not to a newly announced hardened agent framework. OpenAI framed the work as training models to understand business rules, carry out multi-step workflows in specialized software and check completed work against original requirements.

The reported evaluation covered a limited set of 11 research tasks rather than Ironclad workflows generally. OpenAI’s estimated time figures are also not measured productivity outcomes. The post says the estimates were simulated using assumed processing and generation speeds; they do not show how long customers actually saved. That distinction matters when comparing the model’s estimated 19.2 minutes with an experienced user’s estimated 30 to 40 minutes.

OpenAI’s post also invited a small number of software companies to work on tasks current agents cannot reliably complete. It asked potential partners to bring a concrete example of failure, people with deep knowledge of the work, a secure test environment and data suitable for research. The invitation points toward further software-specific training, but does not itself confirm additional partnerships or deployments.

“Agents have to preserve “the controls teams rely on.””

— Sunita Verma, Ironclad CTO

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Evaluation Cannot Establish

The published figures do not show how often Astra completed an entire task correctly, which specific criteria it missed across the 11 tasks, or whether it would perform similarly on a broader set of real-world workflows. Averages can conceal failures that matter greatly in a particular contract or approval process. OpenAI’s reported 94% result applies to one showcase task and should not be treated as the model’s general performance.

The time comparison is simulated, not a measurement of customer work, and the source provides no demonstrated net time saving after human checks. It is also unclear whether OpenAI or Ironclad has tested the system in live customer operations, what level of review would be required for each use case, or whether further software-company partnerships have been agreed. OpenAI’s description of its data practices is attributable to the company; the source does not provide an independent audit of those practices.

Amazon

AI-powered procurement approval software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Software Partners May Test

OpenAI says it is seeking a small number of software-company partners to work on difficult tasks that current agents do not reliably complete. Any further announcement would clarify which products and workflows are included, what data and safeguards are used, and how performance is evaluated beyond a limited task set.

For companies considering agents in contract, finance or customer-record systems, the immediate practical questions are which criteria the model fails, whether any missed requirement blocks use, how a human checks the result, and what audit trail records the agent’s actions. The Ironclad post does not answer those deployment questions. The next meaningful evidence would be a broader evaluation that reports complete-task reliability, specific failure types and measured outcomes under clearly described human oversight.

Amazon

AI contract analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Ironclad in this story?

Ironclad is a contract-management software company. OpenAI’s post describes using hosted copies of its product as a research environment for training and evaluating models on contract-related workflows.

Does GPT-6 Astra complete 55% of tasks?

No. OpenAI reported that Astra met an average 55% of evaluation criteria across the tasks. That is not the share of tasks it completed successfully, and the source does not provide a general complete-task success rate.

Did the model save customers time?

The reported 19.2-minute figure is a simulated estimate per attempt, not measured customer time savings. The source does not establish how long a real workflow would take after a person reviews the model’s work.

What data did OpenAI say it used?

OpenAI said it built synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database and filtered out personal information. It said it used no OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.

Is the system ready to run contract workflows without review?

The published results do not establish that. OpenAI’s post says human oversight still matters, and the reported average criteria score leaves open which requirements were missed. The source does not describe a live, unsupervised deployment.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Building Resilient Systems With Sam Newman

A Pragmatic Engineer podcast episode with Sam Newman covers distributed-system resilience, microservice trade-offs and software development with AI.

The €55,000 Bug in Your AI Agent: It Didn’t Read the Docs

A live benchmark buried a €55,000 fact two documents deep in a company’s files. Only the AI agents that actually read the docs closed the deal — everyone else lost it automatically.

Before an AI Touches Your Workflow, Put It Through a Bad Week

Firmulate’s AI company wargame tests whether models can act on what they know. See how a read-only pilot can expose weak points before agents reach live systems.

Launching Meta Enterprise Platform

Search and coverage interest around “Launching Meta Enterprise Platform” is rising, but the cause and any related Meta announcement remain unconfirmed.