OpenAI Training Agents Inside Software: Read Ironclad’s Terms Carefully
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: OpenAI Training Agents Inside Software: Read Ironclad’s Terms Carefully on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get the little things that make your day delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI described training GPT-6 Astra on 11 legal, commercial and procurement tasks inside hosted copies of contract-management software from Ironclad. Astra met an average 55% of task criteria, and its faster completion estimate was simulated, not measured customer time savings. The work also signals that OpenAI is seeking software partners to help train agents on specialized workflows.

OpenAI said on October 6 that it trained its GPT-6 Astra model on legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software. The report gives an early measure of how agents perform in specialized business applications: Astra met an average of 55% of evaluation criteria across 11 tasks, while OpenAI says its time estimates were simulated rather than measured in customer use.

OpenAI and Ironclad selected 11 tasks drawn from legal, commercial and procurement work. Examples included setting up nondisclosure agreements, creating procurement approval processes and updating a reusable contract clause to reflect a requester’s jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes to complete each task.

Ironclad provided hosted copies of its product for the models to use. OpenAI said it created synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database, with personal information filtered out. The company said the work did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data.

Tasks were scored against rubrics containing eight to 50 criteria, depending on complexity. OpenAI reported that GPT-5.6 Sol met an average 41.6% of criteria, compared with 55.0% for GPT-6 Astra. An internal OpenAI model used in Astra’s development reached 63.7%. On one showcased task, Astra met about 94% of the criteria. These scores measure criteria met, not the percentage of tasks completed successfully.

At a glance
reportWhen: Published October 6; further software p…
The developmentOpenAI published a report on October 6 describing how it trained GPT-6 Astra on professional workflows inside Ironclad’s contract-management software and inviting other software companies to partner on similar work.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Partial Workflow Accuracy Matters

The results point to both progress and a gap between performing parts of a workflow and reliably completing business work. A 55% rubric score means Astra met just over half of the requirements on average; it does not establish that the agent can safely finish a contract or approval process on its own. In workflows with mandatory checks, a missed rule can undermine the whole result.

OpenAI’s example of procurement illustrates the issue: a process may require Finance approval above a spending threshold, Security review for some requests and Legal review for nonstandard terms. An agent that misses one of those controls may route a purchase incorrectly. OpenAI’s report says human oversight still matters when an agent loses track of a business rule during a task.

The time comparison also needs care. OpenAI reported estimated times of 37 minutes for GPT-5.6 Sol and 19.2 minutes for Astra, but said those figures were simulated estimates based on assumed processing and generation speeds. They are not observed customer savings, and the company said they cover the 11 research tasks rather than Ironclad workflows generally. The reported results do not show that customers can yet complete this work faster or with less review.

Amazon

contract management software with AI integration

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Ironclad as a Model Training Site

The report is framed as work on teaching models to understand business rules, perform multi-step actions in specialized software and check whether the result meets the original requirements. Ironclad is a contract-management software company, not the name of a new OpenAI agent framework. The partnership placed models inside a vendor’s hosted product environment to practise tasks tied to real software workflows.

OpenAI’s report also describes a possible next stage: the company is seeking a small number of software partners whose applications contain tasks current agents cannot reliably complete. It asks potential partners to bring a concrete example of failure, people with deep knowledge of the work, a secure testing environment and data that can safely be used for research.

For software vendors, training agents inside their products may help identify where models fail and could make those products more useful if agents become dependable. It also puts attention on what the software provides beyond its screens: business rules, records, audit trails and controls. The report presents the full contracting platform as remaining important, but does not establish how future customer products or commercial arrangements will work.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Evaluation Does Not Show

The report does not establish whether Astra can reliably complete these tasks in live customer environments, how often it would miss individual high-priority requirements, or what level of human review would be needed in routine use. The average rubric score also does not identify, in the material provided, which criteria were missed across all tasks or how errors were weighted by consequence.

The reported time savings are not measured results from customers, and the 11 tasks do not represent every workflow in Ironclad or other business software. OpenAI’s stated data protections describe the inputs used for this research, but the report does not specify the terms future partners would use or how agents trained through later collaborations might be deployed. Claims about eventual productivity gains or changes to software products remain prospective, not demonstrated outcomes.

Amazon

AI-powered contract review software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

OpenAI Seeks More Software Partners

OpenAI says it is inviting a small number of software companies to work on difficult tasks that current agents cannot reliably complete. Potential partners are expected to supply specific failure cases, domain experts, secure test environments and research-safe data. The report does not name additional partners or give a timetable for further evaluations.

For companies considering agents in contract, finance or customer-record systems, the immediate practical question is not only whether an agent can perform a task, but whether it can satisfy every required control and show what it did. Further results would need to clarify criteria-level performance, error handling and review requirements before the reported research could support broader claims about safe, dependable use.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did OpenAI announce about Ironclad?

OpenAI published a report describing how it trained GPT-6 Astra on 11 legal, commercial and procurement tasks using hosted copies of Ironclad’s contract-management software.

What does Astra’s 55% score measure?

It is the average share of rubric criteria met across the evaluated tasks. It is not the share of tasks completed and does not by itself establish that a workflow was safe or ready for use without review.

Did Astra save customers time?

The report does not establish customer time savings. OpenAI’s reported 19.2-minute estimate for Astra was simulated, based on assumed processing and generation speeds, and covered the research tasks.

What data did OpenAI say it used?

OpenAI said it built synthetic tasks from publicly filed contracts in the SEC’s EDGAR database and filtered out personal information. It said it used no OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.

What happens after the Ironclad research?

OpenAI says it is seeking a small number of software companies to help test agents on difficult workflows. It has not named additional partners or provided a timetable for the next results.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The clause. How a contractual definition of AGI met the capital built on top of it.

OpenAI’s 2019 contract’s AGI clause was defused through amendments, transforming from a doomsday trigger into a verification step amid capital pressures.

Anchor. The Schwarz Group model.

An in-depth analysis of Schwarz Group’s €11B AI data center investment and its potential as a scalable European industrial-anchor model.

Best Quiet CPU Coolers for Sustained AI/Compute Loads

Discover top quiet CPU coolers ideal for sustained AI and compute loads, including air and liquid options, with expert recommendations for 2026.

14 Best AI Automation Software Tools for Smarter Workflows in 2026

Discover the 14 best AI automation software tools for 2026, focusing on agent orchestration, workflow integration, and practical application for smarter work.