🔍 Read the full analysis: OpenAI Training Agents Inside Software: Read Ironclad’s Terms Carefully on ThorstenMeyerAI.com
Get the little things that make your day delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
TL;DR
OpenAI described training GPT-6 Astra on 11 legal, commercial and procurement tasks inside hosted copies of contract-management software from Ironclad. Astra met an average 55% of task criteria, and its faster completion estimate was simulated, not measured customer time savings. The work also signals that OpenAI is seeking software partners to help train agents on specialized workflows.
OpenAI said on October 6 that it trained its GPT-6 Astra model on legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software. The report gives an early measure of how agents perform in specialized business applications: Astra met an average of 55% of evaluation criteria across 11 tasks, while OpenAI says its time estimates were simulated rather than measured in customer use.
OpenAI and Ironclad selected 11 tasks drawn from legal, commercial and procurement work. Examples included setting up nondisclosure agreements, creating procurement approval processes and updating a reusable contract clause to reflect a requester’s jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes to complete each task.
Ironclad provided hosted copies of its product for the models to use. OpenAI said it created synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database, with personal information filtered out. The company said the work did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data.
Tasks were scored against rubrics containing eight to 50 criteria, depending on complexity. OpenAI reported that GPT-5.6 Sol met an average 41.6% of criteria, compared with 55.0% for GPT-6 Astra. An internal OpenAI model used in Astra’s development reached 63.7%. On one showcased task, Astra met about 94% of the criteria. These scores measure criteria met, not the percentage of tasks completed successfully.
OpenAI is training agents inside your software. Read the fine print on Ironclad.
Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.
legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses
per task, experienced user (OpenAI estimate)
criteria per task — a rubric, not pass/fail
public SEC filings; no customer or non-public Ironclad data
The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.
37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.
Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.
Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.
Averages hide missed approvals.
Narrowest access; no self-escalation.
METR found agents spoofing tool-call records.
Measure the whole loop.
Public filings, not your contracts.
Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.
Why Partial Workflow Accuracy Matters
The results point to both progress and a gap between performing parts of a workflow and reliably completing business work. A 55% rubric score means Astra met just over half of the requirements on average; it does not establish that the agent can safely finish a contract or approval process on its own. In workflows with mandatory checks, a missed rule can undermine the whole result.
OpenAI’s example of procurement illustrates the issue: a process may require Finance approval above a spending threshold, Security review for some requests and Legal review for nonstandard terms. An agent that misses one of those controls may route a purchase incorrectly. OpenAI’s report says human oversight still matters when an agent loses track of a business rule during a task.
The time comparison also needs care. OpenAI reported estimated times of 37 minutes for GPT-5.6 Sol and 19.2 minutes for Astra, but said those figures were simulated estimates based on assumed processing and generation speeds. They are not observed customer savings, and the company said they cover the 11 research tasks rather than Ironclad workflows generally. The reported results do not show that customers can yet complete this work faster or with less review.
contract management software with AI integration
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Ironclad as a Model Training Site
The report is framed as work on teaching models to understand business rules, perform multi-step actions in specialized software and check whether the result meets the original requirements. Ironclad is a contract-management software company, not the name of a new OpenAI agent framework. The partnership placed models inside a vendor’s hosted product environment to practise tasks tied to real software workflows.
OpenAI’s report also describes a possible next stage: the company is seeking a small number of software partners whose applications contain tasks current agents cannot reliably complete. It asks potential partners to bring a concrete example of failure, people with deep knowledge of the work, a secure testing environment and data that can safely be used for research.
For software vendors, training agents inside their products may help identify where models fail and could make those products more useful if agents become dependable. It also puts attention on what the software provides beyond its screens: business rules, records, audit trails and controls. The report presents the full contracting platform as remaining important, but does not establish how future customer products or commercial arrangements will work.
As an affiliate, we earn on qualifying purchases.
What the Evaluation Does Not Show
The report does not establish whether Astra can reliably complete these tasks in live customer environments, how often it would miss individual high-priority requirements, or what level of human review would be needed in routine use. The average rubric score also does not identify, in the material provided, which criteria were missed across all tasks or how errors were weighted by consequence.
The reported time savings are not measured results from customers, and the 11 tasks do not represent every workflow in Ironclad or other business software. OpenAI’s stated data protections describe the inputs used for this research, but the report does not specify the terms future partners would use or how agents trained through later collaborations might be deployed. Claims about eventual productivity gains or changes to software products remain prospective, not demonstrated outcomes.
AI-powered contract review software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
OpenAI Seeks More Software Partners
OpenAI says it is inviting a small number of software companies to work on difficult tasks that current agents cannot reliably complete. Potential partners are expected to supply specific failure cases, domain experts, secure test environments and research-safe data. The report does not name additional partners or give a timetable for further evaluations.
For companies considering agents in contract, finance or customer-record systems, the immediate practical question is not only whether an agent can perform a task, but whether it can satisfy every required control and show what it did. Further results would need to clarify criteria-level performance, error handling and review requirements before the reported research could support broader claims about safe, dependable use.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did OpenAI announce about Ironclad?
OpenAI published a report describing how it trained GPT-6 Astra on 11 legal, commercial and procurement tasks using hosted copies of Ironclad’s contract-management software.
What does Astra’s 55% score measure?
It is the average share of rubric criteria met across the evaluated tasks. It is not the share of tasks completed and does not by itself establish that a workflow was safe or ready for use without review.
Did Astra save customers time?
The report does not establish customer time savings. OpenAI’s reported 19.2-minute estimate for Astra was simulated, based on assumed processing and generation speeds, and covered the research tasks.
What data did OpenAI say it used?
OpenAI said it built synthetic tasks from publicly filed contracts in the SEC’s EDGAR database and filtered out personal information. It said it used no OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.
What happens after the Ironclad research?
OpenAI says it is seeking a small number of software companies to help test agents on difficult workflows. It has not named additional partners or provided a timetable for the next results.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
