OpenAI tests Astra on Ironclad contracting tasks
An 11 task research evaluation compares Astra and Sol on contracting workflows.

OpenAI detailed an Ironclad research collaboration on October 6, using 11 contracting tasks to train and evaluate AI agents that operate software.
Tasks cover agreement setup, approval rules and reusable legal clauses. Each was scored against 8 to 50 criteria.
OpenAI reports that GPT 6 Astra at Max reasoning averaged a 55.0% rubric score, compared with 41.6% for GPT 5.6 Sol at High reasoning.
Estimated average time per attempt was 19.2 minutes for Astra and 37 minutes for Sol. These are simulated estimates, rather than measured customer savings, and the findings cover only the 11 research tasks.
Synthetic tasks drew on public EDGAR contracts filtered for personal information. OpenAI says its customer data and internal contracts, along with nonpublic Ironclad customer data and contracts, were excluded. Human oversight remains important in complex contracting.
Archival January 2015 photograph by AlexKingCreative via Wikimedia Commons under CC0. Resized and converted to WebP. It illustrates computer use and does not show Ironclad, Astra or the contracting evaluation.



