GPT 6.1 Sol ARC scores show why test setup matters
ARC Prize reports a large gap between GPT 6.1 Sol test harnesses even at the same reasoning setting

GPT 6.1 Sol posts sharply different ARC AGI 3 scores depending on the software setup used to run the test. ARC Prize’s verified results put the model at 52.73 percent with the Standard harness and 96.18 percent with the Provider Adapter harness when both use maximum reasoning effort.
That is a gap of 43.45 percentage points in the published result rows. It makes the harness label essential when judging claims about the model’s reasoning performance.
The same reasoning setting produces two different results
The result page lists ten configurations across five effort levels and two harnesses. At extra high reasoning, the Standard result is 39.93 percent and the Provider Adapter result is 96.37 percent. The adapter’s best listed score therefore comes below the maximum effort setting.

| Reasoning effort | Standard score | Provider Adapter score |
|---|---|---|
| Low | 3.92% | 82.79% |
| Medium | 10.57% | 91.02% |
| High | 26.72% | 94.99% |
| Extra high | 39.93% | 96.37% |
| Max | 52.73% | 96.18% |
Those results describe model and harness combinations. They cannot establish that changing one memory feature alone caused the difference, or that the higher score will carry over to every coding agent.
These appear on ARC Prize’s verified results page. The foundation says it runs ARC AGI 3 evaluations with its open benchmarking software. Its policy distinguishes that process from community submissions, which it does not routinely verify.
What the two harnesses preserve
A harness is the surrounding software that manages an agent’s interaction with a task. ARC Prize’s testing policy describes its Standard condition as a minimal interface shared across providers. It lets the model retain notes it chooses to carry forward through an environment.
The Provider Adapter condition adds context management features designed by the model provider. These include preserving opaque reasoning state between requests and compacting longer conversations so earlier work can remain useful. ARC Prize says the two conditions answer different evaluation questions and reports them separately.
In practical terms, comparing a provider adapter score for one model with a standard score for another would mix the model comparison with a change in the surrounding software.
The score measures more than completed puzzles
ARC AGI 3 asks agents to learn through interaction with unfamiliar environments. Its definition of a 100 percent score includes completing every game as efficiently as human participants. A percentage on this benchmark should therefore not be casually rewritten as the share of games solved.
ARC Prize’s policy uses a single run for each configuration rather than averaging repeated runs. Small score differences should be read with that limitation in mind.
The evaluation offers a specific signal about interactive learning. It does not establish general intelligence or guarantee reliability in a business workflow.
What developers should take from the results
OpenAI positions GPT 6.1 Sol as a cheaper option for complex agent work, with standard API rates of $2 per million input tokens and $10 per million output tokens. OpenAI also cautions that its research evaluations can differ from production behavior because prompts, tools and reasoning settings vary.
The ARC results make that qualification concrete. Teams comparing models should record the harness, reasoning setting and context retention method alongside the score. Testing a representative workflow with the intended production setup will be more useful than treating the highest available number as a universal property of the model.
See our GPT 6.1 Sol launch coverage for its published pricing and access details.
Featured image is an original AI generated editorial illustration. It is conceptual artwork rather than a product screenshot.



