Accessibility Adjustments

Use these optional tools to adjust reading and display preferences. These tools cannot resolve every accessibility barrier. Please contact the website owner if you need assistance.

  • Text adjustments
  • Content scaling 100%
  • Font size 100%
  • Line height 100%
  • Letter spacing 100%
  • Colour adjustments
  • Orientation adjustments

GPT 6.1 Sol ARC scores show why test setup matters

ARC Prize reports a large gap between GPT 6.1 Sol test harnesses even at the same reasoning setting

Listen to this article

GPT 6.1 Sol posts sharply different ARC AGI 3 scores depending on the software setup used to run the test. ARC Prize’s verified results put the model at 52.73 percent with the Standard harness and 96.18 percent with the Provider Adapter harness when both use maximum reasoning effort.

That is a gap of 43.45 percentage points in the published result rows. It makes the harness label essential when judging claims about the model’s reasoning performance.

The same reasoning setting produces two different results

The result page lists ten configurations across five effort levels and two harnesses. At extra high reasoning, the Standard result is 39.93 percent and the Provider Adapter result is 96.37 percent. The adapter’s best listed score therefore comes below the maximum effort setting.

ARC AGI 3 scores for GPT 6.1 Sol showing Standard and Provider Adapter results at five reasoning levels
Chart by ByteForward using ARC Prize verified results. One run per configuration.
Reasoning effortStandard scoreProvider Adapter score
Low3.92%82.79%
Medium10.57%91.02%
High26.72%94.99%
Extra high39.93%96.37%
Max52.73%96.18%
Data behind the chart from the original evaluator

Those results describe model and harness combinations. They cannot establish that changing one memory feature alone caused the difference, or that the higher score will carry over to every coding agent.

These appear on ARC Prize’s verified results page. The foundation says it runs ARC AGI 3 evaluations with its open benchmarking software. Its policy distinguishes that process from community submissions, which it does not routinely verify.

What the two harnesses preserve

A harness is the surrounding software that manages an agent’s interaction with a task. ARC Prize’s testing policy describes its Standard condition as a minimal interface shared across providers. It lets the model retain notes it chooses to carry forward through an environment.

The Provider Adapter condition adds context management features designed by the model provider. These include preserving opaque reasoning state between requests and compacting longer conversations so earlier work can remain useful. ARC Prize says the two conditions answer different evaluation questions and reports them separately.

In practical terms, comparing a provider adapter score for one model with a standard score for another would mix the model comparison with a change in the surrounding software.

The score measures more than completed puzzles

ARC AGI 3 asks agents to learn through interaction with unfamiliar environments. Its definition of a 100 percent score includes completing every game as efficiently as human participants. A percentage on this benchmark should therefore not be casually rewritten as the share of games solved.

ARC Prize’s policy uses a single run for each configuration rather than averaging repeated runs. Small score differences should be read with that limitation in mind.

The evaluation offers a specific signal about interactive learning. It does not establish general intelligence or guarantee reliability in a business workflow.

What developers should take from the results

OpenAI positions GPT 6.1 Sol as a cheaper option for complex agent work, with standard API rates of $2 per million input tokens and $10 per million output tokens. OpenAI also cautions that its research evaluations can differ from production behavior because prompts, tools and reasoning settings vary.

The ARC results make that qualification concrete. Teams comparing models should record the harness, reasoning setting and context retention method alongside the score. Testing a representative workflow with the intended production setup will be more useful than treating the highest available number as a universal property of the model.

See our GPT 6.1 Sol launch coverage for its published pricing and access details.

Featured image is an original AI generated editorial illustration. It is conceptual artwork rather than a product screenshot.

Marcus Reid
Marcus Reid

Marcus Reid is focused on covering the money, rules, and institutional choices shaping AI. He runs from funding rounds and chip deals to regulation, lawsuits, leadership changes, and the business of building enormous computing systems. Marcus follows the incentives behind the announcement. Who pays, who gains leverage, and what changes for everyone else? The voice is direct, measured, and occasionally dry, especially when a grand promise arrives with very little detail.

Leave a Reply

Your email address will not be published. Required fields are marked *

Gravatar profile