GPT 6.1 Sol joins Agent Arena with a lower median task cost
Arena’s October 2 addition puts GPT 6.1 Sol alongside its predecessor, with a lower observed median task cost and uncertainty that matters for interpretation.

Arena’s October 2 leaderboard changelog records the addition of GPT 6.1 Sol at Max reasoning effort to Agent Arena. The entry gives developers another independent view of its behavior on tasks that involve tools, extending the evidence beyond its launch claims and ARC results.
The overall snapshot checked on October 2 at about 9 pm UTC lists a median task cost of $0.57 for GPT 6.1 Sol, compared with $0.91 for GPT 6 Sol, both at Max effort. Their listed token rates are identical at $2 per million input tokens and $10 per million output tokens.
That difference is useful when planning a trial. It suggests that comparing token rates alone can miss differences in what a task costs. These medians describe the work sampled by Arena, and cannot promise the same saving on a particular customer job.
What the current results show
GPT 6.1 Sol appears fifth overall, with 6,278 sessions and estimated net improvement of 11.23 percent, plus or minus 2.76 percentage points. GPT 6 Sol has 6,106 sessions and 9.71 percent, plus or minus 2.37 points. The reported intervals overlap, so the ordering does not establish a decisive performance gap.
Median output volume is 26,400 tokens per task for GPT 6.1 Sol versus 28,500 for GPT 6 Sol.
The dated addition identifies an evaluation entry rather than a model release. It does not establish the time of the first evaluation session. ByteForward had already observed these Sol figures in an earlier October 2 snapshot, before the changelog sentence appeared.
Success means something specific here
Arena’s current signal definitions require an explicit answer to its task completion question for confirmed success. A silent ending contributes no confirmation, even if the work was useful. Ordinary praise is recorded separately. This makes the signal sensitive to how people interact with the completion prompt, not just whether an output looks correct.
The percentages in the leaderboard use Arena’s causal evaluation methodology. It randomizes the orchestrator model inside its agent environment and estimates contributions relative to a baseline distribution. Net improvement combines several behavioral signals. The displayed effect estimates should not be read as raw percentages of all tasks completed.
Arena reports 95 percent confidence intervals and gives more weight to recent observations. Its harness, available tools and sampled user tasks are therefore part of the result. A different application may produce different spending and outcomes even when it selects the same model and reasoning setting.
Why the cost unit matters
A task can include several requests and corrections. Arena’s cost explanation describes splitting sessions into task boundaries, then measuring at most the first three tasks in each session. That limit reduces distortion from the accumulated context that makes later tasks more expensive.
The P50 figure is a median. It is not a spending ceiling or the average bill for a production workload. Arena also offers other percentiles and category views for Code, Chat and Work. Its cost observations come from live use rather than a static test that reruns an identical job for every model.
The practical comparison is to run both Sol versions on work a team already understands. Keep the environment, effort setting and acceptance criteria fixed. Count failed attempts, corrections, total spending and the review needed before using the result. Arena provides a reason to investigate the newer model’s task economics, while the local trial establishes whether the advantage carries over.
ByteForward’s GPT 6.1 Sol launch report covers pricing and availability. Our ARC analysis shows how the surrounding software can change a benchmark result. The separate Sonnet Agent Arena report examines another configuration where token prices and observed task costs tell different stories.
Illustration is original AI generated artwork created for ByteForward. The drafting instruments and resource counters are conceptual and do not depict a product, measured result or exact comparison.



