Sonnet 5.5 enters Agent Arena with a task cost tradeoff
New Agent Arena results place Sonnet 5.5 near the top while showing why effort settings and measured task spending matter alongside token prices.

Arena added Claude Sonnet 5.5 at Max effort to its Agent Arena leaderboard on October 2, bringing a new independent measurement of how the model handles work with tools. The dated changelog records the addition separately from its text and vision results.
The overall leaderboard snapshot reviewed by ByteForward placed Sonnet third, with 5,219 sessions and a net improvement score of 12.52 percent. Its median task cost was $2.78, compared with $1.58 for Opus 5.5 at High effort. The figures were retrieved on October 2 at about 8 pm UTC.
A strong result with visible uncertainty
Fable 5.1 at Max effort led the displayed ranking at 14.31 percent, followed by Opus 5.5 at 13.82 percent. Sonnetโs reported interval was plus or minus 3.09 percentage points, overlapping the intervals for both models above it. Its displayed rank range ran from first through ninth. The ordering should therefore be read alongside the uncertainty rather than as a decisive separation.
Arenaโs methodology evaluates models inside its own agent environment using real user sessions. It randomizes the selected orchestrator and estimates its contribution to outcomes against a baseline distribution. The headline score combines several behavioral signals. A score of 12.52 percent is an estimated net improvement, not a claim that the model completed only that share of tasks.
The estimates include 95 percent confidence intervals, and newer observations receive more weight. These choices make the board responsive to current usage, while also making it a changing measurement rather than a fixed certification of model quality.
Task spending tells a different part of the story
Sonnetโs median output volume was 116,700 tokens per task in this snapshot, compared with 26,200 for Opus. Fableโs median task cost was $4.62. These are observations for the listed effort settings and Arena workloads, rather than quotes for running the same job in another application.
Arenaโs cost documentation explains why it measures tasks instead of relying only on token prices. A session can contain several requests, and later tasks inherit more conversation context. To reduce that distortion, the cost calculation includes at most the first three tasks in a session.
The displayed P50 figure is the median, so it is not an average spending forecast or an upper limit. Arena also provides other cost percentiles and separate categories for coding, chat and office work. The same model can therefore look different depending on the workload and the part of the cost distribution being examined.
Effort settings are part of the comparison
Anthropicโs September 28 launch documentation prices Sonnet at $2 per million input tokens and $10 per million output tokens. Opus is listed at $4 and $20. Anthropic also says Sonnetโs lower effort settings offer its strongest cost advantage, while higher settings can approach Opus performance at similar task costs.
The defaults matter. Anthropic lists Medium effort for Sonnet in its apps and Claude Code, and High for the Claude Platform. Arenaโs new entry uses Max. Comparing that result with Opus at High evaluates two particular configurations. It does not isolate a model family difference or establish the bill a customer should expect at the default setting.
Feedback signals have specific meanings
The current signal definitions distinguish explicit confirmation from ordinary positive feedback. Confirmed success requires the user to answer the task completion prompt. A silent ending contributes no confirmation even if the result was useful. Praise and complaints are measured separately.
Steerability measures the burden of corrections, including whether users need to repeat them. Bash recovery examines how quickly an agent resolves command errors. Tool hallucination concerns calls to tools that do not exist in its available set. Together, these signals describe aspects of the interaction rather than proving that every generated artifact is correct.
For teams considering Sonnet, the useful next step is to test the effort setting they would actually deploy on representative tasks. Record spending, corrections and whether the result passes review. The new ranking offers a practical comparison point, but choosing a configuration still requires matching that evidence to the work.
ByteForwardโs Sonnet migration coverage explains the approaching retirement deadline. Its Argon benchmark analysis examines another case where strong scores and agent performance need separate scrutiny.
Illustration is original AI generated artwork created for ByteForward. It represents different paths through a task and does not depict a product or measured result.



