ThinkingBox tests repeatability of Claude Opus 5.5
ThinkingBox compares average completion and repeated success

Microsoft and Hugging Face’s October 3 ThinkingBox update reports Claude Opus 5.5 results on an existing benchmark for agents handling business records.
Opus 5.5 succeeded on 67.16 percent of attempts across 507 tasks, versus 66.50 percent for Opus 5. Both passed all 20 recorded attempts on 241 tasks, or 47.53 percent.
Average success and repeated success
The average counts successful attempts. The stricter measure counts tasks completed in every recorded trial. Equal totals do not show that both models passed the same tasks.
The blog does not establish statistical significance for the small average gain or detail the new model’s exact run configuration.
What the benchmark measures
The underlying paper, first submitted August 20 and revised October 1, describes synthetic tasks from a private source collection. They are not a random sample of enterprise work.
Trials reset the database, but simulated user phrasing varies. Of 507 tasks, 477 judge only database outcomes and can pass despite an incorrect explanation to the user.
These observed results cannot guarantee future performance. ByteForward inspected the sources without rerunning the benchmark.
Illustrative network equipment photograph by Tyler, published April 8 2023 under the Unsplash License. It does not depict the ThinkingBox experiments.



