Accessibility Adjustments

Use these optional tools to adjust reading and display preferences. These tools cannot resolve every accessibility barrier. Please contact the website owner if you need assistance.

  • Text adjustments
  • Content scaling 100%
  • Font size 100%
  • Line height 100%
  • Letter spacing 100%
  • Colour adjustments
  • Orientation adjustments

RobotWorld tests AI agents across 84 simulated tasks

Five models leave 63 tasks unsolved in the reported evaluation.

Listen to this article

RobotWorld’s October 7 preprint reports five models evaluated on 84 simulated tasks, with one episode per model and task. GPT 6 Astra completed 16 and Claude Opus 5.5 completed 13. Their combined successes covered 21 tasks. The other models added none beyond that set, leaving 63 unsolved.

The other models were Kimi K3, DeepSeek V4.1 Flash and Gemini 3.8 Flash. Tasks spanned manipulation, mobile manipulation, locomotion, driving and aerial control.

Physics paused during reasoning. The results describe these trials rather than repeated reliability or performance on physical robots. Some interaction budgets were calibrated using Astra records. The authors’ explanations of differences between models remain hypotheses.

The arXiv record gives a submission time, which does not establish the first public release clock.

Maya Chen
Maya Chen

Maya Chen is focused on covering AI models, research, and the evidence behind new capabilities. Maya follows model launches, benchmarks, open weights, and scientific uses of AI with one question in mind. What changed, and how would we know? The voice is curious and exacting, with a soft spot for elegant technical ideas and little patience for a leaderboard without context.

Leave a Reply

Your email address will not be published. Required fields are marked *

Gravatar profile