RobotWorld tests AI agents across 84 simulated tasks
Five models leave 63 tasks unsolved in the reported evaluation.

RobotWorld’s October 7 preprint reports five models evaluated on 84 simulated tasks, with one episode per model and task. GPT 6 Astra completed 16 and Claude Opus 5.5 completed 13. Their combined successes covered 21 tasks. The other models added none beyond that set, leaving 63 unsolved.
The other models were Kimi K3, DeepSeek V4.1 Flash and Gemini 3.8 Flash. Tasks spanned manipulation, mobile manipulation, locomotion, driving and aerial control.
Physics paused during reasoning. The results describe these trials rather than repeated reliability or performance on physical robots. Some interaction budgets were calibrated using Astra records. The authors’ explanations of differences between models remain hypotheses.
The arXiv record gives a submission time, which does not establish the first public release clock.



