Inspect Robots study compares AI models on physical tasks
A new preprint evaluates robot task completion and refusals using an existing open source framework.

Researchers describe physical tests of language model based robot policies in an October 5 preprint about Inspect Robots. The evaluation framework had already been available for three months.
What the physical tests measured
The study uses an I2RT YAM robot with two arms. Five models tackle four manipulation tasks. Six models face four potentially unsafe instructions, with 20 trials per task and model.
The authors report manipulation score gains of 33 percentage points for Astra over GPT 5.6 Sol and 24 for Opus 5.5 over Opus 5. Their safety evaluation separately measures whether models refuse instructions. On three of four unsafe tasks, no model refused in more than 5 percent of trials.
A shared record of each experiment
The project repository describes a framework that separates the policy being tested from the robot or simulator running it. Records include the configuration, software versions and model conversation, helping researchers inspect what happened during a trial.
The code is available under the MIT license. Its maintainers warn that development is still early and interfaces can change between releases. They recommend pinning a version before depending on it.
The paper provides no human performance baseline. Results cover one robot setup and a small task set. ByteForward has not reproduced the experiments.
Illustrative archival photograph of the NIST Dexterous Manipulation Testbed by Falco/NIST, taken July 10, 2013. Public domain in the United States. Resized and converted to WebP. The photograph shows different equipment from the study.



