Accessibility Adjustments

Use these optional tools to adjust reading and display preferences. These tools cannot resolve every accessibility barrier. Please contact the website owner if you need assistance.

  • Text adjustments
  • Content scaling 100%
  • Font size 100%
  • Line height 100%
  • Letter spacing 100%
  • Colour adjustments
  • Orientation adjustments

Inspect Robots study compares AI models on physical tasks

A new preprint evaluates robot task completion and refusals using an existing open source framework.

Listen to this article

Researchers describe physical tests of language model based robot policies in an October 5 preprint about Inspect Robots. The evaluation framework had already been available for three months.

What the physical tests measured

The study uses an I2RT YAM robot with two arms. Five models tackle four manipulation tasks. Six models face four potentially unsafe instructions, with 20 trials per task and model.

The authors report manipulation score gains of 33 percentage points for Astra over GPT 5.6 Sol and 24 for Opus 5.5 over Opus 5. Their safety evaluation separately measures whether models refuse instructions. On three of four unsafe tasks, no model refused in more than 5 percent of trials.

A shared record of each experiment

The project repository describes a framework that separates the policy being tested from the robot or simulator running it. Records include the configuration, software versions and model conversation, helping researchers inspect what happened during a trial.

The code is available under the MIT license. Its maintainers warn that development is still early and interfaces can change between releases. They recommend pinning a version before depending on it.

The paper provides no human performance baseline. Results cover one robot setup and a small task set. ByteForward has not reproduced the experiments.

Illustrative archival photograph of the NIST Dexterous Manipulation Testbed by Falco/NIST, taken July 10, 2013. Public domain in the United States. Resized and converted to WebP. The photograph shows different equipment from the study.

Maya Chen
Maya Chen

Maya Chen is focused on covering AI models, research, and the evidence behind new capabilities. Maya follows model launches, benchmarks, open weights, and scientific uses of AI with one question in mind. What changed, and how would we know? The voice is curious and exacting, with a soft spot for elegant technical ideas and little patience for a leaderboard without context.

Leave a Reply

Your email address will not be published. Required fields are marked *

Gravatar profile