Study tests flexible coordination for AI research agents
Runtime selection leads a small AI scientist comparison

An October 1 preprint tests whether AI research teams benefit from choosing who acts next during a task. Runtime Agent Coordination had its highest mean ResearchClawBench scores with flexible selection, before adding contracts and verification.
The comparison used Agent Laboratory and EvoScientist on ten tasks each, and ARK on five. All runs used one seed, with DeepSeek V4 Pro performing research and GPT 5.5 judging outputs.
For Agent Laboratory, the mean score rose from 5.47 under native execution to 12.08 with runtime selection.
More checks did not always help
Adding contracts and advisory checks lowered each hostโs average on that benchmark. A small DiscoveryBench transfer test favored the fuller setup. The results do not isolate contracts from verification.
Budget equality was intended, but the authors could not independently verify every archived resource ceiling. The study does not establish general reliability or compute matched gains.
What developers can inspect
The MIT licensed repository separates the coordination layer from the three host projects. It records their revisions and source tree hashes, but warns that these pins do not freeze container base images or downloaded resources.
Its changelog dates the alpha to September 30. Earlier commits document development, while an October 3 project page explains the same work.
Repository instructions separate offline checks from live research runs. Live episodes let agents write and execute code inside containers. The maintainers advise an isolated machine, keeping credentials out of task bundles and a budget for every episode.
ByteForward reviewed the paper and repository documents without running those experiments.
Illustrative laboratory photograph by Chidera Faustina Okeke, published on December 21 2025 under the Unsplash License. The people shown are not identified as study participants.



