DOE selects four autonomous science robotics testbeds

The Department of Energy selected four laboratory teams on October 8 to develop robotics testbeds for autonomous science. Its announcement describes $30 million in total funding, including $2 million in fiscal 2026. Funding in later years depends on congressional appropriations. Selection begins award negotiations and does not guarantee an award or funding.
For researchers, the practical question is how much of an experiment a machine can manage reliably when conditions change. Moving a sample is one step. Knowing whether a measurement is usable, preserving the evidence behind that decision and recovering from a problem are separate challenges. Combining those responsibilities creates a much harder engineering problem than demonstrating an isolated robot skill.

Four projects tackle different problems
Argonne’s MAESTRO will develop reusable robotic skills and test their transfer between experimental settings. Brookhaven’s DART will examine reliability and scientific validity, using simulated adversarial conditions and a catalog of verified failures.
Oak Ridge’s TRACE will connect instruments, robots, agents and computing through standardized interfaces, with records that let others reproduce evaluations. SLAC’s SPIRE will combine sample handling, detectors, beam controls and computing for photon and electron facilities.
DOE expects outputs including open software, interfaces, datasets, trained models and benchmark tasks. Those planned resources could give other laboratories something concrete to evaluate and reuse.
The distinction matters when deciding what success should look like. A robot could complete its movement while a sample ends up poorly positioned. Software could produce an answer without leaving enough information for another researcher to check it. A useful evaluation needs to catch both kinds of failure.
SLAC has groundwork to build on
In its account of SPIRE, SLAC identifies sample manipulation, beam adjustment and analysis during experiments as the main areas of work. Existing automated beam characterization can already return results in about five minutes. The proposed autonomy layer would coordinate adjustments across existing control systems.
SLAC also describes a battery imaging demonstration in which agents moved data to another DOE laboratory, reconstructed a sample in three dimensions and used a vision model to identify particles. At a second SSRL beamline, agents monitor an instrument while a human retains control.
These examples establish a starting point for SPIRE. They do not establish that its complete laboratory platform is operational. SLAC says existing hardware safety systems will remain in place, with AI operating within those limits and researchers setting scientific goals.
What the testbeds need to measure
The original solicitation, issued May 14, sets out a useful scorecard. It asks for reproducibility, provenance, time to a result and the average interval between human interventions. It also calls for measurement of task success and generalization across tasks or sites.
The proposed infrastructure can combine physical equipment with simulation and digital twins. The call lists safety measures including operating limits, interlocks and safe behavior when something goes wrong. DOE also describes training for operators and researchers, making adoption part of the program rather than treating the robot as a finished appliance.
The solicitation allows approaches such as learning from human demonstrations and adapting foundation models to laboratory tasks. Those are possible methods, not a statement that every selected team will use the same model or training recipe. The original call seeks infrastructure that supports multiple workflows and can be extended over time.
The call excludes software only AI efforts without evaluation on, or joint design with, physical robotic systems in representative laboratory workflows. That requirement anchors the program in experimental operations.
For anyone evaluating the eventual results, the most revealing comparisons will separate speed from supervision. An experiment that finishes sooner but needs an expert to rescue every unusual case may offer limited operational savings. Reporting interventions alongside successful runs would make that tradeoff visible.
Readers should also look for what happens after a failed attempt. Can an operator identify the step that went wrong and understand why the system continued or stopped? That evidence would make a stronger case for dependable laboratory autonomy than a polished demonstration of the easiest run.







