BootLoops brings reusable scientific tools to AI research
Matthew Schwartz releases a toolkit built around checkable calculations and human direction

Matthew Schwartz released BootLoops 1.0 on October 1, bringing a reusable collection of scientific software and validation protocols to AI research. The public toolkit is designed for language models to operate, with instruments for quantitative problems that can be checked against independent calculations.
In an October 1 guest post, Schwartz describes building the project with Claude and then working with specialists to find worthwhile applications beyond physics. He was a visiting researcher at Anthropic during the work. BootLoops is his independently maintained project, rather than an officially supported Anthropic product.
Reusable software is the central idea
The harness documentation describes a collection that combines existing scientific engines, ports and newly written tools. It is designed to work with different language models. Each new problem can leave behind improved software and instructions that another project can reuse. That makes the lasting asset the accumulated methods and checks, alongside the model that operates them.
The repository specifies Python 3.12, with some Julia components. Its release testing covered Linux environments, while macOS was not tested. A provided validation script runs package checks, but some require missing external engines or data. Researchers should inspect those results before treating an installation as ready.
An exact answer still needs a meaningful question
The project overview emphasizes independent numerical checks, explicit error bounds and scripts that others can rerun. For suitable integral calculations, its standard calls for both a complete functional answer and code that evaluates it to arbitrary precision. Other tasks need different tests, so the toolkit does not turn every research question into a problem with one mechanically verifiable answer.
Schwartz says specialists often found early results technically impressive but scientifically uninteresting. Their input redirected the work. His account also warns that automated checks remain fallible and describes several highlighted applications as undergoing further exploration and verification. Those qualifications matter when evaluating the broader discovery claims.
Economics shows both the promise and the limits
A September working paper by Schwartz, Isaiah Andrews and Jesse Shapiro provides one concrete example. Their workflow examined 4,452 published replication packages from five economics journals through code translation and numerical checks. The authors report making a calculation more than ten times faster in 496 articles.
These are reported results from an ongoing research workflow. The paper identifies limits from rounding and unavailable data, and says the workflow does not aim to evaluate articles. A calculation running faster does not by itself verify its assumptions or establish the scientific value of the result.
Start with a result someone else can check
A sensible first trial would pair a published calculation with a known answer and an independent reference implementation. Record the software version, inputs, resource cost and failed checks. The useful question is whether another researcher can reproduce the result and explain its significance without trusting the agentโs own verdict.
There is also a practical security limit. The repository warns that input files can execute commands. Its integrity checks are not a security boundary, so unfamiliar inputs need isolation rather than confidence in a numerical certificate.
Featured image is an original AI generated editorial illustration.



