BrickBench puts AI brick builders to the test

The BrickBench preprint asks coding agents to design LEGO assemblies from 300 prompts. Its evaluation covers eleven agents across small models, larger sets and a fixed parts inventory. Two specialist baselines cover only small models. Scores cover simulated validity, prompt alignment and design.
Submitted on October 8, the paper appeared in arXiv’s October 9 listing.
What a passing score means
In the authors’ tool ablation, OpenAI’s Astra passed simulated validity on every prompt with BrickAgent. Removing those tools cut that rate to 40 percent. Connection strength is untested, so passing assemblies could still sag or separate when built.
Five raters identified human designs in 323 of 360 judgments. Those pairs matched part counts rather than prompts, and excluded the later Opus 5.5 and GPT 6.1 Sol additions. That limits conclusions about performance on identical design briefs.
What reproducing it requires
Peter Kulits’s MIT licensed repository includes the environment, prompts, validation and evaluation scripts, and a Dockerfile. Its validator names the parts involved in collisions, connection mismatches, inventory overruns or instability.
PyPI lists BrickAgent 0.1.0 as released on October 8, with a source archive and wheel. The package declares Python 3.10 or newer. The wheel is about 590 kB.
The setup has several distinct layers. Collision and stability checks need roughly 1 GB of BrickNet collider meshes. Preview rendering also needs Node, Mesa, a compatible C++ runtime, viewer dependencies and the specified LDraw snapshot.
The documented scoring path renders eight views and uses Gemma 4 31B, with approximately 62 GB of GPU memory. That renderer requires Python 3.13, a stricter requirement than the package itself.
For submissions, the documented form accepts a ZIP up to 10 MB, with one folder per task containing build.py, model.mpd or both. Missing tasks count as invalid. ByteForward has not run the environment or submitted results.
A useful first experiment
For developers, a sensible first step is one saved build with a recorded tool configuration. Inspect the assembly, keep the validation output, then decide whether a full scoring run justifies the hardware budget. A reproducible artifact gives another team something concrete to challenge.
A physical LEGO pyramid model shown for illustration. Lego MOC Great Pyramid of Giza by Hasan Kabalak, via Wikimedia Commons, licensed under CC BY-SA 2.0.







