Loopie research tests whether repeated layers improve AI efficiency

Looped Transformers keep losing the boring argument: if you spend Nร the pre-training compute looping, why not just buy a bigger non-looped model? A new paper, Loop the Loopies! (arXiv:2607.16051), says that fight is winnable โ if you loop layers, not whole stacks, and you match training cost the way operators actually measure it.
The claim is sharp. Independent replication is still open.
What layer-loop recurrence means here
Loopie is a Mixture-of-Experts series built on a Qwen3-MoE-like backbone. The twist is scheduling. Prior looped LLMs such as Ouro and Huginn mostly use model-loop: run the full stack, then repeat. Loopie uses layer-loop: each Attention/MoE block is applied recurrently before handing off to the next layer.
For a three-layer, two-step toy: model-loop is L1โL2โL3โL1โL2โL3. Layer-loop is L1โL1โL2โL2โL3โL3. The authors argue that local iteration is friendlier to MoE scaling, pipeline locality, and coherent parameter sharing across adjacent effective depths. Flagship configs use only two loop steps โ the smallest nontrivial recurrence they find worth the compute multiplier in large-scale pre-training.
Compute-matched comparisons claimed
The paper’s center of gravity is the Loopie Recipe. Start from a non-recurrent MoE reference (a Qwen3-like 30B-A3B), halve stored depth, set R=2, then spend the activation-memory headroom on a larger microbatch and extra width/depth so measured optimizer-step time matches the baseline in Megatron-LM.
Important hedge from the authors themselves: they match realized wall-clock training cost, not exact theoretical FLOPs. Loopie may do more nominal work per token; efficiency from the schedule is what closes the gap. That is a stronger claim than “same parameter count,” and a harder one for outsiders to audit without the same stack, checkpointing policy, and hardware grid search.
Loopie-20B and 6B headline results
Two models headline the series: Loopie-20B-A2B (20B total / 2B active) and Loopie-6B-A0.6B (6B / 0.6B active). Against a compute-matched vanilla 30B-A3B trained on the same budget, the authors say Loopie-20B lags early, then overtakes after roughly 600B tokens and holds the lead. A smaller scaling ladder across four rungs also favors Loopie under matched step time, which is the paper’s answer to “does this only work once?”
Post-training is aggressive: a “supervised pre-training” stage on about 2T tokens of instruction and reasoning data (SFT-style loss at pre-training scale), then math-then-code RL. Author-reported Loopie-20B Thinking numbers include AIME 2024 around 92 (avg@8) and competitive scores versus similarly sized MoE reasoners โ on a claimed ~3.5T pre-training token budget, far below some Nemotron comparisons in their table. Treat every leaderboard cell as author-reported until someone else runs the evals.
What independent replications should stress-test
Analysis: the paper’s real stake is methodological. If layer-loop plus hardware-aware matching beats wider vanilla MoEs under equal step time, recurrence stops being a parameter-efficiency parlor trick and becomes a third scaling axis next to width and data. What to stress-test first: (1) whether wall-clock matching reproduces on other clusters and frameworks; (2) whether gains survive held-out eval harnesses, not only the authors’ eight-benchmark average; (3) whether R=2 remains the sweet spot once someone else pays for the ablations. Until then, Loopie is a serious recipe paper โ not a settled rewrite of scaling laws.



