Loopie research tests whether repeated layers improve AI efficiency

Looped Transformers keep losing the boring argument: if you spend N× the pre-training compute looping, why not just buy a bigger non-looped model? A new paper, Loop the Loopies! (arXiv:2607.16051), says that fight is winnable — if you loop layers, not whole stacks, and you match training cost the way operators actually measure it.
The claim is sharp. Independent replication is still open.
What layer-loop recurrence means here
Loopie is a Mixture-of-Experts series built on a Qwen3-MoE-like backbone. The twist is scheduling. Prior looped LLMs such as Ouro and Huginn mostly use model-loop: run the full stack, then repeat. Loopie uses layer-loop: each Attention/MoE block is applied recurrently before handing off to the next layer.
For a three-layer, two-step toy: model-loop is L1→L2→L3→L1→L2→L3. Layer-loop is L1→L1→L2→L2→L3→L3. The authors argue that local iteration is friendlier to MoE scaling, pipeline locality, and coherent parameter sharing across adjacent effective depths. Flagship configs use only two loop steps — the smallest nontrivial recurrence they find worth the compute multiplier in large-scale pre-training.
Compute-matched comparisons claimed
The paper’s center of gravity is the Loopie Recipe. Start from a non-recurrent MoE reference (a Qwen3-like 30B-A3B), halve stored depth, set R=2, then spend the activation-memory headroom on a larger microbatch and extra width/depth so measured optimizer-step time matches the baseline in Megatron-LM.
Important hedge from the authors themselves: they match realized wall-clock training cost, not exact theoretical FLOPs. Loopie may do more nominal work per token; efficiency from the schedule is what closes the gap. That is a stronger claim than “same parameter count,” and a harder one for outsiders to audit without the same stack, checkpointing policy, and hardware grid search.
Loopie-20B and 6B headline results
Two models headline the series: Loopie-20B-A2B (20B total / 2B active) and Loopie-6B-A0.6B (6B / 0.6B active). Against a compute-matched vanilla 30B-A3B trained on the same budget, the authors say Loopie-20B lags early, then overtakes after roughly 600B tokens and holds the lead. A smaller scaling ladder across four rungs also favors Loopie under matched step time, which is the paper’s answer to “does this only work once?”
Post-training is aggressive: a “supervised pre-training” stage on about 2T tokens of instruction and reasoning data (SFT-style loss at pre-training scale), then math-then-code RL. Author-reported Loopie-20B Thinking numbers include AIME 2024 around 92 (avg@8) and competitive scores versus similarly sized MoE reasoners — on a claimed ~3.5T pre-training token budget, far below some Nemotron comparisons in their table. Treat every leaderboard cell as author-reported until someone else runs the evals.
What independent replications should stress-test
Analysis: the paper’s real stake is methodological. If layer-loop plus hardware-aware matching beats wider vanilla MoEs under equal step time, recurrence stops being a parameter-efficiency parlor trick and becomes a third scaling axis next to width and data. What to stress-test first: (1) whether wall-clock matching reproduces on other clusters and frameworks; (2) whether gains survive held-out eval harnesses, not only the authors’ eight-benchmark average; (3) whether R=2 remains the sweet spot once someone else pays for the ablations. Until then, Loopie is a serious recipe paper — not a settled rewrite of scaling laws.



