llms-foundation-models · 2026-09-27 · Tier 2

Sparse Layers are Critical to Scaling Looped Language Models

Sparse Layers are Critical to Scaling Looped Language Models

Source: arXiv 2605.09165 (Ryan Lee, Jonathan May, Edward Hu et al.), accepted at NeurIPS 2026. Surfaced via the X home feed (@_ryantlee), 2026-09-26.

TL;DR

Looped language models repeat a set of layers through depth, which saves memory and gives natural early-exit points, but dense looped models scale worse than ordinary transformers. This paper compares dense and MoE (mixture-of-experts, where each token runs through a few selected expert sub-networks) models with and without looping. Looped-MoE models scale better than the standard baseline, and dense looped ones do not. The reason is routing divergence: on each pass through the same shared layers, different experts fire, which recovers the expressivity that plain weight sharing loses, at no extra parameter cost. Loop boundaries also make better early-exit points than arbitrary layers, because each loop ends with the same layers that produce the final output.

flowchart LR
  X[Token] --> B1[Shared block<br/>pass 1]
  B1 -->|router picks<br/>experts A,C| B2[Shared block<br/>pass 2]
  B2 -->|router picks<br/>experts B,D| E{Exit at<br/>loop boundary?}
  E -->|confident| O[Output early]
  E -->|not yet| B3[Pass 3]
  classDef input fill:#dbeafe,stroke:#3b82f6,color:#1e3a8a
  classDef decision fill:#fef3c7,stroke:#f59e0b,color:#78350f
  classDef output fill:#d1fae5,stroke:#10b981,color:#065f46
  class X input
  class E decision
  class O output

Relation to prior wiki pages

  • Confirms SMELT. SMELT (09-02) already found that MoE looping beats unlooped baselines under matched budgets. This is a second, independent group reaching the same conclusion, and it names a mechanism (per-pass expert divergence). Two papers now say the loop needs sparsity to pay off.
  • Complements FlashLoop. FlashLoop (09-27) cuts the serving cost of extra loops; this paper says which looped models are worth serving.
  • Concept page: looped-transformers.

Gaps

  • Early-exit savings are argued from output convergence; wall-clock serving numbers under batching are not the focus.
  • Whether per-pass routing divergence survives expert-parallel serving (where each token's experts may sit on other GPUs) is untested.