Looped Models: Step Size, Placement, and Supervising the Thought
Source: HuggingFace Daily Papers, listed 2026-09-30 · TAPS, arXiv 2609.36653 · What Makes Recurrence Effective, arXiv 2609.36636 · REST, arXiv 2609.36159 Raw: TAPS · Recurrence · REST
TL;DR
Three papers on looped (recurrent-depth) models, which reuse the same layers several times to buy extra compute without extra parameters.
- TAPS (Trajectory Adaptive Progress-Fluctuation Scheduler): every loop applies its update at a fixed size of 1. That is too timid when updates keep making progress and too aggressive when they oscillate. TAPS splits the loss sensitivity to step size exactly into a "persistent progress" part and a "fluctuation" part, and adapts the step size online. Training-free, it raises final accuracy on structured reasoning tasks. Folded into training, it reaches baseline accuracy up to 1.56x faster in wall-clock.
- What Makes Recurrence Effective: a controlled study. Extra loops help reasoning beyond the training horizon but hurt knowledge recall, and harder problems do not reliably gain more. Where you put the unshared input and output layers matters, so effective depth alone does not predict behavior. Non-looped output layers make models robust to running fewer loops. The standard trick of re-injecting the initial state each loop is weak. Channel-wise history-state injection plus a timestep signal keeps knowledge intact under longer unrolling.
- REST (Representation-Supervised Thoughts): latent reasoning (loops over hidden states, or agents passing hidden states) is usually trained only on the final answer's cross-entropy. That lets thoughts collapse across different questions or carry junk. REST adds four losses (causality, minimality, separability, stability). Up to +7.5 points on seven benchmarks and 30% better convergence, with no inference cost.
The loop now has a speed dial, not just a count
TAPS sizes each loop's step from how steady the progress is.
flowchart LR
H["Hidden state<br/><small>loop t</small>"] --> U["Shared block<br/><small>proposed update</small>"]
U --> D{"TAPS<br/><small>progress vs noise</small>"}
D -->|steady| B["Bigger step<br/><small>fewer loops</small>"]
D -->|oscillating| S["Smaller step<br/><small>damp it</small>"]
B --> N["Next state<br/><small>loop t+1</small>"]
S --> N
classDef input fill:#d0ebff,stroke:#1971c2,color:#1b1b1b,stroke-width:2px
classDef core fill:#e5dbff,stroke:#6741d9,color:#1b1b1b,stroke-width:2px
classDef loop fill:#fff3bf,stroke:#f08c00,color:#1b1b1b,stroke-width:2px
classDef exit fill:#d3f9d8,stroke:#2f9e44,color:#1b1b1b,stroke-width:2px
classDef err fill:#ffe3e3,stroke:#e03131,color:#1b1b1b,stroke-width:2px
class H input
class U core
class D loop
class B exit
class S err
class N exit
linkStyle 2 stroke:#2f9e44,stroke-width:2px
Blue is the state, purple is the shared block, amber is the scheduler, green speeds up, red damps.
How it relates to prior wiki pages
- A fifth control knob on looped-transformers. The page's serving stack had cheaper loops (FlashLoop, 09-27), robust any-depth inference (LoopFormer, 09-28), batchable exits (Continuous Depth Batching, 09-29), and a learned per-token loop count plus speculative decoding (TaH2 and WaveFront, 09-30). TAPS adds update size. Loop count times step size is the real compute budget, and nobody has learned both jointly.
- A caution for TaH2's gains. TaH2 (09-30) showed its gain grows with depth on AIME. "What Makes Recurrence Effective" says extra depth trades knowledge for reasoning. That fits TaH2 being a math result, and predicts weaker gains on knowledge-heavy evaluations.
- REST links to the CLM and latent-communication threads: supervising what the hidden "thought" contains also makes it easier to decode, which matters for monitoring.
Links
- Concept: Looped transformers · Test-time compute allocation
- Digest: 2026-10-01