Looped Transformers as Programmable Computers
DeepLoop:Depth Scaling for Looped Transformers
The Recurrent Transformer:Greater Effective Depth and Efficient Decoding
SMELT:Scaling Laws for Compute-Matched MoE Looped Transformers
Thinking Deeper, Not Longer:Depth-Recurrent Transformers for Compositional Generalization
Loop, Think, & Generalize:Implicit Reasoning in Recurrent-Depth Transformers
Relaxed Recursive Transformers:Effective Parameter Sharing with Layer-wise LoRA