Universal Transformers

Looped Transformers as Programmable Computers

DeepLoop:Depth Scaling for Looped Transformers

The Recurrent Transformer:Greater Effective Depth and Efficient Decoding

SMELT:Scaling Laws for Compute-Matched MoE Looped Transformers

Thinking Deeper, Not Longer:Depth-Recurrent Transformers for Compositional Generalization

Loop, Think, & Generalize:Implicit Reasoning in Recurrent-Depth Transformers

Relaxed Recursive Transformers:Effective Parameter Sharing with Layer-wise LoRA