Paper: DeepLoop: Depth Scaling for Looped Transformers

Topic: Looped Transformer / Recurrent Depth / Residual Scaling / Parameter Sharing / Reasoning


1. 一句话总结

DeepLoop 研究 Looped Transformer 中的深度扩展问题:当少量 Transformer Blocks 被反复复用时,由于参数共享导致 Gradient Aggregation + Repeated Read-back,传统 DeepNorm 的 $N^{1/4}$ scaling 可能失效,因此提出更强的 $N^{1/2}$ residual scaling。

核心结论:

$\boxed{p=\frac14 \rightarrow p=\frac12}$


2. 为什么研究 Looped Transformer?

传统 Transformer:

$Block1→Block2→⋯→BlockNBlock_1 \rightarrow Block_2 \rightarrow \cdots \rightarrow Block_N$

如果想增加模型深度:$N\uparrow$

通常需要:$Parameters\uparrow$


Looped Transformer

只使用少量 Blocks:

$K \text{ physical blocks}$

然后循环执行:

$R \text{ times}$

因此有效深度:

$\boxed{N=K\times R}$