A Princeton Researcher Proposes Recurrent Looped Transformer (RLT): a decoder that carries its full state across every prompt and response token, with 96 logical blocks per token and no reset at the boundary.
Here's how it works. 👇
1. One state, never reset
A causal encoder builds KV memory in parallel. A recurrent decoder merges each token's encoder feature with the previous final hidden state, then runs sliding-window attention, cross-attention to encoder memory, and an FFN. The state H_t = (s_t, C^D_t) crosses the serving split untouched.
→ Proposition 3.1: moving the prompt-response split does not change the conditional distribution
2. Fixed work, growing depth
Reference config ties 48 encoder and 48 decoder layers with shared attention and FFN weights.
→ 96 logical blocks per token, fixed
→ 48t decoder blocks on the state path after t tokens
3. RL replay under current parameters
The sampler records behavior log-probs. The trainer rebuilds encoder memory, recurrent output, and every SWA cache from H_0 before scoring each action.
→ Old rollout states are never reused; weight updates invalidate all caches
The key takeaway: a design spec that makes pretraining, SFT, sampling, and RL replay share one transition.
Full analysis: https://www.marktechpost.com/2026/09/13/a-princeton-researcher-proposes-recurrent-looped-transformer-rlt/
Paper: https://github.com/yifanzhang-pro/recurrent-looped-tranformer/blob/master/Recurrent_Looped_Transformer.pdf
Git Repo: https://github.com/yifanzhang-pro/recurrent-looped-tranformer
Amazing explanation in author's research blog: https://yifanzhang-pro.github.io/recurrent-looped-tranformer/
Post #1548
762

- 👍 1