Conventional language models read text as one long strip.
What's closer to the beginning of attention - that's "more important".
What's further - the model sees it worse. And here a problem arises: if an important fact is hidden somewhere far away in the noise, the model may simply not use it.
It spends attention on everything, instead of focusing on the main thing.
🟢 What Sakana AI has figured out?
Sakana AI proposed a solution - RePo (Context Re-Positioning).
The idea is very clear: the model gets a module that allows to dynamically "re-position" the context.
Like a person:
you read a long document, realize that the important part was 20 pages back - and mentally re-read it, ignoring the rest.
What RePo does
- pulls important pieces of information closer
- pushes away noise and excess text
- helps the model's attention focus on what's needed
In the model, there's a trainable module that re-assigns the positions of tokens according to meaning, not order
✅ important = what helps reduce the model's error and solve the task correctly
❌ secondary = what doesn't help (noise), so it's "pushed away" in terms of positions
As a result, the model with such a memory starts to work better where LLMs usually suffer:
- when the context is long
- when there's a lot of noise
- when important details are scattered far from each other
- when the data is structured (tables, lists, rules)
The authors show that RePo gives a noticeable increase in robustness, without worsening the overall quality.
▶️ Robustness to noise (Noisy Context)
Average result on 8 noisy benchmarks:
- Regular RoPE: 21.07
- RePo: 28.31
🟡 Increase: +7.24 points (strongly)
The authors separately note a key figure:
on noisy-eval (4K context) RePo is better than RoPE by +11.04 points.
🔥 Examples of increase on specific tasks
(everywhere RePo > RoPE)
- TriviaQA: 61.47 → 73.02 (+11.55)
- GovReport: 6.23 → 16.80 (+10.57)
- 2WikiMultihopQA: 23.32 → 30.86 (+7.54)
- MuSiQue: 7.24 → 13.45 (+6.21)
This is a step towards models that don't just "read what they're given", but are able to organize their working memory themselves.
Details, Article • #RePo #SakanaAI #LLM #AI #AIAgents #Context #LongContext #Attention
••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore