TGViewer
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security Data eXplore : Data Science, ML, Big Data, LLMs and AI Security @dataxplore · 583 subscribers
Post #1985 173
How to let LLMs sort context by importance themselves?

Conventional language models read text as one long strip.

What's closer to the beginning of attention - that's "more important".
What's further - the model sees it worse. And here a problem arises: if an important fact is hidden somewhere far away in the noise, the model may simply not use it.

It spends attention on everything, instead of focusing on the main thing.

🟢 What Sakana AI has figured out?
Sakana AI proposed a solution - RePo (Context Re-Positioning).

The idea is very clear: the model gets a module that allows to dynamically "re-position" the context.

Like a person:
you read a long document, realize that the important part was 20 pages back - and mentally re-read it, ignoring the rest.

What RePo does
- pulls important pieces of information closer
- pushes away noise and excess text
- helps the model's attention focus on what's needed

In the model, there's a trainable module that re-assigns the positions of tokens according to meaning, not order

✅ important = what helps reduce the model's error and solve the task correctly
❌ secondary = what doesn't help (noise), so it's "pushed away" in terms of positions

As a result, the model with such a memory starts to work better where LLMs usually suffer:
- when the context is long
- when there's a lot of noise
- when important details are scattered far from each other
- when the data is structured (tables, lists, rules)

The authors show that RePo gives a noticeable increase in robustness, without worsening the overall quality.

▶️ Robustness to noise (Noisy Context)
Average result on 8 noisy benchmarks:

- Regular RoPE: 21.07
- RePo: 28.31

🟡 Increase: +7.24 points (strongly)

The authors separately note a key figure:
on noisy-eval (4K context) RePo is better than RoPE by +11.04 points.

🔥 Examples of increase on specific tasks
(everywhere RePo > RoPE)

- TriviaQA: 61.47 → 73.02 (+11.55)
- GovReport: 6.23 → 16.80 (+10.57)
- 2WikiMultihopQA: 23.32 → 30.86 (+7.54)
- MuSiQue: 7.24 → 13.45 (+6.21)

This is a step towards models that don't just "read what they're given", but are able to organize their working memory themselves.


Details, Article • #RePo #SakanaAI #LLM #AI #AIAgents #Context #LongContext #Attention

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
More from @dataxplore
  1. Oct 1, 2026Monitoring and debugging "silent" token drift in LLM pipelines: frequency analysis of embe…
  2. Sep 30, 2026Online detection of feature collisions in TDA transformation When using Topological Data A…
  3. Sep 18, 2026Am going to announce something big (for me, it's really big) on October 11, 2026.
  4. Sep 14, 2026Post #2188
  5. Aug 31, 2026I joined a Russian community on Telegram. They share some Russian startup and technology u…
  6. Aug 22, 2026Post #2185
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →