Instead replacing old information, it integrates new learning inside what’s already known, like a layer within a layer. The model retains prior skills, adapts to new tasks, and distinguishes the context it’s operating in.
🟢 How Nested Learning works?
1️⃣ The authors formalize the model as a set of optimization problems: each has its own stream of information it learns from and its own update frequency. For example, components with a high update frequency are responsible for adapting to the current context, those with a low frequency handle some basic knowledge, and so on.
2️⃣ But the model won’t just magically know what and when to update. Therefore, the authors propose making the optimizer itself trainable. That is, the algorithm responsible for updating weights stops being just a formula and turns into a neural network itself. This is called Deep Optimizers.
3️⃣ Formally, the optimizer is considered associative memory that learns to associate gradients with the correct weight changes. In this sense, conventional SGD or Adam are the simplest special cases.
Google extended their old TITAN architecture (which handled short and long memory) into HOPE, supporting unlimited levels of in-context learning.
Resulting Lower(↓) perplexity, higher(↑) accuracy in common-sense reasoning, and stronger long-context memory.
🤖 Data Science, ML & Big Data with @DataXplore