ML Nested Learning Approach
Recently Google published quite interesting research "Introducing Nested Learning: A new ML paradigm for continual learning" .
So let's check what it is about and why it can be interesting for ML society.
Modern LLMs tend to loose efficiency in old tasks while learning new tasks. This effect is called catastrophic forgetting (CF). Researchers and data scientists spend significant amount of time to invent some architectural tweaks or better optimizations to deal with it.
Authors define the root cause of this issue as a separation of model's architecture and optimization algorithms for two different things.
The Nested Learning treats ML model as a system of interconnected, multi-level learning problems that are optimized simultaneously. And in that approach the model's architecture and the rules used to train it are fundamentally the same concepts.
The overall idea is heavily based on associative memory concept: the ability to map and recall one thing based on another (like recalling a name when you see someone's face):
🔸 The training process is modeled as an associative memory. The model learns to map a given data point to the value of its local error, which serves as a measure of how "surprising" or unexpected that data point was.
🔸 Key architectural components (like transformers) are also formalized as simple associative memory modules that learn the mapping between tokens in a sequence.
Proof-of-concept tests shows superior performance in language modeling and demonstrates better long-context memory management than in existing models (you can check the measurements in the article or in the full text of the research there ).
ML keeps evolving. What’s interesting is that the best architecture ideas come from systems theory and attempts to copy human brain behavior.
#ai #architecture
Post #239
253