This is a new version of "LLM microscope", or more precisely a set of tools (interpretability tools), designed for interpreting the behavior of LLMs. Specifically, from the Gemma 3 family.
🟢 Why it matters?
Scope works on the basis of SAEs - sparse autoencoders. These are models that untangle the activations of LLMs and extract interpretable concepts from them.These are called "features": they can be things from the real world (bridges, cows) or abstractions (lying, responsiveness).
In essence, by analyzing these features, we can see what the model was actually thinking when generating a particular output. For example, it generates seemingly harmless code, but "thinks" about the concept of a "cyberattack". And this tells us something.
SAEs, by the way, were proposed for use by Anthropic in 2023 (here's our analysis of their article that made the approach popular). But it was Google that brought autoencoders to the production level. Now, this is actually the first and only open tool for such detailed interpretation of LLMs.
The first version of Scope was released in 2024. Then it only worked for small models and simple queries. Now, the approach has been scaled up even for a 27B model.
Plus, the tool has now become more versatile. If the original Scope only existed for a limited number of layers, now it's possible to analyze complex dialogue mechanisms in their entirety.
According to the article, this was mainly achieved by adding Skip-transcoders and Cross-layer transcoders to the model. These are modules that help to see the connections between distant layers and facilitate the analysis of distributed computations. And, by the way, SAEs were trained using the matryoshka method, like Gemma 3n (we wrote about this method here).
If you want to try and delve into the thoughts of models: Huggingface, Colab notebook, technical report, documentation
•••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
