Stanford introduced Matryoshka Attribution, a new attribution method which uses gradient descent to find which parts of a neural network are responsible for a behaviour.
MAttr is #1 on the Mechanistic Interpretability Benchmark by a wide margin (2.9× the runner up).
GitHub.
Post #4498
295