SymTorch: a library that translates deep learning models into human-readable equations.
Attached a short video showing how SymTorch works.
I have a background in physics, and when I think about understanding a system, I think about EQUATIONS.
Equations are great: they precisely show how inputs map to outputs, which variables are important, and how the system behaves in OOD situations. Let's apply this to model interpretability.
The main principle of SymTorch is simple. For any arbitrary component of the neural network in your large architecture, we record the input and output activations on some data examples and use symbolic regression with PySR to find an equation that approximately describes the behavior of this component.
All the engineering overhead (GPU/CPU data transfer, native PyTorch model serialization, I/O caching, etc.) is already handled by SymTorch.
We've demonstrated SymTorch on a wide range of cases and architectures: from solving PDEs with PINN to understanding LLM outputs.
Paper, Website, GitHub
••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore