Why ordinary transformers are almost incapable of multi-digit multiplication and how to fix it?
MIT + Harvard + Google DeepMind Trained 2 small Transformers to perform 4-digit × 4-digit multiplication.
1️⃣ The First used Implicit chain-of-thought (ICoT) method:
the model first sees all the intermediate calculation steps, and then these steps are gradually removed.
In other words, the model is forced to “think internally” rather than rely on visible hints.
Result: 100% accuracy on all examples.
2️⃣ The second used regular training:
input → answer, without intermediate steps.
Result: about 1% correct answers.
Why is that?
- Multi-digit multiplication requires long-range dependencies
- It is necessary to remember and carry over the “sum + carry” between different positions
- The model must store intermediate partial products and return to them later
- A working model forms a “running sum” and carry, like a human
- Inside attention, a structure resembling a small binary tree appears
- Digit representations form a special space (five-pointed prism + Fourier code)
Regular training captures the “edge” digits and gets stuck — it cannot connect the middle.
ICoT provides the correct inductive bias: it forces the model to build an internal algorithm rather than guess a pattern.
AI needs a computational process to do arithmetic and logic, not just more data.
🤖 Data Science, ML & Big Data with @DataXplore
