Training-free language model (without training) which, according to him, generates Shakespeare 250 times faster than nanoGPT.
What's inside, in theory?
This is the "unbounded n-gram" approach: you can use very large n without running into exponential memory like with classic n-grams.
Instead of a huge n-gram table, a suffix array is used: it simulates n-gram lookup of any size with logarithmic access time.
Previously, such things were hardly used for generation, because sampling broke down: infinite perplexity and frequent verbatim copying.
The author claims that he solved this with a new method called Selective Back-off Interpolation Sampling: it mixes probability distributions from several levels of n-grams to maintain a balance between quality and novelty.
If you like non-standard LM approaches and "fast, simple, without training" - it's worth checking out the analysis and code.
Link to detailed write-up
••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore