TGViewer
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security Data eXplore : Data Science, ML, Big Data, LLMs and AI Security @dataxplore · 583 subscribers
Post #2001 154
Came across an interesting project: tiny-infini-gram.

Training-free language model (without training) which, according to him, generates Shakespeare 250 times faster than nanoGPT.

What's inside, in theory?

This is the "unbounded n-gram" approach: you can use very large n without running into exponential memory like with classic n-grams.

Instead of a huge n-gram table, a suffix array is used: it simulates n-gram lookup of any size with logarithmic access time.

Previously, such things were hardly used for generation, because sampling broke down: infinite perplexity and frequent verbatim copying.

The author claims that he solved this with a new method called Selective Back-off Interpolation Sampling: it mixes probability distributions from several levels of n-grams to maintain a balance between quality and novelty.

If you like non-standard LM approaches and "fast, simple, without training" - it's worth checking out the analysis and code.


Link to detailed write-up

••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
More from @dataxplore
  1. Oct 1, 2026Monitoring and debugging "silent" token drift in LLM pipelines: frequency analysis of embe…
  2. Sep 30, 2026Online detection of feature collisions in TDA transformation When using Topological Data A…
  3. Sep 18, 2026Am going to announce something big (for me, it's really big) on October 11, 2026.
  4. Sep 14, 2026Post #2188
  5. Aug 31, 2026I joined a Russian community on Telegram. They share some Russian startup and technology u…
  6. Aug 22, 2026Post #2185
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →