TGViewer
Data science/ML/AI Data science/ML/AI @datascience_bds · 14K subscribers
Post #1328 1.22K
🚨 UC Berkeley just open-sourced FreeToken.

It claims 2–4× faster local LLM inference than Ollama, and the wild part is the models it can run:
• Qwen3.6-35B on 8GB VRAM → 39.3 tok/s
• DeepSeek-V4-Flash 284B on 32GB VRAM → 22 tok/s
• GLM-5.2 753B on 96GB VRAM → 14.9 tok/s

How? These are Mixture-of-Experts models. A 35B model doesn't actually use all 35B parameters for every token.

FreeToken keeps the experts in system RAM and intelligently decides whether a missing expert should be sent to the GPU or computed on the CPU.

The good part is the best strategy depends on your exact machine. A 5090 desktop and an 8GB laptop may want completely opposite approaches.

It also checkpoints agent context, so coding agents don't repeatedly prefill thousands of unchanged tokens.

Open weights don't mean much if nobody can afford the hardware to run them. FreeToken is attacking that gap.

📄 Paper: https://arxiv.org/pdf/2608.16157
💻 Repo: https://github.com/FlashML-org/FreeToken
  • ❤ 3
  • 🔥 3
More from @datascience_bds
  1. Oct 9, 20268 RAG architectures for AI Engineers, visually explained:
  2. Oct 8, 2026document post
  3. Oct 7, 2026🧮 NumPy: Why axis=0 and axis=1 Feel Backwards You've probably seen: np.mean(X, axis=0) an…
  4. Oct 6, 2026document post
  5. Oct 5, 2026📊 Pandas Cheatsheet Every Data Analyst Should Save Pandas is one of the most important to…
  6. Oct 4, 2026SQLBolt: Interactive SQL You can learn SQL by writing real queries directly in the browser…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →