TGViewer
All about AI, Web 3.0, BCI All about AI, Web 3.0, BCI @alwebbci · 3.89K subscribers
Post #4395 624
Cloudflare showed how to run massive open models like Moonshot’s Kimi and Z.ai’s GLM faster, cheaper, and safer without sacrificing quality.

These are powerful long-context MoE models, but they’re extremely memory-hungry.

Cloudflare shared 3 practical techniques that let them fit more of them onto the same GPUs in Workers AI.
Cloudflare Blog Smaller, faster, safer: running Kimi and GLM at scale Serving frontier models like Kimi and GLM means fighting for GPU memory. Here's how we quantize KV caches, compress model weights, and add integrity checks to serve them faster, cheaper, and safely.
  • ❤ 7
  • 🔥 4
  • 🙏 3
More from @alwebbci
  1. Sep 25, 2026New From DeepSeek: a sandbox platform running 3 million AI agent environments per day. Dee…
  2. Sep 25, 2026Super interesting paper from Google and colleagues. It studies where it's possible to dist…
  3. Sep 24, 2026DeepMind Institute presented 3 new essays: 1. How can we control misbehaviour in agent swa…
  4. Sep 24, 2026Stanford introduced Matryoshka Attribution, a new attribution method which uses gradient d…
  5. Sep 23, 2026Claude has discovered a previously unknown enzyme system hidden in the DNA of bacteriophag…
  6. Sep 23, 2026Xiaomi released MiMo-V2.6 - Pro & Flash 2 omnimodal models, advancing through scaled reinf…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →