TGViewer
HN Best Comments HN Best Comments @hn_best_comments · 4.16K subscribers
Post #33888 479
Re: Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

I'm a little skeptical of going below 4-bit quants due to the potential for significant degradation in quality. I'm running 4-bit quants on an RTX Pro 6000 rented for approximately $1/hour and getting about 1.2 million tokens out and 40 million tokens in per hour with caching. The quality of 4-bit quant is good enough for difficult but well-scoped coding tasks. Here is the inference stack I am using: https://www.reddit.com/r/BlackwellPerformance/s/FrKwk3GoDK

a11r, 14 hours ago
More from @hn_best_comments
  1. Oct 5, 2026Re: The Tao of Backup > The novice became suspicious and said: "Master, is all this 'Tao o…
  2. Oct 5, 2026Re: What is going on with ceiling fans When an LED bulb dies, it's typically not the LED,…
  3. Oct 5, 2026Re: What is going on with ceiling fans Technology Connections nerd-snipes ceiling fans in…
  4. Oct 5, 2026Re: Pixel 11 doesn't yet meet the GrapheneOS security standards and may be skipped As othe…
  5. Oct 5, 2026Re: Denmark Data Breach Exposes 8.8M People's Personal Data I am at a point where I simply…
  6. Oct 5, 2026Re: Improper redaction reveals Google Data Center water and electricity usage As other com…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →