TGViewer
Data science/ML/AI Data science/ML/AI @datascience_bds · 14K subscribers
Post #1338 1.2K
5 LLM quantization techniques, clearly explained:

1. RTN: ignores them. Rounds every weight to the nearest grid level with no calibration data. Cheapest option, weakest at low bit widths.

2. GPTQ: repairs after rounding. Quantizes a layer column by column and adjusts the remaining weights to absorb the error before moving on.

3. AWQ: protects before rounding. Finds the ~1% of weight channels that matter most and scales them up so they survive quantization. Everything still ends up in plain INT4.

4. LLM. int8(): isolates at inference. Outlier dimensions run in FP16, the other 99.9% run in INT8, and the results are merged.

5. QAT: solves it during training. The model is fine-tuned with rounding baked into every forward pass, so it adapts to the damage before quantization is actually applied.

All five produce the same artifact, a model at a fraction of its trained precision. They differ only in where the outlier problem gets addressed.

The visual above nicely summarises these techniques.

#LLM
  • ❤ 8
More from @datascience_bds
  1. Oct 8, 2026document post
  2. Oct 7, 2026🧮 NumPy: Why axis=0 and axis=1 Feel Backwards You've probably seen: np.mean(X, axis=0) an…
  3. Oct 6, 2026document post
  4. Oct 5, 2026📊 Pandas Cheatsheet Every Data Analyst Should Save Pandas is one of the most important to…
  5. Oct 4, 2026SQLBolt: Interactive SQL You can learn SQL by writing real queries directly in the browser…
  6. Oct 3, 2026Tools vs MCP vs Skills: 3 Layers That Power Production AI Agents
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →