An excellent tool to estimate how much VRAM your LLM really needs.
You change the hardware configuration, quantization, etc. and immediately see:
generation speed (tokens/sec)
exact memory allocation
system throughput and more
Try Here
••••••••••••••••••••••••••••••••••••••
🤖 @DataXplore
Post #1991
173