On a 24 GB Mac Mini, MLX Beats llama.cpp by 1.9×
A controlled 24 GB M4 Mac mini test ran the same Qwen DFlash 2 speculative-decoding setup through MLX and llama.cpp. MLX delivered 1.8–1.9× speedups, while the tested llama.cpp build was roughly 50% slower and exhausted memory.
🔗 tools.cooconsbit.com
#LocalLLM #MLX #Benchmark
Post #309
449

- 👍 3
- 🔥 2
- ❤ 1