⚡️ Transformers can now load llama.cpp GGUF quants directly via from_pretrained
Hugging Face added native GGUF support for Qwen3.5 on Apple Silicon using ggml's Metal kernels, matching llama.cpp throughput while staying inside the regular Transformers API.
♻️ Subscribe for free now!
Post #2780
14