Post #1811
297
Forwarded from Alexander S
GitHub GitHub - Tencent/WeDLM: WeDLM: The fastest diffusion language model with standard causal attention and native KV cache compatibility… WeDLM: The fastest diffusion language model with standard causal attention and native KV cache compatibility, delivering real speedups over vLLM-optimized baselines. - Tencent/WeDLM