🟢 What's new?
☞ FlexAttention has become even more widely available. Added support for Intel GPU, plus flash-decoding optimizations on x86 CPU (speeding up transformer inference without special CUDA cores).
☞ In torch.compile, you can now switch behavior on graph breaks, meaning you can choose to either throw an error or continue.
☞ Official binaries now come immediately with versions for ROCm (AMD), XPU (Intel GPU), and CUDA 13. This greatly simplifies GPU startup out of the box outside NVIDIA, especially on ROCm.
☞ Updated stable libtorch ABI. There should now be fewer surprises when building and updating third-party C++/CUDA extensions.
☞ Rolled out symmetric memory. a new memory model designed to make writing multi-GPU kernels and code easier.
By the way, the release was built from 3216 commits by 452 contributors. Open source is power 💪
🤖 Data Science, ML & Big Data with @DataXplore
