TGViewer
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security Data eXplore : Data Science, ML, Big Data, LLMs and AI Security @dataxplore · 583 subscribers
Post #1890 210
AI has learned to write CUDA kernels more efficiently than NVIDIA engineers.

DeepReinforce research group has developed a fully automatic GPU code generation system for matrix multiplication called CUDA-L2.

🟢 How does it achieve a 10–30% faster over NVIDIA's highly-optimized cuBLAS & cuBLASLt libraries?

Such libraries are manually created by people who use ready-made kernel templates. Autotuners only tweak parameters, such as tile sizes.

But DeepReinforce believes that even critically important and deeply optimized tasks like HGEMM can be improved using an LLM working in conjunction with RL.

In the CUDA-L2 system, the language model literally writes CUDA source code from scratch for each matrix size. It doesn't just change parameters; it can change the code structure, loops, tiling strategy, padding, and even swizzle patterns. Moreover, it chooses the programming style itself—whether raw CUDA, CuTe, CUTLASS, or inline PTX.

The process works like this: an RL loop runs the generated kernels on real hardware, measures speed and correctness, and then updates the LLM. Over time, the model derives its own performance rules instead of relying on human knowledge.

The generator used was the DeepSeek 671B model. It was further trained on a mixture of CUDA kernel arrays and quality code from PyTorch, ATen, CUTLASS libraries, and examples from NVIDIA.

WHAT THIS MEANS?

For pretraining and fine-tuning, most GPU time is spent on HGEMM matrix multiplication operations. If these kernels are sped up by the 10–30% promised by CUDA-L2, the entire training process becomes noticeably cheaper and faster.

Since CUDA-L2 handles about 1000 real matrix sizes, not just a few manually tuned ones, the acceleration works across a wide range of architectures. This means that with the same GPU budget, you can fit more training tokens, more SFT or RLHF runs, etc.


TESTS

HGEMM kernels created by CUDA-L2 are consistently faster than standard libraries.

In the so-called "offline scenario," CUDA-L2 runs about 17–22% faster than torch.matmul, cuBLAS, and cuBLASLt. It even beats cuBLASLt AutoTuning by 11%, which itself uses kernel search.

In the "server" scenario, which simulates real inference with pauses between calls, the difference is even greater: a boost of 24–29% compared to torch.matmul and cuBLAS.


Project is not limited to simple research; in the GitHub repository, optimized 32-bit HGEMM A100 kernels for 1000 configurations.

#AI #ML #CUDA #DeepReinforce

•••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
More from @dataxplore
  1. Oct 1, 2026Monitoring and debugging "silent" token drift in LLM pipelines: frequency analysis of embe…
  2. Sep 30, 2026Online detection of feature collisions in TDA transformation When using Topological Data A…
  3. Sep 18, 2026Am going to announce something big (for me, it's really big) on October 11, 2026.
  4. Sep 14, 2026Post #2188
  5. Aug 31, 2026I joined a Russian community on Telegram. They share some Russian startup and technology u…
  6. Aug 22, 2026Post #2185
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →