TGViewer
C++ - Reddit C++ - Reddit @r_cpp · 229 subscribers
Post #25669 16
AVX-512 Intrinsics Optimization: Achieving ~7% speedup over libpopcnt on Sapphire Rapids

Benchmarking random 64B cache-line jumps on Intel Xeon Sapphire Rapids showed the custom kernel achieving 7.09 ns/line, compared to libpopcnt (7.58 ns) and std::popcount (46.10 ns).

While contiguous memory performance ties libpopcnt at 0.47 ns/word, the performance gain on this gather hot-path comes from maintaining accumulator states inside ZMM registers across steps and applying 2-accumulator unrolling to break dependency chains.

Target is v40. Looking for insights on whether software prefetching (_mm_prefetch) provides measurable benefits for this gather pattern or if latency is memory-bound.

https://redd.it/1v86xcs
@r_cpp
Reddit From the cpp community on Reddit Explore this post and more from the cpp community
More from @r_cpp
  1. Sep 26, 2026Token Sequence Injection & Modern Macros: The Most Game-Changing Compile-Time Feature in C…
  2. Sep 25, 2026myStringStream.str("") Considered Harmful Under C++20 I was recently looking at some (rath…
  3. Sep 20, 2026A clever branch free optimization I'm the developer of memlz which is an extremely fast co…
  4. Sep 15, 2026Inside Boost.PolyCollection https://bannalia.blogspot.com/2026/09/inside-boostpolycollecti…
  5. Sep 11, 2026C++26: Standard Library Hardening Experiments https://www.cppstories.com/2026/hardening-ex…
  6. Sep 11, 2026MSVC C++23: constexpr cmath with LLVM Libc https://devblogs.microsoft.com/cppblog/msvc-c23…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →