Benchmarking random 64B cache-line jumps on Intel Xeon Sapphire Rapids showed the custom kernel achieving 7.09 ns/line, compared to
libpopcnt (7.58 ns) and std::popcount (46.10 ns).While contiguous memory performance ties
libpopcnt at 0.47 ns/word, the performance gain on this gather hot-path comes from maintaining accumulator states inside ZMM registers across steps and applying 2-accumulator unrolling to break dependency chains.Target is v40. Looking for insights on whether software prefetching (
_mm_prefetch) provides measurable benefits for this gather pattern or if latency is memory-bound.https://redd.it/1v86xcs
@r_cpp