Built a dpdk-inspired lock-free queue and found some surprising perf behavior
Built a lock-free ring queue in C++ inspired by DPDK `rte_ring`.
Expected to spend most of the time investigating cache coherency and inter-core synchronization.
Didn’t expect compiler codegen and tiny abstractions to matter this much.
Some interesting findings:
* avoiding grouped producer/consumer structs reduced \~300k instructions, \~200k branches and \~100k cache misses
* adding a single variable(for batching) and using template abstraction measurably changed generated assembly
* two specialized classes generated better code than one generic template-heavy abstraction
* GCC was much more sensitive to benchmark-harness overhead while Clang optimized through it better (32-byte payload throughput: \~8M ops/sec vs \~52M ops/sec before fixing the harness)
Benchmarked against:
* Rigtorp
* Drogalis
* Boost SPSC
* MoodyCamel
Current focus:
* SPSC latency
* batching/publication tuning
* cache coherency behavior
Note: Have not much explored the mpmc yet neither exposed its benchmarking strategy.
I felt like although my code was relatively easier to comprehend, although my throughput is better than them, my latency is to be improved currently 10-20 more cycles than rigtorp and drogalis.
Would appreciate feedback from people experienced with low-latency systems / lock-free structures.
[https://github.com/shuraih775/lfqueue](https://github.com/shuraih775/lfqueue)
https://redd.it/1t96fqb
@r_cpp
Post #25166
10