Achieving 56.5 ns cross-language IPC latency: Defeating false sharing and bypassing the kernel.
Hi,
I recently open-sourced Tachyon, a low-latency shared-memory IPC library I’ve been working on. The goal was to reach RAM-speed communication between processes (C++, Rust, Python, etc.) without any serialization overhead or kernel involvement on the hot path.
I managed to hit a p50 round-trip time of 56.5 ns (for 32-byte payloads) and a throughput of \~13M RTT/sec on an i7-12650H, which is about 150x faster than ZeroMQ inproc.
Here are a few architectural choices I made to achieve this, which I thought might interest this sub:
Strict SPSC & No CAS: I went with a strict Single-Producer Single-Consumer topology. There are no compare-and-swap loops on the hot path. acquire_tx and acquire_rx are just a load, a mask, and a branch using memory_order_acquire/release.
Hardware Sympathy: Every control structure (message headers, atomic indices) is padded to 64-byte or 128-byte boundaries. False sharing between the producer and consumer cache lines is structurally impossible.
Hybrid Wait Strategy: The consumer spins for a bounded threshold (cpu_relax()), then sleeps via SYS_futex (Linux) or __ulock_wait (macOS).
Zero-Copy: The hot path is entirely in the memfd shared memory segment after an initial Unix Domain Socket handshake.
The core is C++23 (compiled with GCC 14+/Clang 17+), and it currently has bindings for 6 other languages.
Repository: https://github.com/riyaneel/Tachyon
I’d love to get some feedback from the C++ community on the architecture, especially regarding the memory model implementation and the hybrid futex spin-wait strategy.
Thanks!
https://redd.it/1sn8a1y
@r_cpp
Post #24990
22