IPC benchmark: ~3M msg/s and ~7 GB/s on macOS. Missing bare-metal Linux numbers
I've been optimizing the IPC layer of areg-sdk, an open-source C++ framework I maintain. Here are the numbers and methodology.
**What was measured**:
Numbers are taken at `mtrouter` (the framework's message router), which sees both directions simultaneously and is more accurate point in the pipeline.
* **macOS (Apple M3 Pro, LPDDR5):** \~2.5–3.0M msg/s at \~0.5 KB payload, \~6.5–7.0 GB/s at \~3 MB payload
* **Windows 11 (Intel i7-13700H, DDR4):** \~1.0–1.2M msg/s at \~0.5 KB payload, \~2.4–2.6 GB/s at \~3 MB payload
* **WSL2 (Intel i7-13700H, DDR4):** \~450–520K msg/s at \~0.5 KB payload, \~5.0–5.6 GB/s at \~3 MB (after network tuning)
* **Linux VM on macOS:** results close to macOS native
*What is included in these numbers:* TCP `localhost` loopback, 1:1 (provider -> mtrouter -> consumer). The stack includes: event creation, multithreading, event queuing, event dispatching, and socket communication. Service discovery, automatic framing, and thread dispatch are all active.
*Note:* the benchmark example (`23_pubdatarate`) is optimized for throughput pressure. It pre-builds messages to maximize network load.
*What is not included:* The consumer has a lower stable dispatch ceiling before its internal queue grows unbounded and memory climbs: *\~300–400K msg/s* on Windows, *\~500–600K msg/s* on macOS. Above those rates the dispatch thread becomes the bottleneck. I'm working on improving it.
**Missing: bare-metal Linux**
All Linux measurements so far come from VMs (WSL2 and a Linux VM on macOS). VM overhead varies too much to extrapolate confidently. Based on macOS results, bare-metal x86-64 Linux should be (estimated) \~*6.0–7.5 GB/s* for large payloads and \~*2.0–3.0M msg/s* for small payloads, *\~500K msg/s* stable dispatch.
The benchmark is self-contained. If anyone runs it on bare-metal Linux, the results would be interesting to compare. README has exact commands to set:
[https://github.com/aregtech/areg-sdk/tree/master/examples/23\_pubdatarate](https://github.com/aregtech/areg-sdk/tree/master/examples/23_pubdatarate)
**Useful data points:**
* distro/kernel
* CPU model and class (mobile/desktop/server)
* RAM type (DDR4/DDR5/LPDDR5)
* measured msg/s and data rate per recipe
* stable consumption before memory growth starts.
Happy to answer questions about methodology or implementation.
P.S. About areg-sdk: C++17 Service-oriented framework with ORPC for distributed development. Targets: Linux/Windows/macOS; RTOS is the next step. GitHub: [https://github.com/aregtech/areg-sdk](https://github.com/aregtech/areg-sdk)
https://redd.it/1til14z
@r_cpp
Post #25237
17