Hi everyone,
I’ve been working on a high-performance infrastructure project where the core data engines are written in Rust and C++ (using Polars and custom C-style binary protocols), but the orchestration and business logic are handled by Python 3.13+.
In the C++ world, we often avoid Python for anything "mission-critical" due to the GIL and the overhead of the Python C-API. However, I’ve been experimenting with a Shared Memory IPC approach to decouple the high-level logic from the data plane.
I’d love to get some feedback from the C++ community on this architecture:
1. Zero-Copy IPC via Shared Memory: Instead of using Protobuf or JSON over a socket, I'm allocating segments in the Windows Kernel and using
struct packing to write raw bytes. For those of you building C++ engines that need to talk to high-level "glue" languages, do you still prefer Unix Domain Sockets/Named Pipes, or has Shared Memory become your standard?2. Memory Alignment and Padding: When interfacing Python’s
struct module (C-style) with C++ structs, I've had to be extremely careful with Little Endianness and memory alignment. Is there a more robust way to handle this without bringing in heavy dependencies like FlatBuffers?3. CPU Affinity: I'm pinning the Python consumer to specific cores to avoid context switching when reading from the shared buffer. In a hybrid system, do you usually reserve specific cores for the C++ engine and leave others for the management scripts, or do you let the OS scheduler handle it?
The Goal:
I'm trying to build a system where Python handles the "thinking" (logic/orchestration) while the "doing" (I/O and computation) happens at $O(1)$ or $O(\\log n)$ at the hardware level.
I'm curious: What is your "threshold" for moving a component out of Python/Go and into pure C++? Is it strictly latency, or is it about memory safety and deterministic behavior?
https://redd.it/1syrus9
@r_cpp