High-throughput log parsing (~500K lines/sec) in C++ without regex — looking for performance ideas
I’m building a log ingestion + parsing pipeline in C++ and trying to push throughput as far as possible.
Current setup:
\- \~500K lines/sec
\- custom tokenizer (no regex)
\- string_view everywhere to avoid copies
\- batch processing
\- append-only write path
Next step:
I want to optimize the query side using:
\- SIMD for substring search
\- possibly precomputed token patterns
Questions:
\- Best SIMD strategies for substring / token matching?
\- Any experience with AVX2/AVX512 for log-like workloads?
\- At what point does memory bandwidth become the bottleneck?
Also curious if anyone has benchmarked SIMD vs naive scan for log-style data.
Any pointers or war stories appreciated.
https://redd.it/1t37nr3
@r_cpp
Post #25123
15