Project A backend-agnostic C++17 document intelligence pipeline with measurable OCR, layout and table stages
I have been working on an early-stage open-source document intelligence engine in C++17.
The goal is not to own every OCR or layout model. The core keeps typed pipeline boundaries and normalizes PDF text, OCR, layout blocks, tables, reading order, coordinates, confidence and provenance into a stable document model.
Current components include:
\- PDFium rendering and native text
\- PaddleOCR ONNX and optional Tesseract
\- RF-DETR DocLayNet and Paddle PP-DocLayoutV3
\- Table Transformer detection and structure recognition
\- JSON/Markdown/HTML output
\- small public regression datasets
\- an optional FastAPI/Redis Streams/persistent C++ Worker/React inspection platform
It is still an alpha. The current small benchmark subsets are regression checks, not production accuracy claims.
I am looking for contributors interested in C++ architecture, ONNX session reuse, OCR/CV evaluation, Redis Streams recovery, reading order, or Web-based bounding-box diagnostics.
Repository:
https://github.com/ChNanAn/technical-doc-parser
https://redd.it/1v41u7b
@r_cpp
Post #25650
25