TGViewer
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security Data eXplore : Data Science, ML, Big Data, LLMs and AI Security @dataxplore · 583 subscribers
Post #1965 182
Convert PDF files into clean data, ready for LLM.

Dolphin is a document parsing framework that converts PDFs into structured formats: Markdown, HTML, LaTeX, and JSON.

Works and Features:

Stage 1️⃣ Detailed analysis of the layout at the page level. Elements and their order are determined according to the natural reading order.

Stage 2️⃣ Parallel parsing of elements using different types of anchors and task-specific prompts.

Key features:

» Open-source
» A two-stage approach of analyze-then-parse based on a single VLM
» Encouraging performance on document parsing tasks
» Generation of a sequence of elements in the natural reading order
» Heterogeneous anchor prompts for different types of document elements
» An efficient parallel parsing mechanism


GitHub

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
More from @dataxplore
  1. Oct 1, 2026Monitoring and debugging "silent" token drift in LLM pipelines: frequency analysis of embe…
  2. Sep 30, 2026Online detection of feature collisions in TDA transformation When using Topological Data A…
  3. Sep 18, 2026Am going to announce something big (for me, it's really big) on October 11, 2026.
  4. Sep 14, 2026Post #2188
  5. Aug 31, 2026I joined a Russian community on Telegram. They share some Russian startup and technology u…
  6. Aug 22, 2026Post #2185
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →