🟢 Why it matters?
The model leads among models with up to 1 billion parameters and encodes queries 7 times faster on regular CPUs.
Unlike decoders that read text left to right and cannot revisit earlier tokens, ModernVBERT uses a bidirectional text encoder trained on word masking and a small visual module.
Each page image is split into patches that are mapped into the same space as the text and then combined with word tokens.
The late interaction mechanism retains vectors of all tokens, allowing each query token to find the most precise match. This combination of bidirectional attention and late interaction outperforms decoder architectures in document retrieval.
Higher page resolution and a short "high-resolution cooldown" phase improve retrieval accuracy, although they may degrade performance on regular images. Adding "text-only" pairs in contrastive learning helps the model effectively unify text and visual spaces.
ColModernVBERT remains compact, demonstrates high benchmark scores, and runs efficiently even on standard CPUs.
🤖 Data Science, ML & Big Data with @DataXplore
