An omnimodal model that capable of simultaneously understanding and processing different types of information: text, images, video, and sound.
🟢 How did NVIDIA make a model this efficient?
The model is extremely efficient despite being trained on only 200 billion tokens (which is 6 times less than Qwen2.5-Omni - 1.2 trillion). This was made possible thanks to architectural features and a careful approach to data preparation.
OmniVinci is based on 3 components:
☞ Temporal Embedding Grouping (TEG) - organizes embeddings from video and audio by timestamps.
☞ Constrained Rotary Time Embedding (CRTE) - encodes absolute time.
☞ OmniAlignNet - aligns video and audio embeddings in a common latent space using contrastive learning.
Ablation showed that each element plays its important role: the base model with simple token concatenation scores an average of 45.51 points. Adding TEG raises the result to 47.72 (+2.21), CRTE — to 50.25 (+4.74 from the base), and the final layer in the form of OmniAlignNet brings the average score to 52.59, which in total gives an increase of 7.08 points.
Training data consisted of 24 million dialogues, which were processed through a system where a separate LLM analyzes and combines descriptions from multiple modalities, creating a unified and accurate annotation.
The final dataset consisted of 36% images, 21% sounds, 17% speech, 15% mixed data, and 11% video.
In benchmarks, OmniVinci outperformed all competitors. On Worldsense, the model scored 48.23 points versus 45.40 for Qwen2.5-Omni. On Dailyomni - 66.50 versus 47.45. In audio tasks, OmniVinci also performed well: 58.40 in MMAR and 71.60 in MMAU.
In speech recognition, the model showed a WER of 1.7% on the LibriSpeech-clean dataset.
The model's application was tested in practice. In the task of classifying semiconductor wafer defects, OmniVinci achieved an accuracy of 98.1%, which is better than the specialized NVILA (97.6%) and the larger 40-billion-parameter VILA (90.8%).
Code licensing: Apache 2.0 License.
Licensing: NVIDIA One Way Noncommercial License.
Project page, Model, GitHub
#AI #ML #NVIDIA #OmniVinci
🤖 Data Science, ML & Big Data with @DataXplore