The new unified model from NVIDIA is trained on diverse multimodal data and can combine different types of input signals into a common vector representation.
🟢 What are the Features?
- Searching supports all data types: text, image, audio, video.
- Based on the Qwen Omni architecture (Thinker module, without text generation).
- Context up to 32,768 tokens, embedding size — 2048.
- Optimized for GPU, supports FlashAttention 2.
This makes it ideal for…
- cross-modal search (searching text by video or image);
- enhancing RAG projects;
- multimodal content understanding systems.
Simple, fast, and efficient - all in one open solution.
Open model on Huggingface
#crossmodal #retrieval #openAI #NVIDIA #OmniEmbed #multimodal #AIModels #OpenSource #Search #UnifiedEmbedding
🤖 Data Science, ML & Big Data with @DataXplore