From multimodality to full omnmodal understanding and generation: Speech, Text, Images, Video, audio-video interactions.
🟢 What developers demonstrated?
How to evolutionarily transform ordinary dense LLMs into efficient MoE models capable of working with all modalities simultaneously?
🧠 Architecture
1️⃣ Omnimodality 3D RoPE + Dynamic Capacity MoE
- Unifies alignment of speech, text, images, and video in spatiotemporal dimensions
- Dynamically allocates computation depending on task complexity
2️⃣ Deeply fused multimodal encoder-decoder
- Any combinations of input and output modalities
- True omnmodal interaction and generation
🛠️ Training
1️⃣ Progressive training strategy
Cross-modal alignment → Expert warm-up → MoE + RL → Generative training
- Scales dense LLMs into MoE models
- Only 75B tokens
- Stable convergence, especially on RL
2️⃣ Language foundation for understanding and generation tasks
- All tasks reduce to language generation
- Breaks barriers between modalities
🎨 Capabilities
✔ Generation and interaction through speech
✔ Image generation and editing
✔ Image and video understanding
✔ Audiovisual reasoning
✔ 10+ multimodal tasks
🔥 Results
The model outperformed Qwen2.5-Omni (1.2T tokens) in 50+ of 76 tasks, having only 75B tokens:
- Video understanding: +5%
- Omnimodal understanding: +7%
- Speech QA: +4.3%
- Image processing: +7%
Open Source Model, Code, Homepage
••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore