🔬 ML PAPERS Дайджест (Image,video,text - arXiv, Feb 13 2026)
🔥 MonarchRT — Efficient attention for real-time video generation via Monarch matrix factorization. Makes autoregressive video DiT viable.
→ arxiv.org/abs/2602.12271
🔥 DreamID-Omni — Unified human-centric audio-video gen. Multi-person identity + voice disentanglement in one framework.
→ arxiv.org/abs/2602.12160
UniT — Unified multimodal CoT with test-time scaling
→ arxiv.org/abs/2602.12279
UniDFlow — Discrete flow matching for multimodal understanding + generation + editing
→ arxiv.org/abs/2602.12221
DeepGen 1.0 — Lightweight unified model for image gen & editing
→ arxiv.org/abs/2602.12205
FAIL — Adversarial imitation learning for flow matching post-training (no reward model needed)
→ arxiv.org/abs/2602.12155
GigaBrain-0.5M — VLA from world model RL (robotics)
→ arxiv.org/abs/2602.12099
Post #4887
3.41K