🔬 ML PAPERS (arXiv, Mar 5-6)
🔥 Helios — 14B autoregressive diffusion model for REAL-TIME long video generation. No conventional optimization tricks needed.
📄 https://arxiv.org/abs/2603.04379
🔥 CalibAtt — Training-free 1.58x speedup for video diffusion (tested on Wan 2.1 14B, Mochi 1). Identifies skippable attention patterns offline.
📄 https://arxiv.org/abs/2603.05503
🔥 RealWonder — First real-time action-conditioned video gen from single image. 13.2 FPS at 480x832. Physics sim as bridge. Code + weights open.
📄 https://arxiv.org/abs/2603.05449
• FaceCam — Portrait video camera control via scale-aware conditioning. Custom camera trajectories for face videos. CVPR 2026.
📄 https://arxiv.org/abs/2603.05506
🌐 https://weijielyu.github.io/FaceCam
• CompACT — Compress observations to 8 tokens for world models. Orders of magnitude faster planning. CVPR 2026.
📄 https://arxiv.org/abs/2603.05438
• LSP Scheduler — 3.4x faster inference for Diffusion Language Models (LLaDA-8B, Dream-7B). ICLR 2026.
📄 https://arxiv.org/abs/2603.05454
• VLM Hallucination Detection — Detect hallucinations BEFORE generating tokens. 0.93 AUROC on Gemma-3, Phi-4-VL, Molmo.
📄 https://arxiv.org/abs/2603.05465
Post #4956
3.58K