Spatial-TTT adapts "fast weights" to capture and structure spatial information from long video streams. This allows models to form structured 3D spatial memory over time.
Key ideas:
☞ Efficient streaming memory
Fast weights work as compact spatial memory.
Memory growth is sublinear even for videos longer than 7000 frames, while computations are reduced by more than 40%.
☞ Spatial-predictive mechanism
TTT layers with 3D spatial-temporal convolution capture geometric correspondences and temporal continuity.
☞ SOTA results
The model shows the best results on long-term spatial understanding of video (VSI-Bench) tasks.
Work took 1st place in the Daily Papers ranking on Hugging Face on March 13th.
Project, GitHub, Article, Models and data on HuggingFace
••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore