๐ฌ Kandinsky 6.0 Video: Training Details and Architecture
The new Kandinsky 6.0 Video lineup โ the flagship Pro (29B) and the lightweight Lite (3B) โ uses a dual-stream architecture: a video stream and an audio stream connected by bidirectional cross-attention.
๐ Training methodology:
โข Audio stream: pretrained from scratch on 40 million audio tracks; joint training: 7 million audio-video segments
โข Selection criteria: high-quality sound, close-up shots of speaking people
โข Objective: precise synchronization of speech and facial expressions
๐ Architecture:
Two connected streams:
1. Video stream
2. Audio stream
The streams are connected by bidirectional cross-attention for temporal and semantic alignment.
๐ฏ Key capabilities:
โข Lip-sync technology
โข Wide audio range (speech, effects, ambient)
โข Consumer GPU deployment (block offloading, 16โ32 GB VRAM presets)
โข Realistic physics modeling
๐ Performance metrics:
โข vs Kandinsky 5.0: clear advantage across all criteria (visual quality, prompt adherence, camera movement)
โข vs LTX 2.5: preferred in visual quality and speech quality
โข vs Veo 3.1 Fast: superior in image animation tasks
๐ง Specifications:
โข Resolutions: SD, HD, Full HD
โข Audio quality: 44 kHz
โข Duration: up to 5 seconds
โข License: MIT
๐ Hugging Face
Post #4655
395

- โค 2