Qwen has officially released Qwen3-TTS and fully opened up the entire line of models - Base / CustomVoice / VoiceDesign.
🟢 What's inside?
- 5 models (0.6B and 1.8B classes)
- Free-form Voice Design - generating/editing a voice based on a description
- Voice Cloning - cloning a voice
- 10 languages
- 12Hz tokenizer - strong audio compression without significant quality loss
- full support for fine-tuning
- claim SOTA quality on a number of metrics
Previously, the best generators were in closed APIs, but now a full-fledged open-source TTS stack is emerging, where you can:
- train for a domain,
- create custom voices,
- and not depend on a provider.
GitHub | HuggingFace | Demo | Blog | Paper • #AI #TTS #Qwen #OpenSource #SpeechAI
•••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
