Qwen has released an open source TTS model that can clone voices, create new ones, and control speech through natural language.
You can just ask
Speak in a cheerful tone with a hint of nervousness
and it will actually do it.
And without all this complex audio engineering.
GitHub
••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore