🚀 NeuTTS Air - on-device TTS with instant voice cloning
The first realistic speech synthesis model running on-device, without an API.
Format - GGML, which allows it to work on phones, laptops, and even Raspberry Pi.
Voice cloning in 3 seconds: a short audio snippet is enough to construct a voice for subsequent syntheses.
Based on a lightweight language core (0.5 B) + NeuCodec neural codec, providing a balance between quality and speed.
Generated audio is watermarked using Perceptual Threshold Watermarker to combat misuse.
GitHub
🤖 Data Science, ML & Big Data with @DataXplore
Post #1706
120
