MAI-Transcribe-1 for speech recognition, MAI-Voice-1 for voice synthesis, and MAI-Image-2 for generating images from textual descriptions.
All of them are positioned as a solution for those who need production-level solutions with competitive inference costs.
➡️ Which three models?
MAI-Transcribe-1
A speech-to-text model with high-speed transcription for 25 languages, including Russian. On the FLEURS benchmark, it shows the best Word Error Rate among competitors: the average value is 3.86%.
The model outperforms Whisper in all 25 languages, and Gemini 3.1 Flash - in 22 out of 25. Supports WAV, MP3, and FLAC formats.
Real-time transcription, diary, and context biasing are not yet available - these features are planned for the future.
Cost: $0.36 per hour of audio.
MAI-Voice-1
A TTS model that generates realistic speech with emotional coloring, natural intonation, and the ability to clone a voice based on a reference.
Access to cloning requires Microsoft's approval and the upload of a recorded consent from the voice owner.
The declared generation speed is 1 minute of audio per second. The model supports emotional management at the level of individual phrases via SSML and is designed for long-form content: audiobooks, podcasts, lectures.
It currently only works with English, but support for more than 10 languages is planned in the future. Available in 3 Azure regions: Central US, Japan West, and Sweden Central.
Cost: $22 per 1 million characters.
MAI-Image-2
A diffusion model for generating images based on a textual prompt, which Microsoft tested in beta testing from March 20th.
The model contains from 10 to 50 billion parameters (excluding embeddings), accepts a context of up to 32K tokens, and generates images with a maximum resolution of 1024×1024 pixels.
According to internal estimates via the Elo rating, MAI-Image-2 scores 1190 ± 8 points compared to 1093 ± 4 for its predecessor MAI-Image-1, particularly strong in photo-realistic and portrait categories (1201 points). On the ArenaAI leaderboard, the model entered the top 3.
Cost: $5 per 1 million text input tokens, $33 per 1 million output tokens (images).
All models are available through Microsoft Foundry. To try them in an interactive environment MAI Playground, you can currently only do so from the USA.
••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
