TGViewer
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security Data eXplore : Data Science, ML, Big Data, LLMs and AI Security @dataxplore · 578 subscribers
Post #2119 214
Microsoft introduced 3 models of the MAI family.

MAI-Transcribe-1 for speech recognition, MAI-Voice-1 for voice synthesis, and MAI-Image-2 for generating images from textual descriptions.

All of them are positioned as a solution for those who need production-level solutions with competitive inference costs.

➡️ Which three models?
MAI-Transcribe-1

A speech-to-text model with high-speed transcription for 25 languages, including Russian. On the FLEURS benchmark, it shows the best Word Error Rate among competitors: the average value is 3.86%.

The model outperforms Whisper in all 25 languages, and Gemini 3.1 Flash - in 22 out of 25. Supports WAV, MP3, and FLAC formats.

Real-time transcription, diary, and context biasing are not yet available - these features are planned for the future.

Cost: $0.36 per hour of audio.

MAI-Voice-1

A TTS model that generates realistic speech with emotional coloring, natural intonation, and the ability to clone a voice based on a reference.

Access to cloning requires Microsoft's approval and the upload of a recorded consent from the voice owner.

The declared generation speed is 1 minute of audio per second. The model supports emotional management at the level of individual phrases via SSML and is designed for long-form content: audiobooks, podcasts, lectures.

It currently only works with English, but support for more than 10 languages is planned in the future. Available in 3 Azure regions: Central US, Japan West, and Sweden Central.

Cost: $22 per 1 million characters.

MAI-Image-2

A diffusion model for generating images based on a textual prompt, which Microsoft tested in beta testing from March 20th.

The model contains from 10 to 50 billion parameters (excluding embeddings), accepts a context of up to 32K tokens, and generates images with a maximum resolution of 1024×1024 pixels.

According to internal estimates via the Elo rating, MAI-Image-2 scores 1190 ± 8 points compared to 1093 ± 4 for its predecessor MAI-Image-1, particularly strong in photo-realistic and portrait categories (1201 points). On the ArenaAI leaderboard, the model entered the top 3.

Cost: $5 per 1 million text input tokens, $33 per 1 million output tokens (images).

All models are available through Microsoft Foundry. To try them in an interactive environment MAI Playground, you can currently only do so from the USA.


••••••••••••••••••••••••••••••••••••••
🤖 Data & ML |
@DataXplore
More from @dataxplore
  1. Sep 18, 2026Am going to announce something big (for me, it's really big) on October 11, 2026.
  2. Sep 14, 2026Post #2188
  3. Aug 31, 2026I joined a Russian community on Telegram. They share some Russian startup and technology u…
  4. Aug 22, 2026Post #2185
  5. Aug 21, 2026Deep systemic analysis of AI constraints from context to internal weight editing. 📂 PDF #…
  6. Aug 17, 2026Adaptive Gradient Thresholding Why Fixed Gradient Clipping Kills Deep RecSys When Feedback…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →