Microsoft AI released MAI-Transcribe-2 on September 3, 2026, calling it the lab’s most capable transcription model and the cheapest among the rivals it named. The model is priced at $0.10 per hour of audio as a limited-time offer through the end of the year. It is available to demo in Microsoft Foundry, the MAI Playground, and Open Router.

The company said MAI-Transcribe-2 ranks first on FLEURS across 60 languages with an average word-error rate of 5.2 percent. On Artificial Analysis, it sits on the accuracy-and-latency Pareto frontier and ranks second on that firm’s word-error-rate leaderboard. Microsoft’s own comparison says the model is 10 times faster than OpenAI’s GPT-Transcribe, 7 times faster than ElevenLabs’ Scribe v2, and 5 times faster than Gemini 3.5 Transcribe, while posting higher accuracy.

Features aimed at messy, long audio

The launch list is built for production queues rather than clean studio files. Speaker diarization attributes words to the right person. Word-level timestamps support alignment, search, and editing. Keyword biasing is meant to catch names, abbreviations, and domain terms. A verbatim style keeps fillers and false starts for compliance work. A clean style strips those for captions and notes. The model also handles code switching, including Hinglish and Spanglish, plus automatic language identification so callers do not have to label the language first.

Microsoft said long-form inference is up to 10 times faster than leading competitors, and that quality holds up in noisy rooms. The pitch is one model for clinical notes, legal records, accessibility, and captions, instead of a stack of language-specific engines.

The release follows earlier MAI audio work. Microsoft had already shipped MAI-Transcribe-1 alongside MAI-Image-2 and MAI-Voice-1. Transcribe-2 is the version the lab is willing to price as a volume product.

Decoded Take

Ten cents an hour is a dumping price designed to move transcription off Whisper-class self-hosting and off OpenAI’s meter. The interesting part is not the FLEURS average. It is diarization, timestamps, and verbatim mode in the same endpoint, which is what contact centers and clinics actually buy. The risk is the usual Microsoft sandwich: a sharp MAI model that later gets folded into Copilot SKUs with less transparent metering. Watch whether the $0.10 rate survives January, whether independent ASR boards confirm the 5.2 percent FLEURS number, and whether Foundry customers can keep the raw model when a bundled Microsoft 365 transcript is sitting one checkbox away.