Every meeting you sit through produces audio that almost nobody ever listens to again. Same with interviews, lectures, and voice memos. Converting that audio into searchable text is called AI transcription, and for years it cost real money. Microsoft just made it dramatically cheaper, and claims its new model is also the fastest and most accurate one in the world.
That’s a bold triple threat. Here’s what actually happened, what the numbers say, and what it means for you.
Why AI transcription costs money
Transcribing audio used to mean either hiring a human (slow, expensive) or running clunky software that choked on accents and background noise.
Then AI models got good at it. Whisper, OpenAI’s open transcription model, proved that machines could turn speech into text with near-human accuracy. Companies built businesses on top of it, charging per minute or per hour. Speech-to-text became a real industry, and like every AI industry, the pricing started high and has been sliding ever since.
What Microsoft just did
On September 3, Microsoft released MAI-Transcribe-2, priced at 10 cents per hour of audio, according to VentureBeat. That’s about 72% below its own previous model, and it undercuts OpenAI, Google, and ElevenLabs on the same workloads.
The company’s pitch: it’s the fastest, most accurate, and cheapest speech recognition model in the world. Those are fighting words, so let’s look at the evidence instead of the marketing.
According to Artificial Analysis, an independent evaluation outfit that Microsoft cites, MAI-Transcribe-2 runs about 10 times faster than OpenAI’s GPT-Transcribe, seven times faster than ElevenLabs’ Scribe v2, and five times faster than Google’s Gemini 3.5 Transcribe. Voice is a hot battlefield right now: Google just added voice features to Gmail and Docs, and we covered Gemini’s Mac voice commands earlier this year.
On accuracy, Microsoft says the model tops the FLEURS multilingual benchmark across 60 languages, with an average word error rate around 5.2%. It also added speaker diarization, which means it can tell who said what in a conversation, plus word-level timestamps and a keyword biasing feature that lets companies train it on medical or legal vocabulary without fine-tuning.
The price math that matters
Ten cents an hour sounds abstract, so do the math on a real business. A call center processes 100,000 hours of recorded calls per year. That’s a modest volume for a large bank or telecom. At Microsoft’s old price, that bill came to about $36,000 a year. At 10 cents an hour, it drops to $10,000.
Gentle reminder: AI transcription got 72% cheaper in five months. That’s the kind of price war that ends with consumers winning. We looked at what OpenAI’s own chip plans mean for AI prices like these a while back, and this is exactly the scenario that piece predicted.
How it compares to the competition
Here’s where things stand in the speech-to-text market right now:
| Service | Maker | What you should know |
|---|---|---|
| MAI-Transcribe-2 | Microsoft | 10 cents per audio hour (early-bird price); 60 languages; #1 on FLEURS benchmark; speaker diarization + timestamps; available via Azure, Microsoft Foundry, MAI Playground, and OpenRouter |
| GPT-Transcribe | OpenAI | OpenAI’s paid transcription API; Microsoft claims MAI-Transcribe-2 is 10x faster |
| Scribe v2 | ElevenLabs | ElevenLabs’ latest transcription model; Microsoft claims MAI-Transcribe-2 is 7x faster |
| Gemini 3.5 Transcribe | Google’s transcription model, launched in August; Microsoft claims MAI-Transcribe-2 is 5x faster | |
| Whisper V3-Large | OpenAI | The open-source model that started the wave; free if you host it yourself, but you need the hardware |
One honest gap in Microsoft’s claims: it does not claim to beat Alibaba’s models on accuracy, which have led several independent leaderboards through 2026. The company’s answer to Chinese rivals is basically “comparable accuracy, much faster, cheaper, and a vendor your compliance team already trusts.”
Where you can actually try it
You don’t need to be a corporation to poke at this thing. MAI-Transcribe-2 is available through:
- Microsoft Foundry, the company’s AI deployment platform
- Azure Speech, for businesses already on Azure
- The MAI Playground, a public web playground where anyone can test Microsoft’s models
- OpenRouter, which is how many independent developers and hobbyists will reach it
The 10 cents per hour is an early-bird launch price, and Microsoft hasn’t said exactly when it ends. If you’re building something, assume the easy money window is this year.
What it means for normal people
You’ll probably never call a transcription API directly. That’s fine. You’ll feel this through the apps you already use.
Meeting-note apps and dictation tools all run on transcription engines bought on this wholesale market. When the price collapses, one of three things happens: the apps get cheaper, the free tiers get more generous, or the same price buys better accuracy.
Microsoft is also quietly swapping its own models into Word, Excel, and Teams, which Bloomberg reported on in July. That means dictation and meeting transcription inside Microsoft apps should get better and cheaper to run. In a lucky coincidence, we have a guide to the video search tool Clipto that shows what happens when AI gets good at finding things inside recordings, and the same reasoning applies here.
For side hustlers, this is interesting in a different way. Transcription is a classic beginner service business: people pay to turn interviews, podcasts, and old recordings into text. If you’re using a tool that bills you per hour of audio, a 72% price cut changes your margins overnight. The economics of “I’ll transcribe your podcast for you” just got a lot friendlier.
Why Microsoft is doing this
It’s not charity, and it’s not just about transcription. Microsoft has spent two years building its own family of AI models, called MAI, to stop renting intelligence from OpenAI. At its Build conference in June it announced seven new MAI models in a single keynote.
Each one follows the same playbook: target a specific capability, optimize aggressively for cost, price below the frontier labs, sell through Foundry, and quietly swap the model into Microsoft’s own products. Transcription is just the first place this strategy fully matured, according to Mustafa Suleyman, the Microsoft AI chief.
He told The Verge back in April that the first transcription model came from a small, focused team of about 10 people, running at roughly half the GPU cost of other state-of-the-art models. And in April Microsoft formally ended its exclusive access to OpenAI’s technology and stopped the revenue-share payments. The partnership isn’t over, but the training wheels are off.
The honest caveats
Before you go all in, three things worth knowing.
Those speed numbers come from evaluations by Artificial Analysis, cited by Microsoft itself. Nobody’s confirmed “yes, we’re that much slower,” and vendor-published comparisons always deserve a raised eyebrow. Independent testers will sort it out in the coming weeks.
Real-time transcription isn’t addressed. The release focuses on batch processing and long audio files, which is fine for meetings you’ve recorded, less relevant if you need live captions in a call.
And per-language accuracy varies. A 5.2% average error rate across 60 languages hides the fact that some languages do much better than others. English is practically always the best supported.
The takeaway
AI transcription just became a commodity, and Microsoft is the one that made it one. Ten cents per hour of audio, with best-in-class speed claims to back it up, resets the market that OpenAI, Google, and ElevenLabs were happily charging more for.
If you’re a business that processes lots of audio, re-quote your transcription vendors this quarter. If you’re a regular person, expect your meeting-note apps to get cheaper or better, ideally both. And if you’re building anything with voice, this is the moment the economics of AI transcription stopped being an argument against your idea.
The price war is on, and for once, the customer is winning it.