Founder's Take

The Best Speech-to-Text Model Keeps Changing

Meta Muse Voice Transcribe and Microsoft MAI-Transcribe-2 show why dictation products need model choice more than a permanent house model.

A cyber-samurai redirects a glowing route from an older train to a new high-speed engine in a futuristic Japanese rail yard

This week made one thing obvious: the best speech to text model is a moving target.

Meta released Muse Voice Transcribe on September 1. It handles live transcription, endpointing, and speaker diarization for more than 20 speakers. Two days later, Microsoft released MAI-Transcribe-2 with a 2.0% Word Error Rate on Artificial Analysis, very high batch speed, and a temporary price of $0.10 per hour.

If those results hold up in normal use, any dictation company whose main advantage is “we own the speech model” has a problem. Meta, Microsoft, Google, and other large labs can move the frontier faster than most application companies can afford to.

Two releases, one week

The speech recognition frontier moved again

META MUSE VOICE

20+ speakers

Real-time transcription, endpointing, and diarization in one streaming model. Meta says it ranked first on Artificial Analysis streaming speech-to-text as of September 1. source

MICROSOFT MAI-TRANSCRIBE-2

2.0% AA-WER

Second place for accuracy on the current Artificial Analysis non-streaming leaderboard, with a 410.7x median speed factor. source

LIMITED 2026 PRICE

$0.10 per hour

Microsoft says the launch price runs until the end of 2026. The company has not announced what the price will be after that. source

Microsoft did not win every column. It did something harder

Microsoft calls MAI-Transcribe-2 the fastest, most accurate, and cheapest speech recognition model in the world. The independent table is more nuanced.

Artificial Analysis currently places it second for accuracy at 2.0% WER. Nova-3 is faster. Several models are cheaper. What makes MAI-Transcribe-2 impressive is the combination: near-best accuracy, 410.7x median processing speed, and $1.67 per 1,000 minutes at the launch price.

Independent benchmark snapshot

MAI-Transcribe-2 combines accuracy, speed, and price

Artificial Analysis non-streaming API results checked September 5, 2026. Lower WER and price are better. Higher speed is better. source
Model
AA-WER
Median speed
Price / 1,000 min
MAI-Transcribe-2
2.0%
410.7x
$1.67
Scribe v2
2.2%
53.2x
$3.67
Gemini 3.5 Transcribe
2.6%
90.2x
$5.00
GPT Transcribe
3.3%
36.0x
$4.50
Nova-3
5.2%
603.3x
$4.30

That table is only for non-streaming transcription. Meta's release is built for live audio, where delay after someone stops speaking matters as much as batch throughput. Meta says Muse processes audio in 80 millisecond chunks, adjusts its delay word by word, and reaches the speed-accuracy Pareto frontier. Its public demo shows eight speakers talking in the same room while the model labels them in real time.

Owning a model can help. It is not a permanent moat

Some dictation companies are making serious investments here. Aqua built Avalon around developer speech and technical terms. Superwhisper says its S1 models were trained and fine-tuned in house. Wispr Flow layers proprietary models on top of speech recognition. Willow starts with ASR and invests in the edit model that turns raw speech into usable writing.

These are not identical strategies, and building a specialized model can make sense. Aqua's focus on phrases such as git checkout dev is a good example. A company can tune for its users, control latency, protect margins, and improve from product-specific feedback.

The risk is treating model ownership as the whole product. A model that leads today can become average after one release from a larger lab. Users will not accept worse transcription because your margins are better or because training the old model was expensive.

MachinesFluent is built around that reality

I made a different bet with MachinesFluent. I sell the Windows application. I do not need one speech model to remain the winner forever.

MachinesFluent already lets users choose local or cloud speech-to-text and switch speech models from the same app. Local options matter for privacy, offline use, and control. Cloud options matter when a hosted model offers better speed or accuracy. The right choice depends on the job and it can change.

MAI-Transcribe-2 is not in MachinesFluent as I write this. That is the point. I can evaluate it on real dictation, look at the access terms and long-term price, and add it if it earns a place. I do not have to pretend an older model is still better because I own it.

The same rule applies to Muse Voice Transcribe or whatever Google, Meta, Microsoft, NVIDIA, ElevenLabs, Deepgram, or an open-source team releases next. Add strong options. Retire weak ones when they stop earning their place. Let the user choose between local control and cloud performance.

The product has to survive the next leaderboard

No benchmark proves that one model is best for every microphone, language, accent, room, or workflow. WER can also hide the mistakes that matter most to a particular user. A technical term transcribed incorrectly may be far more expensive than a missing filler word. If you want the measurement details, read How to Read a Speech-to-Text Benchmark.

Still, the direction is hard to miss. Speech recognition is becoming infrastructure. The frontier will keep moving, prices will keep changing, and the winner will keep rotating.

I would rather give MachinesFluent users access to the best available work than spend years defending one house model. The app should get better when the industry gets better. It should not be threatened by it.

If you want dictation that gives you local and cloud speech options in one Windows application, try MachinesFluent.

Sources checked

Related reading