In a post on X, @ai_for_success says Microsoft has released MAI-Transcribe-2-Streaming, a real-time speech-to-text model. The post reports 2.5% word error rate (WER) for both the final transcript and the first partial transcript, a 0.13-second delay after speaking stops, and a $0.54-per-hour streaming price.

Image shared by @ai_for_success on X
Image shared by @ai_for_success on X

Image credit: @ai_for_success on X

The supplied post does not cite independent benchmark or product documentation for these figures. It also does not establish API availability, supported languages, evaluation conditions or production performance.

What the report says about the model

The post describes MAI-Transcribe-2-Streaming as a real-time transcription model and reports first place for both final-transcript and first-partial-transcript accuracy. The attached image visibly shows the final-transcript leaderboard, labeled “Streaming Speech to Text Leaderboard,” where MAI-Transcribe-2-Streaming appears first with a 2.5 AA-WER streaming index.

The image places Grok Voice Transcribe 2.0 second at 2.7, followed by other entries. The leaderboard’s methodology and evaluation conditions are not provided, so the displayed ranking does not independently establish comparative performance.

How to read the reported accuracy and latency

WER means word error rate. It is calculated from substitutions, deletions and insertions divided by the number of words in the reference transcript. Lower is better. A reported 2.5% WER therefore indicates fewer errors than 2.7% only if both figures were produced under comparable test conditions.

The post gives 2.5% WER for the final transcript and 2.5% for the first partial transcript. A final transcript is the text produced after the relevant speech has been processed. A first partial transcript is an early result produced while speech is still being processed. The same reported WER does not establish that the early and final transcript text was similar; it only gives the same score for two measures whose calculation details are not provided.

The reported 0.13-second figure is described as the delay after the speaker stops talking. That could matter for interactive applications if measured consistently, but the report does not say whether it represents end-to-end latency, processing time or another interval. It also does not provide the audio conditions, hardware, network setup or evaluation procedure needed to interpret the number fully.

MAI-Transcribe-2-Streaming versus Grok Voice Transcribe 2.0

The post compares the two models on reported WER and delay:

Item

Reported accuracy

Reported latency

Reported price

MAI-Transcribe-2-Streaming

2.5% WER for final and first partial transcripts

0.13 seconds after speaking stops

$0.54 per hour for streaming

Grok Voice Transcribe 2.0

2.7% WER

0.49 seconds

Not supplied

Non-streaming version

Not supplied

Not supplied

$0.10 per hour

On the figures reported in the post, MAI-Transcribe-2-Streaming has a lower WER and shorter post-speech delay than Grok Voice Transcribe 2.0. That comparison does not show that one model performs better across languages, accents, background noise, audio formats or production workloads.

It is also unclear whether the two models were tested with the same dataset or evaluation settings. The ranking is therefore a useful reported comparison, but not a complete purchasing or deployment assessment.

Reported pricing: streaming versus non-streaming transcription

The post reports a price of $0.54 per hour for MAI-Transcribe-2-Streaming. It separately reports $0.10 per hour for a non-streaming version.

These figures describe different operating modes, so the lower non-streaming price should not be treated as the cost of the streaming model. The report does not specify billing units beyond the hourly figures, minimum charges, pricing conditions, included features or where to obtain the models. The figures support a reported price comparison, not a confirmed cost estimate for a real application.

What remains unconfirmed

The evidence does not independently confirm that Microsoft released the model, that it is available through an API or that the reported prices can currently be used by developers. It also does not establish the benchmark dataset, languages, audio conditions, scoring procedure, latency definition or production suitability. Those details are necessary before using the figures to compare transcription models or decide whether to deploy one.

Sources