All models
Microsoft

MAI-Transcribe-2

Microsoft's fast batch transcription model for long recordings, speaker labels, word timing, keyword biasing, and clean or verbatim transcripts.

Plain-English overview

What MAI-Transcribe-2 actually is

MAI-Transcribe-2 is aimed at turning large audio files into usable text quickly. Microsoft pairs multilingual recognition with diarization, word-level timestamps, automatic language detection, and controls for whether the transcript preserves filler words or returns cleaner prose.

Access runs through Microsoft Foundry, which makes the model a natural fit for organizations already managing identity, regions, and data controls in Azure. The launch price is explicitly limited-time, so cost projections should retain a margin for a future standard rate.

Good fit for

  • High-volume recorded calls and meetings
  • Media archives that need speaker and word timing
  • Azure teams that want a managed transcription model

Category comparison

The facts that matter for transcription models

These are provider-published specifications, not Cody benchmark scores. Follow the linked sources for current limits and endpoint-specific exceptions.

Transcription price
$0.10 / audio hour limited-time rateCurrent provider price per audio hour or minute for the listed processing route.
Live or batch
Fast batch and long-form transcriptionWhether the model handles realtime streams, uploaded recordings, or both.
Languages
60 languagesProvider-published language coverage, separating trained or advertised coverage from specifically verified languages where needed.
Speaker labels
YesWhether the model identifies who spoke and any published speaker limit.
Timestamps
Word-level timestampsAvailable word-, segment-, or utterance-level timing information.
Vocabulary control
Keyword biasing, clean/verbatim output, automatic language detectionKeyword boosting, custom spelling, context, prompting, or other ways to improve domain terms.
Where to use it
Microsoft Foundry / AzureDirect API, cloud catalog, application, or regional endpoint documented by the provider.

Pricing & comparisons

Estimate your cost

Set your usage. Your estimate updates as you type.

Uses the recording length in hours: 30 minutes = 0.5 hours. Extra features and minimum charges may change the bill.

MAI-Transcribe-2

Microsoft

Estimated total (USD)

No reviewed rate

For the usage above · USD · API pricing, not a subscription

How this estimate works

Estimates exclude taxes, tools, cache storage/writes, free allowances and custom discounts. Image estimates cover output only, not prompt or reference-image charges. Quality modes differ by model. Unlisted settings are not treated as free.

API and provider access

Where to get MAI-Transcribe-2

Availability

Regions and access stage

Available through Microsoft Foundry for batch and long-form speech recognition.

Deployment availability follows the live Microsoft Foundry model catalog and the Azure region chosen by the customer.

Check live availability

Data and training

The route matters.

Microsoft says prompts, outputs, embeddings, and training data submitted to Foundry Models are not available to model providers and are not used to train foundation models without permission. Retention and abuse-monitoring details depend on the deployed service.

This is a concise reading of the cited provider material, not legal advice. A third-party gateway can have different storage, routing, training, and residency terms from the model maker's direct API.

Read the provider policy

Frequently asked questions

MAI-Transcribe-2 FAQ

What is MAI-Transcribe-2?

Microsoft's fast batch transcription model for long recordings, speaker labels, word timing, keyword biasing, and clean or verbatim transcripts. MAI-Transcribe-2 is aimed at turning large audio files into usable text quickly. Microsoft pairs multilingual recognition with diarization, word-level timestamps, automatic language detection, and controls for whether the transcript preserves filler words or returns cleaner prose.

When was MAI-Transcribe-2 released?

MAI-Transcribe-2 was released on September 3, 2026 according to the cited provider materials.

Where can I access MAI-Transcribe-2?

Available through Microsoft Foundry for batch and long-form speech recognition. The access routes listed in this guide are Microsoft AI and Microsoft Foundry.

How much does MAI-Transcribe-2 cost?

$0.10 / audio hour for the limited-time preview rate. Microsoft labels this as limited-time pricing. Confirm the live Foundry catalog before forecasting sustained production cost.

Where is MAI-Transcribe-2 available?

Available through Microsoft Foundry for batch and long-form speech recognition. Deployment availability follows the live Microsoft Foundry model catalog and the Azure region chosen by the customer.

Is my MAI-Transcribe-2 API data used for training?

Microsoft says prompts, outputs, embeddings, and training data submitted to Foundry Models are not available to model providers and are not used to train foundation models without permission. Retention and abuse-monitoring details depend on the deployed service. The policy belongs to the provider route and account terms, so verify it again before production use.