Back to Speech recognition & transcription

Model comparison

MAI-Transcribe-2 vs Chirp 3

Compare MAI-Transcribe-2 and Chirp 3 using the same provider-sourced speech recognition & transcription rubric. No mystery score and no invented benchmark ranking.

Facts checked September 4, 2026

Estimate your cost

Set your usage. Your estimate updates as you type.

Uses the recording length in hours: 30 minutes = 0.5 hours. Extra features and minimum charges may change the bill.

How this estimate works

Estimates exclude taxes, tools, cache storage/writes, free allowances and custom discounts. Image estimates cover output only, not prompt or reference-image charges. Quality modes differ by model. Unlisted settings are not treated as free.

Estimate your cost
ModelEstimated total (USD)
Chirp 3Google CloudNo reviewed rate
MAI-Transcribe-2MicrosoftNo reviewed rate

Quick take

MAI-Transcribe-2

Microsoft's fast batch transcription model for long recordings, speaker labels, word timing, keyword biasing, and clean or verbatim transcripts.

Best for

  • High-volume recorded calls and meetings
  • Media archives that need speaker and word timing
  • Azure teams that want a managed transcription model

Watch out for

Microsoft's one-hour-in-about-ten-seconds figure is a provider claim, not a guaranteed SLA. Benchmark your own audio and confirm the post-preview price, deployment region, quotas, and retention configuration.

Chirp 3

Google Cloud's general speech-to-text model for streaming, file, and low-cost dynamic-batch transcription across many languages and regions.

Best for

  • Multilingual cloud transcription
  • Teams mixing realtime and batch workloads
  • Google Cloud applications that need regional endpoints

Watch out for

Check the Chirp 3 table for your precise language and region. Speaker labels, timestamps, adaptation, and batch behavior have route-specific constraints that a single coverage number can hide.

Compare the published facts

MAI-Transcribe-2 vs Chirp 3

Values use each provider's own published units and limits. A blank means the provider did not publish a directly comparable value in the sources reviewed.

Speech recognition & transcriptionMAI-Transcribe-2Chirp 3
Transcription priceCurrent provider price per audio hour or minute for the listed processing route.$0.10 / audio hour limited-time rate$0.016/min standard · $0.003/min dynamic batch
Live or batchWhether the model handles realtime streams, uploaded recordings, or both.Fast batch and long-form transcriptionStreaming, synchronous, and batch
LanguagesProvider-published language coverage, separating trained or advertised coverage from specifically verified languages where needed.60 languages85+ languages and locales
Speaker labelsWhether the model identifies who spoke and any published speaker limit.YesSupported for a documented subset of languages
TimestampsAvailable word-, segment-, or utterance-level timing information.Word-level timestampsWord-level timestamps with route constraints
Vocabulary controlKeyword boosting, custom spelling, context, prompting, or other ways to improve domain terms.Keyword biasing, clean/verbatim output, automatic language detectionSpeech adaptation, custom vocabulary, language detection, denoising
Where to use itDirect API, cloud catalog, application, or regional endpoint documented by the provider.Microsoft Foundry / AzureGoogle Cloud Speech-to-Text V2

How to choose

Compare the job, not the hype.

Start with the job you need to complete, then validate cost, access, and policy details on your exact provider route.

MAI-Transcribe-2

Microsoft says prompts, outputs, embeddings, and training data submitted to Foundry Models are not available to model providers and are not used to train foundation models without permission. Retention and abuse-monitoring details depend on the deployed service.

Chirp 3

Google says Speech-to-Text content is not used beyond providing the service unless the customer opts into data logging. Streaming and synchronous content is handled in memory; asynchronous results may be retained temporarily as documented.

Frequently asked questions

MAI-Transcribe-2 vs Chirp 3 FAQ

What is the main difference between MAI-Transcribe-2 and Chirp 3?

MAI-Transcribe-2: Microsoft's fast batch transcription model for long recordings, speaker labels, word timing, keyword biasing, and clean or verbatim transcripts. Chirp 3: Google Cloud's general speech-to-text model for streaming, file, and low-cost dynamic-batch transcription across many languages and regions.

Should I choose MAI-Transcribe-2 or Chirp 3?

Consider MAI-Transcribe-2 when your priority is High-volume recorded calls and meetings. Consider Chirp 3 when your priority is Multilingual cloud transcription. Test both with your own data and provider route before committing.

Is this MAI-Transcribe-2 vs Chirp 3 comparison based on Cody benchmarks?

No. This comparison aligns provider-published facts for the Speech recognition & transcription category. It does not claim a universal winner or combine incompatible third-party benchmark scores.