Model comparison
Chirp 3 vs Universal-3.5 Pro
Compare Chirp 3 and Universal-3.5 Pro using the same provider-sourced speech recognition & transcription rubric. No mystery score and no invented benchmark ranking.
Facts checked September 4, 2026
Model comparison
Compare Chirp 3 and Universal-3.5 Pro using the same provider-sourced speech recognition & transcription rubric. No mystery score and no invented benchmark ranking.
Facts checked September 4, 2026
Set your usage. Your estimate updates as you type.
Uses the recording length in hours: 30 minutes = 0.5 hours. Extra features and minimum charges may change the bill.
Estimates exclude taxes, tools, cache storage/writes, free allowances and custom discounts. Image estimates cover output only, not prompt or reference-image charges. Quality modes differ by model. Unlisted settings are not treated as free.
| Model | Access | Estimated total (USD) |
|---|---|---|
| Chirp 3Google Cloud | Google Cloud | No reviewed rate |
| Universal-3.5 ProAssemblyAI | AssemblyAI | No reviewed rate |
Quick take
Google Cloud's general speech-to-text model for streaming, file, and low-cost dynamic-batch transcription across many languages and regions.
Check the Chirp 3 table for your precise language and region. Speaker labels, timestamps, adaptation, and batch behavior have route-specific constraints that a single coverage number can hide.
AssemblyAI's accuracy-focused transcription model for files and live audio, with code-switching, speaker labels, timestamps, and contextual prompting.
Async and realtime routes have different prices, and retention depends on endpoint and account controls. Verify optional feature charges, EU routing, training opt-out, BAA status, and zero-retention eligibility.
Compare the published facts
Values use each provider's own published units and limits. A blank means the provider did not publish a directly comparable value in the sources reviewed.
| Speech recognition & transcription | Chirp 3 | Universal-3.5 Pro |
|---|---|---|
| Transcription priceCurrent provider price per audio hour or minute for the listed processing route. | $0.016/min standard · $0.003/min dynamic batch | $0.21/hour async · $0.45/hour streaming |
| Live or batchWhether the model handles realtime streams, uploaded recordings, or both. | Streaming, synchronous, and batch | Uploaded files and realtime streaming |
| LanguagesProvider-published language coverage, separating trained or advertised coverage from specifically verified languages where needed. | 85+ languages and locales | 18 languages at launch with native code-switching |
| Speaker labelsWhether the model identifies who spoke and any published speaker limit. | Supported for a documented subset of languages | Yes |
| TimestampsAvailable word-, segment-, or utterance-level timing information. | Word-level timestamps with route constraints | Word-level timestamps |
| Vocabulary controlKeyword boosting, custom spelling, context, prompting, or other ways to improve domain terms. | Speech adaptation, custom vocabulary, language detection, denoising | Contextual prompting and keyterm prompting |
| Where to use itDirect API, cloud catalog, application, or regional endpoint documented by the provider. | Google Cloud Speech-to-Text V2 | AssemblyAI US and EU APIs |
How to choose
Start with the job you need to complete, then validate cost, access, and policy details on your exact provider route.
Google says Speech-to-Text content is not used beyond providing the service unless the customer opts into data logging. Streaming and synchronous content is handled in memory; asynchronous results may be retained temporarily as documented.
Provider and API links
AssemblyAI's retention and model-training behavior depends on endpoint, contract, BAA status, EU processing, and opt-out settings. Its docs describe zero-data-retention options for eligible realtime use; verify the exact account configuration.
Provider and API links
Frequently asked questions
Chirp 3: Google Cloud's general speech-to-text model for streaming, file, and low-cost dynamic-batch transcription across many languages and regions. Universal-3.5 Pro: AssemblyAI's accuracy-focused transcription model for files and live audio, with code-switching, speaker labels, timestamps, and contextual prompting.
Consider Chirp 3 when your priority is Multilingual cloud transcription. Consider Universal-3.5 Pro when your priority is Meetings and calls with language switching. Test both with your own data and provider route before committing.
No. This comparison aligns provider-published facts for the Speech recognition & transcription category. It does not claim a universal winner or combine incompatible third-party benchmark scores.