All models
Google Cloud

Chirp 3

Google Cloud's general speech-to-text model for streaming, file, and low-cost dynamic-batch transcription across many languages and regions.

Plain-English overview

What Chirp 3 actually is

Chirp 3 sits inside Google Cloud Speech-to-Text V2, so it supports several operational patterns rather than one fixed app. Teams can use it for live streams, short synchronous requests, or asynchronous batches and can add language detection, speech adaptation, and denoising where the selected region and language support them.

Its feature matrix is not uniform across every language. Diarization and some adaptation options apply to documented subsets, and endpoint location affects availability. The right comparison is therefore an exact language, region, and request mode—not the headline language count alone.

Good fit for

  • Multilingual cloud transcription
  • Teams mixing realtime and batch workloads
  • Google Cloud applications that need regional endpoints

Category comparison

The facts that matter for transcription models

These are provider-published specifications, not Cody benchmark scores. Follow the linked sources for current limits and endpoint-specific exceptions.

Transcription price
$0.016/min standard · $0.003/min dynamic batchCurrent provider price per audio hour or minute for the listed processing route.
Live or batch
Streaming, synchronous, and batchWhether the model handles realtime streams, uploaded recordings, or both.
Languages
85+ languages and localesProvider-published language coverage, separating trained or advertised coverage from specifically verified languages where needed.
Speaker labels
Supported for a documented subset of languagesWhether the model identifies who spoke and any published speaker limit.
Timestamps
Word-level timestamps with route constraintsAvailable word-, segment-, or utterance-level timing information.
Vocabulary control
Speech adaptation, custom vocabulary, language detection, denoisingKeyword boosting, custom spelling, context, prompting, or other ways to improve domain terms.
Where to use it
Google Cloud Speech-to-Text V2Direct API, cloud catalog, application, or regional endpoint documented by the provider.

Pricing & comparisons

Estimate your cost

Set your usage. Your estimate updates as you type.

Uses the recording length in hours: 30 minutes = 0.5 hours. Extra features and minimum charges may change the bill.

Chirp 3

Google Cloud

Estimated total (USD)

No reviewed rate

For the usage above · USD · API pricing, not a subscription

How this estimate works

Estimates exclude taxes, tools, cache storage/writes, free allowances and custom discounts. Image estimates cover output only, not prompt or reference-image charges. Quality modes differ by model. Unlisted settings are not treated as free.

API and provider access

Where to get Chirp 3

Availability

Regions and access stage

Generally available in Google Cloud Speech-to-Text V2 with streaming, synchronous, and batch recognition routes.

Features and supported languages vary by Google Cloud region; choose a regional endpoint from the live Chirp 3 documentation.

Check live availability

Data and training

The route matters.

Google says Speech-to-Text content is not used beyond providing the service unless the customer opts into data logging. Streaming and synchronous content is handled in memory; asynchronous results may be retained temporarily as documented.

This is a concise reading of the cited provider material, not legal advice. A third-party gateway can have different storage, routing, training, and residency terms from the model maker's direct API.

Read the provider policy

Frequently asked questions

Chirp 3 FAQ

What is Chirp 3?

Google Cloud's general speech-to-text model for streaming, file, and low-cost dynamic-batch transcription across many languages and regions. Chirp 3 sits inside Google Cloud Speech-to-Text V2, so it supports several operational patterns rather than one fixed app. Teams can use it for live streams, short synchronous requests, or asynchronous batches and can add language detection, speech adaptation, and denoising where the selected region and language support them.

When was Chirp 3 released?

Chirp 3 was released on October 13, 2025 according to the cited provider materials.

Where can I access Chirp 3?

Generally available in Google Cloud Speech-to-Text V2 with streaming, synchronous, and batch recognition routes. The access routes listed in this guide are Google Cloud.

How much does Chirp 3 cost?

$0.016/min standard · $0.003/min dynamic batch. These are current Google Cloud Speech-to-Text rates for the relevant recognition routes; eligible volume tiers and cloud contracts can change effective cost.

Where is Chirp 3 available?

Generally available in Google Cloud Speech-to-Text V2 with streaming, synchronous, and batch recognition routes. Features and supported languages vary by Google Cloud region; choose a regional endpoint from the live Chirp 3 documentation.

Is my Chirp 3 API data used for training?

Google says Speech-to-Text content is not used beyond providing the service unless the customer opts into data logging. Streaming and synchronous content is handled in memory; asynchronous results may be retained temporarily as documented. The policy belongs to the provider route and account terms, so verify it again before production use.