Universal-3.5 Pro
AssemblyAI's accuracy-focused transcription model for files and live audio, with code-switching, speaker labels, timestamps, and contextual prompting.
AssemblyAI's accuracy-focused transcription model for files and live audio, with code-switching, speaker labels, timestamps, and contextual prompting.
Plain-English overview
Universal-3.5 Pro is a speech model for product teams that want a focused transcription API instead of a broad cloud catalog. It handles uploaded recordings and realtime streams, supports native code-switching, and can use context or keyterms to improve company names and specialist language.
AssemblyAI separates the base transcript from optional intelligence features and offers US and EU endpoints. That makes the bill and data posture configurable, but it also means teams should compare the complete request—not only the advertised per-hour base rate.
Category comparison
These are provider-published specifications, not Cody benchmark scores. Follow the linked sources for current limits and endpoint-specific exceptions.
Pricing & comparisons
Set your usage. Your estimate updates as you type.
Uses the recording length in hours: 30 minutes = 0.5 hours. Extra features and minimum charges may change the bill.
Universal-3.5 Pro
AssemblyAI
Estimated total (USD)
For the usage above · USD · API pricing, not a subscription
Estimates exclude taxes, tools, cache storage/writes, free allowances and custom discounts. Image estimates cover output only, not prompt or reference-image charges. Quality modes differ by model. Unlisted settings are not treated as free.
API and provider access
Availability
Available through AssemblyAI's US and EU APIs for asynchronous files and realtime streaming.
US and EU endpoints are documented; feature and retention behavior can differ by endpoint and account agreement.
Check live availabilityData and training
AssemblyAI's retention and model-training behavior depends on endpoint, contract, BAA status, EU processing, and opt-out settings. Its docs describe zero-data-retention options for eligible realtime use; verify the exact account configuration.
This is a concise reading of the cited provider material, not legal advice. A third-party gateway can have different storage, routing, training, and residency terms from the model maker's direct API.
Read the provider policyFrequently asked questions
AssemblyAI's accuracy-focused transcription model for files and live audio, with code-switching, speaker labels, timestamps, and contextual prompting. Universal-3.5 Pro is a speech model for product teams that want a focused transcription API instead of a broad cloud catalog. It handles uploaded recordings and realtime streams, supports native code-switching, and can use context or keyterms to improve company names and specialist language.
Universal-3.5 Pro was released on July 7, 2026 according to the cited provider materials.
Available through AssemblyAI's US and EU APIs for asynchronous files and realtime streaming. The access routes listed in this guide are AssemblyAI.
$0.21 / audio hour async · $0.45 / hour streaming. Async and streaming use different routes and rates even when the same model family is selected. Optional intelligence features can add cost.
Available through AssemblyAI's US and EU APIs for asynchronous files and realtime streaming. US and EU endpoints are documented; feature and retention behavior can differ by endpoint and account agreement.
AssemblyAI's retention and model-training behavior depends on endpoint, contract, BAA status, EU processing, and opt-out settings. Its docs describe zero-data-retention options for eligible realtime use; verify the exact account configuration. The policy belongs to the provider route and account terms, so verify it again before production use.
Related comparisons
Meta's streaming speech-recognition model for live captions, long audio, many speakers, code-switching, and domain-aware transcription.
Open full comparisonMicrosoft's fast batch transcription model for long recordings, speaker labels, word timing, keyword biasing, and clean or verbatim transcripts.
Open full comparisonGoogle Cloud's general speech-to-text model for streaming, file, and low-cost dynamic-batch transcription across many languages and regions.
Open full comparisonElevenLabs' multilingual speech-to-text model for detailed transcripts with many speakers, word timing, audio events, and large custom keyterm lists.
Open full comparison