Mistral OCR 4.1
Mistral's current Document AI OCR model for preserving reading order, tables, structure, locations, and confidence across multilingual files.
Mistral's current Document AI OCR model for preserving reading order, tables, structure, locations, and confidence across multilingual files.
Plain-English overview
Mistral OCR 4.1 turns PDFs, office documents, and images into structured machine-readable content. It can return Markdown, separate tables, paragraph-level bounding boxes, structural labels such as title or footer, and confidence at page, block, or word level.
The base OCR and structured-annotation prices are published separately, which helps teams distinguish simple document ingestion from schema-driven extraction. Online and batch workflows use the same Document AI stack.
Category comparison
These are provider-published specifications, not Cody benchmark scores. Follow the linked sources for current limits and endpoint-specific exceptions.
Pricing & comparisons
Set your usage. Your estimate updates as you type.
Counts pages, not files. A 10-page PDF counts as 10 pages. The estimate uses the document-processing rate available for each model.
Mistral OCR 4.1
Mistral AI
Estimated total (USD)
For the usage above · USD · API pricing, not a subscription
Estimates exclude taxes, tools, cache storage/writes, free allowances and custom discounts. Image estimates cover output only, not prompt or reference-image charges. Quality modes differ by model. Unlisted settings are not treated as free.
API and provider access
Availability
Generally available through Mistral's OCR API and Document AI workflow, with batch and structured-annotation support.
Optional EU and US regional endpoints are available only when the selected model and features are supported; verify OCR coverage before relying on residency.
Check live availabilityData and training
Mistral's commercial terms exclude model training by default except opt-in and preview exceptions. Standard API input and output retention is generally 30 rolling days unless zero data retention is enabled.
This is a concise reading of the cited provider material, not legal advice. A third-party gateway can have different storage, routing, training, and residency terms from the model maker's direct API.
Read the provider policyFrequently asked questions
Mistral's current Document AI OCR model for preserving reading order, tables, structure, locations, and confidence across multilingual files. Mistral OCR 4.1 turns PDFs, office documents, and images into structured machine-readable content. It can return Markdown, separate tables, paragraph-level bounding boxes, structural labels such as title or footer, and confidence at page, block, or word level.
Mistral OCR 4.1 was released on July 16, 2026 according to the cited provider materials.
Generally available through Mistral's OCR API and Document AI workflow, with batch and structured-annotation support. The access routes listed in this guide are Mistral AI.
$4 / 1,000 pages · $5 annotated. Mistral lists $4 per 1,000 OCR pages and $5 per 1,000 annotated pages for structured extraction.
Generally available through Mistral's OCR API and Document AI workflow, with batch and structured-annotation support. Optional EU and US regional endpoints are available only when the selected model and features are supported; verify OCR coverage before relying on residency.
Mistral's commercial terms exclude model training by default except opt-in and preview exceptions. Standard API input and output retention is generally 30 rolling days unless zero data retention is enabled. The policy belongs to the provider route and account terms, so verify it again before production use.
Related comparisons
Cohere's compact document parser for turning enterprise PDFs, presentations, and scans into retrieval-ready Markdown and layout metadata.
Open full comparisonGoogle Cloud's Gemini-assisted parser for preserving document hierarchy and creating context-rich chunks for enterprise search and RAG.
Open full comparisonGoogle Cloud's Gemini-powered Document AI processor for extracting the exact fields and derived values defined in a business schema.
Open full comparison