Cohere Rerank 4 Fast
Cohere's lower-latency Rerank 4 option for high-volume search, recommendations, and RAG retrieval.
Cohere's lower-latency Rerank 4 option for high-volume search, recommendations, and RAG retrieval.
Plain-English overview
Rerank 4 Fast uses the same basic workflow as Pro: retrieve a candidate set first, then ask the reranker to move the most useful results to the top. It supports multilingual text and semi-structured JSON with the same extended context.
Fast is the operational choice when traffic, response time, and cost carry more weight than squeezing out the strongest possible ordering. The shared API shape makes it practical to test Fast and Pro on the same real queries before deciding.
Category comparison
These are provider-published specifications, not Cody benchmark scores. Follow the linked sources for current limits and endpoint-specific exceptions.
API and provider access
Availability
Available through Cohere's Rerank API, with enterprise deployment choices that can bring the model closer to private data.
Hosted and private-deployment regions depend on the selected Cohere route and agreement.
Check live availabilityData and training
Cohere enterprise customers can opt out of training; SaaS prompts and generations are generally deleted after 30 days. Approved zero-data-retention accounts and private or third-party deployments offer stronger controls.
This is a concise reading of the cited provider material, not legal advice. A third-party gateway can have different storage, routing, training, and residency terms from the model maker's direct API.
Read the provider policyFrequently asked questions
Cohere's lower-latency Rerank 4 option for high-volume search, recommendations, and RAG retrieval. Rerank 4 Fast uses the same basic workflow as Pro: retrieve a candidate set first, then ask the reranker to move the most useful results to the top. It supports multilingual text and semi-structured JSON with the same extended context.
Cohere Rerank 4 Fast was released on December 11, 2025 according to the cited provider materials.
Available through Cohere's Rerank API, with enterprise deployment choices that can bring the model closer to private data. The access routes listed in this guide are Cohere.
$2 / 1,000 searches. One search unit covers one query against up to 100 documents. Longer documents can be split into extra billable chunks.
Available through Cohere's Rerank API, with enterprise deployment choices that can bring the model closer to private data. Hosted and private-deployment regions depend on the selected Cohere route and agreement.
Cohere enterprise customers can opt out of training; SaaS prompts and generations are generally deleted after 30 days. Approved zero-data-retention accounts and private or third-party deployments offer stronger controls. The policy belongs to the provider route and account terms, so verify it again before production use.
Related comparisons
Cohere's quality-first reranker for improving search and RAG results across multilingual text and semi-structured JSON.
Open full comparisonVoyage AI's quality-focused preview reranker, billed by all query and candidate tokens processed in each request.
Open full comparisonJina's compact listwise reranker for long, multilingual, structured, legal, financial, and domain-specific candidate sets.
Open full comparison