Resources

Top AI Embedding Models for Search & RAG

Embedding models turn text, code, images, audio, and documents into vectors for semantic search, recommendations, clustering, and RAG. Compare cost, context, dimensions, accepted inputs, retrieval controls, and hosting choices.

Showing 1–7 of 7 models

What to compare

Embedding price
Current provider price per million input tokens or the closest published billing unit.
Input capacity
Maximum content accepted in one embedding input, using the provider's documented token basis.
Vector dimensions
Supported output sizes; smaller vectors reduce storage while larger vectors may preserve more information.

How to use this library

Compare the job, not the hype.

A curated cross-category set of frontier, specialist, and category-defining models with practical API information for real product decisions.

Cody does not run a universal quality leaderboard. Provider claims are labeled, pricing excludes taxes and optional tools, and production buyers should verify the live provider terms before choosing a model.

  1. 01

    Use first-party documentation for technical limits, lifecycle, pricing, regional access, and data policies whenever it exists.

  2. 02

    Compare models only inside the same category and retain each provider's measurement basis.

  3. 03

    Show Not published when a provider has not published a comparable fact instead of estimating it.

  4. 04

    Treat prices and availability as dated snapshots and link every page to live provider documentation.

Speed keeps its units

Voice latency, text throughput, video render time, and world-model frame rate are not interchangeable.

Unknown stays unknown

A missing release date, country list, or latency number is labeled as unpublished instead of estimated.

Every fact has a date

Model pages link to the underlying source and show when pricing, access, and policies were last checked.

Put the model to work

Pick the model, then give it a better brief.

Use Cody's expanded prompt playbooks and curated agent skills to turn a model choice into a repeatable workflow.