The comprehensive pricing and accuracy guide for dense vector embeddings. Compare cost per 1M tokens, MTEB retrieval benchmarks, and dimension reduction tradeoffs.
The baseline standard for semantic search. Unbelievably inexpensive at $0.020 per 1M tokens. Supports flexible dimension truncation from 1536 down to 512 dimensions to cut Pinecone / Qdrant vector storage fees.
Google's premier lightweight embedding model. Native 768-dimension vectors mean 50% lower RAM and vector index storage footprints in vector databases like pgvector and Milvus.
Engineered specifically for enterprise RAG and code search. Crushes OpenAI in dense domain retrieval and technical documentation search at an economical $0.070/1M rate.
Trained with compression awareness. Native support for int8 and binary embedding compression, reducing vector database memory footprints by up to 96%.
| Model Name | Cost / 1M Tokens | Vector Dimensions | Max Context (Tokens) | MTEB Retrieval Score | Storage Cost Impact | Best For |
|---|---|---|---|---|---|---|
| OpenAI text-embedding-3-small | $0.020 | 1536 (or 512) | 8,191 | 62.3% | Very Low (truncatable) | General purpose search, massive datasets |
| Google text-embedding-004 | $0.025 | 768 | 2,048 | 63.1% | Low (native 768 dims) | Google Cloud native RAG pipelines |
| Voyage AI voyage-3-lite | $0.070 | 512 | 32,000 | 65.8% | Lowest (512 dims) | Long context documents & legal RAG |
| Cohere Embed v3 (int8) | $0.100 | 1024 | 512 | 64.5% | Ultra-Low (int8 quantized) | Enterprise compliance & compressed storage |
| Voyage AI voyage-3 | $0.120 | 1024 | 32,000 | 68.2% | Moderate | Frontier retrieval, zero hallucination RAG |
| OpenAI text-embedding-3-large | $0.130 | 3072 (or 1024) | 8,191 | 64.6% | High (3072 dims) | Deep multi-lingual semantic matching |
Developers frequently obsess over whether embeddings cost $0.02/M or $0.10/M. In reality, generating 10 million tokens of embeddings costs between $0.20 and $1.00 as a one-time API fee.
The ongoing recurring cost is storing those vectors in memory inside Pinecone, Qdrant Cloud, or pgvector. A 3,072-dimension vector takes 4x more RAM and disk space than a 768-dimension vector. By choosing OpenAI text-embedding-3-small truncated to 512 dimensions or Voyage-3-lite (native 512 dimensions), your monthly vector database hosting bill drops by 75%.