/ Vector DB RAG Calculator
Search Directory (144) Vector DB Hub Master Matrix
✦ HNSW MEMORY OVERHEAD & RAG ECONOMICS 2026

Vector Database RAG
Unit Economics & RU Calculator

Estimate RAM consumption with 40% HNSW graph index bloat. Compare Pinecone Serverless Read Units vs self-hosted Qdrant and Weaviate.

Accounts for 2.5x HNSW index overhead
5 providers: Pinecone, Qdrant, Weaviate, Chroma, pgvector
1536 / 3072 / 1024 / 768 dimension models

Your RAG Workload Parameters

E.g. 5M chunks = ~1,250 books at 400 words/chunk
Storage Footprint Preview
Raw vector data 30.7 GB
With HNSW index overhead 76.8 GB
Monthly read units (Pinecone) 20M RU
Pinecone Serverless Monthly Cost
$185 / mo
Storage: $25.34 · Reads: $160.00

Provider Comparison

Qdrant Cloud (Managed)
Rust-native · RAM-based billing · HNSW on-disk
$95 / mo
Save $90 / mo (−48.6%)
Weaviate Cloud (Serverless)
Hybrid vector + BM25 keyword · GraphQL API
$135 / mo
Save $50 / mo
Supabase pgvector
PostgreSQL extension · Ideal for <1M vectors
$25 / mo
Save $160 / mo on small workloads
Chroma (Self-Hosted)
Free OSS · EC2 c5.xlarge hosting cost
$122 / mo
Self-host savings at scale
Architecture Verdict For 5M vectors at 1536 dims + 10M monthly queries, Qdrant Cloud is the best balance of cost and performance. pgvector is cheapest but may degrade in recall quality at this vector count without careful IVFFlat tuning.
Storage Math

Vector Storage Footprint by Model & Scale

Each vector is stored as 32-bit floating point numbers (4 bytes each). For 1536-dimension embeddings, one vector is 6,144 bytes. With HNSW indexing at 2.5x overhead, 1M vectors occupies ~15.4 GB on disk before metadata and payload storage.

Vectors 1536-dim (GPT/Ada) 3072-dim (Large) 1024-dim (Voyage) 768-dim (BERT) Pinecone Cost (1536d)
100K0.6 GB raw · 1.5 GB indexed1.2 GB raw · 3.1 GB0.4 GB raw · 1.0 GB0.3 GB raw · 0.7 GB~$0.50/mo
1M6.1 GB raw · 15.4 GB indexed12.2 GB raw · 30.7 GB4.1 GB · 10.2 GB3.1 GB · 7.7 GB~$5/mo
5M30.7 GB raw · 76.8 GB indexed61.4 GB raw · 153.5 GB20.5 GB · 51.2 GB15.4 GB · 38.4 GB~$25/mo
50M307 GB raw · 768 GB indexed614 GB raw · 1.5 TB205 GB · 512 GB154 GB · 384 GB~$253/mo
500M3.1 TB raw · 7.7 TB indexed6.1 TB raw · 15.3 TB2.0 TB · 5.1 TB1.5 TB · 3.8 TB~$2,541/mo
Indexed storage = raw × 2.5x HNSW overhead. Pinecone cost based on $0.33/GB-month storage only (excludes read unit billing).
Provider Profiles

Vector Database Providers In-Depth

Pinecone Serverless
Storage$0.33 / GB-month
Read Units$0.000008 / RU (2 RU/query)
Write Units$0.000002 / WU
Free Tier5GB · 100K vectors
Index TypeProprietary ANN
Best ForManaged, no-ops teams
Scales to billions of vectors. No infrastructure management. Cold start latency on serverless can be 1–3s for idle indexes. Dedicated pods are available at $0.096/hr per p1 pod.
Qdrant Cloud
Pricing ModelRAM cluster-based
4GB RAM Cluster~$25 / month
8GB RAM Cluster~$50 / month
Free Tier1 GB storage, 1 cluster
Self-HostedFree, MIT license
Query Latency<10ms p99 in-mem
Rust-based engine. Fastest ANN latency. Quantization (scalar, product, binary) reduces RAM 8x–32x. Supports on-disk HNSW for cost-efficient large indexes.
Weaviate Cloud
Storage$0.00 (included)
Dimensions Unit$0.05 / 1M dim-queries
Sandbox (Free)14-day free cluster
Hybrid SearchVector + BM25
Multi-ModalText + Image + Video
Self-HostedFree, BSD 3-Clause
Unique dimension-based billing: total cost = queries × dimensions × $0.00000005. Best for hybrid semantic + keyword search use cases like enterprise search.
Supabase pgvector
Pro Plan$25 / month
Storage Included8 GB
Extra Storage$0.125 / GB
Index TypeHNSW / IVFFlat
Max Vectors~1M (practical)
SQL InterfaceFull PostgreSQL
Best value for small RAG apps and prototypes. Combined with existing Supabase auth/db usage. Performance degrades above 1M vectors without sharding or dimension reduction.
Chroma (Self-Hosted)
LicenseApache 2.0 (Free)
EC2 c5.xlarge~$122 / mo (on-demand)
EC2 c5.xlarge RI~$73 / mo (1yr reserved)
Storage (EBS)$0.10 / GB-month
Chroma CloudComing soon (beta)
Language SDKsPython, JS/TS
Zero per-query cost. Full control over hardware. Ideal for on-premise compliance workloads or very high query volumes where Pinecone RU costs exceed hosting.
Milvus (Self-Hosted)
LicenseApache 2.0 (Free)
Zilliz CloudFrom $65 / mo
Disk Index (DiskANN)Up to 10B vectors
GPU SupportNVIDIA CUVS
Index TypesHNSW, IVF, DiskANN
FilteringMetadata + scalar hybrid
Industry standard for large-scale production RAG at 100M+ vectors. Complex Kubernetes deployment but unmatched throughput with GPU acceleration.
Architecture Guide

How to Choose the Right Vector Database for Your RAG Stack

Decision Tree by Scale

# Choose based on vector count & query volume: if vectors < 500K and team < 5: → Supabase pgvector Cheapest, already have PostgreSQL elif vectors < 10M and queries/mo < 50M: → Qdrant Cloud or Weaviate 40-60% cheaper than Pinecone elif vectors < 100M and no-ops needed: → Pinecone Serverless Best fully-managed experience elif vectors > 100M or compliance needed: → Milvus self-hosted (Zilliz Cloud) DiskANN scales to 10B+ vectors

Cost Reduction Techniques

# 1. Dimension Reduction via Matryoshka # text-embedding-3-* supports truncation: embedding = openai.embed(text, dimensions=512) # 512d vs 1536d: 3x storage savings # 2. Scalar Quantization (Qdrant) # float32 → int8 = 4x RAM reduction # 3. Product Quantization (Qdrant/Milvus) # float32 → ~32-bit code = 32x compression # Recall tradeoff: 0.97 → 0.92
The Hidden Cost of Read Units (Pinecone) At 10M queries/month, Pinecone charges 10M × 2 RU × $0.000008 = $160/month just in read units often exceeding storage costs. If you're query-heavy (search-first, write-rarely), Qdrant Cloud's RAM-based pricing becomes dramatically cheaper above 5M monthly queries.
FAQ

Vector Database Pricing Common Questions

How much does Pinecone Serverless cost per month?
Pinecone Serverless charges $0.33 per GB-month for stored vectors (after HNSW 2.5x overhead is applied to raw size) plus $0.000008 per Read Unit (each ANN query consumes 2 RU by default). For 5M vectors at 1536 dimensions with 10M monthly queries: ~$25 storage + $160 reads = ~$185/month. The free tier includes 5GB storage and 100,000 vectors.
Is Qdrant cheaper than Pinecone?
For most production workloads, yes. Qdrant Cloud charges by cluster RAM size, not by read units making it 40–60% cheaper for query-heavy workloads. A 4GB RAM cluster handles approximately 500K–1M 1536-dim vectors in memory at ~$25/month with unlimited queries. At 10M+ monthly queries, Qdrant's flat cluster pricing vs. Pinecone's $0.000008/RU model generates significant savings.
What is write amplification in vector databases?
HNSW (Hierarchical Navigable Small World) index construction creates multiple levels of graph connections for fast ANN search. Each inserted vector writes its own data plus graph pointers, consuming approximately 2.5x the raw float32 storage. For 5M vectors at 1536 dimensions: raw data = 30.7 GB, but after HNSW indexing = ~76.8 GB. This directly multiplies storage billing at Pinecone ($0.33/GB) and affects RAM cluster sizing at Qdrant.
Can I use pgvector instead of Pinecone to save money?
Yes, for small workloads. Supabase pgvector starts at $25/month on the Pro plan with 8GB included storage vs $185+/month for equivalent Pinecone Serverless. However, pgvector's HNSW performance begins degrading above ~1M vectors without careful IVFFlat tuning and sharding. For prototypes, internal tools, and workloads under 500K vectors, pgvector delivers excellent cost efficiency with the bonus of full SQL access for filtering.
How does vector dimension size affect monthly cost?
Vector dimensions multiply storage linearly. At 5M vectors: 768-dim = 15.4 GB raw, 1536-dim = 30.7 GB raw, 3072-dim = 61.4 GB raw. With 2.5x HNSW overhead these become 38 GB, 77 GB, and 154 GB respectively. At Pinecone's $0.33/GB rate, switching from 1536-dim to 3072-dim adds ~$25/month in storage cost alone. OpenAI's text-embedding-3-small supports Matryoshka dimension reduction to 512 dims with minimal quality loss a 3x storage saving.
What is Weaviate's pricing model?
Weaviate Cloud Serverless uses a dimension-query billing model: cost = (queries × dimensions) × $0.00000005. For 10M queries at 1536 dimensions: 10,000,000 × 1,536 × $0.00000005 = ~$768/month significantly more expensive than alternatives at high query volume. Weaviate is better positioned for moderate query volumes where hybrid vector + BM25 search capabilities justify the premium. Self-hosted Weaviate is BSD-3 licensed and free.