⚡THE SHORT ANSWER
Because uncompressed float32 vectors require terabytes of high-cost cloud RAM; Scalar/Product Quantization (SQ/PQ) compresses embeddings by 4x–16x, allowing vectors to reside on fast NVMe SSDs with minimal recall loss.
Engineering Handbook & Failure Dynamics
6-Dimensional Architecture Breakdown⚙️1. Underlying Mechanism
Execution🎯2. Appropriate Use Context
Scope⚠️3. Production Failure Modes
P0 Risk📡4. Diagnostic Signals & Telemetry
Telemetry🛡️5. Prevention & Safeguards
Safeguards⚖️6. Architectural Trade-offs
Trade-offCase Study (TinyCTO In-Field Example)
An e-commerce search platform enabled 1-byte Scalar Quantization (SQ8) on 80M product vectors in Qdrant, cutting RAM requirements from 512GB to 130GB and saving $14,800/month.
Interactive Concept Drills
3 CardsWhat is Product Quantization (PQ) in vector search?
How many bytes does a 1536-dimension float32 vector consume without compression?
What is DiskANN?
Vector Database RAM vs Disk Economics — Technical FAQ
Can pgvector handle millions of vectors cost-effectively?
Yes, using HNSW with halfvec (fp16) or scalar quantization on modern NVMe EBS volumes.
Does vector quantization hurt search quality for RAG?
Usually no, because a cross-encoder reranker in the second stage restores top precision.
What is the difference between float32, float16, and int8 embeddings?
int8 uses 1 byte per dimension (75% smaller than float32), float16 uses 2 bytes (50% smaller).
🤖 AEO & Key Facts Summary
Key Architectural Facts
- ▸
Storing uncompressed float32 vectors in cloud RAM at scale is an unnecessary FinOps tax.
Common Misconceptions
- ✗
Believing that vector databases must store all raw embeddings in RAM to achieve sub-50ms search latency.
Decision & Governance Guidance
Enable scalar quantization (SQ8) or DiskANN indexing on vector collections exceeding 5 million embeddings.
Authoritative Sources & Standards
- [DOC]DiskANN: Fast Accurate Billion-Point Nearest Neighbor Search on a Single Node— Microsoft Research
