Topic hub
RAG & Vector Search
Master retrieval-augmented generation, vector databases, embeddings, and semantic search systems.
All posts · 35

RAG & VECTOR SEARCH01
Embedding Dimensions by Model: 384 to 4096 (2026 Table)Exact output dimensions for 30 embedding models: BERT-base 768, bge-m3 1024, Qwen3-8B 4096, plus bytes per vector and the fix for dimension mismatch errors.
RAG & VECTOR SEARCH02
Deep-Search Agents Score Just 33/100 on Enterprise DataOn a 39,190-artifact enterprise corpus, the best agentic RAG scores 32.96 out of 100. The bottleneck is retrieval, not the model.
RAG & VECTOR SEARCH03
MTEB Leaderboard Lies: Embedding Recall in ProductionMTEB leaderboard rank rarely predicts production recall. A 15+ NDCG drop from benchmark to your corpus flags contamination, per ZeroEntropy's analysis.
RAG & VECTOR SEARCH04
Visual RAG vs OCR: ColPali for PDFs, Tables, ChartsColPali scores 81.3 nDCG@5 on ViDoRe vs 66.1 for a strong OCR pipeline. When visual RAG beats OCR for PDFs, tables, and charts, and when it doesn't.
RAG & VECTOR SEARCH05
pgvector HNSW Tuning for 10M+ Rows Without Ripping It Outpgvector HNSW survives 10M+ rows if you size RAM to the index and tune ef_search. Get the params, the maintenance_work_mem fix, and when DiskANN wins.
RAG & VECTOR SEARCH06
Self-Hosted vs API Embeddings: Qwen3, Voyage 2026Qwen3-Embedding-8B tops MTEB multilingual at 70.58 and self-hosts at ~1/20th API cost. A decision framework for self-hosted vs API embeddings in 2026.
RAG & VECTOR SEARCH07
DeepEval vs RAGAS vs TruLens: Pick Your RAG Eval StackRAGAS for fast experiments, DeepEval for CI gates, TruLens for production tracing. The metric-by-metric comparison plus the 2026 production thresholds to set.
RAG & VECTOR SEARCH08
Vector Search at a Billion Vectors: The Cost-Per-QPS MathRecall converges at 95-99% across HNSW engines, so cost at scale is throughput-per-dollar. ScyllaDB hits 252K QPS at 2ms P99 on 1B vectors. Here's the math.
RAG & VECTOR SEARCH09
Reranker Models Compared: Cohere vs Voyage vs Jina vs BGEJina Reranker v3 hits 81.33% Hit@1 at 188ms, the only top-tier sub-200ms model. The latency and NDCG breakdown across Cohere, Voyage, Jina, and BGE.Explore other topics