Skip to main content

> RAG_MANUAL_04

Manual 04: Hybrid Search & Cross-Encoder Reranking

BM25 lexical search, Reciprocal Rank Fusion (RRF), and deep cross-encoder token-level reranking.

Canonical Engineering Manual #04|TinyCTO RAG Bible

Hybrid Search & Cross-Encoder Reranking

BM25 lexical search, Reciprocal Rank Fusion (RRF), and deep cross-encoder token-level reranking.

#1. The Failure of Pure Dense Retrieval

Pure dense vector search struggles with out-of-vocabulary technical acronyms, exact product SKUs, part numbers, and legal citations. Conversely, pure keyword search (BM25) fails on synonymy and semantic paraphrase. High-reliability RAG demands a hybrid approach.

Reciprocal Rank Fusion (RRF)

RRF combines ranked candidate lists from heterogeneous retrieval algorithms without requiring calibrated score normalization:

RRF Score(d∈D)=∑m∈M1k+rm(d)\text{RRF Score}(d \in D) = \sum_{m \in M} \frac{1}{k + r_m(d)}

where kk is a smoothing constant (typically 60) and rm(d)r_m(d) is the rank of document dd in system mm.


#2. Two-Stage Retrieval with Cross-Encoders

  • Stage 1 (Candidate Generation): Fast approximate search over vector and BM25 indices fetches top 50 to 100 candidate chunks.
  • Stage 2 (Cross-Encoder Reranking): A deep transformer (e.g. bge-reranker-large) simultaneously evaluates the concatenated query-document pair, capturing cross-attention token interactions to filter the top 5 to 10 highest-quality passages.