Back to Home

Introduction to Rerank: Why do we have to "re rank" after RAG retrieval

September 20, 2026 at 08:02 AMSource: RunByAI0 comment(s)TechGuide

When the Retrieval Enhanced Generative (RAG) system was first launched, many people felt that the answer was "irrelevant": the knowledge base clearly had the correct paragraph, but the model referenced another seemingly relevant text. The problem often lies not in the large model, but in the sorting of the retrieval process - which is exactly what Rerank aims to solve.

1、 The innate shortcomings of vector retrieval

The common first layer search in RAG is vector similarity search: encoding the problem and document into vectors separately, and then using cosine similarity to pick out the closest blocks. It is fast, but has two characteristics:

-Twin tower structure: The problem and document are encoded separately, and they never "meet" during the encoding phase, so it is impossible to model fine-grained interactions between words. Content such as' Apple's revenue 'and' Apple's nutritional value ', which have highly overlapping literal meanings but completely different semantics, may have very close vector distances.

-Equal treatment: Similarity only looks at the overall direction and is not sensitive to which part is the key information.

The result is that recall is usually good, but the top few may not necessarily be the most relevant - that is, the precision is not enough.

2、 How did Rerank fill the position

Rerank's approach is a two-stage search:

Phase 1: Use vector search (or keyword search such as BM25) to quickly extract the top 50 to 100 candidates from the entire knowledge base, pursuing high recall and preferring to have too many candidates.

Phase 2: Use a Cross Encoder model to concatenate the "problem+single candidate document" into one input and feed it into the model. Directly output the relevance score, and then reorder according to the score. Only the top few are handed over to the larger model.

The biggest difference between cross encoders and vector retrieval is that cross encoders allow for full interaction between the question and document within the model, enabling token level attention comparison. Therefore, it is much more accurate to determine whether the question is truly answered in this section.

3、 Why can't we use a cross encoder throughout the entire process

Because it's too slow. The computational complexity of the cross encoder is directly proportional to the number of candidates, and each candidate needs to run model inference once; If we have to score hundreds of thousands of documents one by one, both delay and cost are unacceptable. The vectors for vector retrieval can be pre computed offline, and only approximate nearest neighbor search can be performed online.

So the division of labor between the two is very clear: vector retrieval is responsible for retrieving possible fish from the sea, while Rerank is responsible for picking out the most valuable ones from the web.

4、 Engineering trade-offs

-The larger the number of candidates N: N, the more complete the recall, but the higher the delay. In actual projects, it is common to start from 50 to 100 and then adjust according to the delayed budget.

-Mixed search+Rerank: First, use keyword search and vector search to take a batch each, then remove duplicates and hand it over to Rerank for unified sorting, which is usually more stable than single search.

-Delay and cost: Rerank models are generally not large, with parameter sizes often in the billions, and can be deployed using GPUs or dedicated inference services; Alternatively, candidates can be truncated first and only the top parts can be sorted accurately.

-Evaluation method: Don't just focus on whether the final answer is good or not, you need to measure the recall rate and ranking quality (such as MRR, NDCG and other ranking indicators) separately in order to locate where the problem lies.

5、 Common Misconceptions

-Misconception 1: Thinking that a larger generative model can solve the problem of irrelevant answers. Unable to retrieve, even the strongest model can only be compiled.

-Misconception 2: The larger the cut, the better. The block is too large, with multiple themes mixed into one block, making it difficult for Rerank to determine; The block is too small and the context is incomplete. The slicing strategy needs to be adjusted together with Rerank.

-Misconception 3: Treating Rerank as a panacea. If there are no correct documents in the first stage recall, Rerank is powerless - it can only rearrange, not recall out of thin air.

In summary, Rerank is a cost-effective transformation in the RAG system - it does not change the structure of the knowledge base, nor does it rely on larger generative models. It simply inserts a layer of refinement between retrieval and generation, which can significantly reduce the situation of "the knowledge base has answers, but the model answers incorrectly".

[Reference source] Comprehensive compilation of publicly published academic papers and retrieval framework documents.

Search enhanced generationvector databaseRAG
Discussion

Comments (0)

No comments yet. Be the first!

Leave a Comment