Cross-Encoder Reranking for RAG
RAG can retrieve a useful document and still produce an off-topic answer if the best evidence appears too far down the list. A Cross-Encoder scores the query together with each candidate and reorders the material before context selection. Its boundary is simple: reranking cannot recover evidence that never entered the candidate pool.
Contents
This follows RRF for hybrid search. Retrieval fusion assembles candidates; reranking examines their match to the current question. 中文版
Where the reranker belongs
A common pipeline is lexical and vector retrieval → merge and deduplicate → rerank → select context → generate. Retrieval searches a large collection. Reranking evaluates only the candidate pool, avoiding expensive pairwise scoring across the entire corpus.

A standard text Bi-Encoder encodes queries and documents separately, allowing document embeddings to be stored in advance. A Cross-Encoder processes the query and candidate together so they can interact during computation. It produces a pair score; a query-independent document embedding cannot directly replace that pairwise computation.
A score is not an accuracy estimate
Write the score for query q and candidate d as s(q,d), then sort in descending order. The range depends on the model and output processing. The Sentence Transformers usage guide explicitly notes that some MS MARCO models return logits.
Applying sigmoid puts values between 0 and 1, but a value of 0.9 does not automatically mean the passage has a 90% probability of being correct. A monotonic transformation can preserve ordering; a probability interpretation needs calibration and evaluation. Raw scores from different models should not be mixed without justification.
Better order, unchanged candidate recall
Suppose B and E are relevant, but retrieval returns only A, C, B and D. Moving B to the first position exposes useful evidence earlier. E remains missing.
| Metric | Before | After |
|---|---|---|
| Candidate-pool recall | 1/2 | 1/2 |
| First relevant rank | 3 | 1 |
| Reciprocal rank for this query | 1/3 | 1 |
RR is 1 divided by the first relevant rank, or zero when no relevant result appears. MRR averages RR over multiple queries. This table contains one query and is not a system benchmark. Candidate-pool recall divides the number of retrieved relevant documents by the number of all labeled relevant documents.
The following Python example was executed and its assertions passed. The scores are invented; no model inference was performed. It checks the logic, not the effectiveness of a trained reranker.
def reciprocal_rank(order, relevant):
return next((1 / rank for rank, doc in enumerate(order, 1)
if doc in relevant), 0.0)
def recall(order, relevant):
if not relevant:
raise ValueError('relevance labels are required')
return len(set(order) & relevant) / len(relevant)
candidates = ['A', 'C', 'B', 'D']
relevant = {'B', 'E'}
# Invented scores illustrate ordering, not model predictions.
scores = {'A': 0.6, 'C': 0.1, 'B': 0.9, 'D': 0.3}
reranked = sorted(candidates, key=lambda d: (-scores[d], d))
assert set(reranked) == set(candidates)
assert reranked == ['B', 'A', 'D', 'C']
assert recall(candidates, relevant) == recall(reranked, relevant) == 0.5
assert reciprocal_rank(candidates, relevant) == 1 / 3
assert reciprocal_rank(reranked, relevant) == 1.0
assert reciprocal_rank([], relevant) == 0.0
print('order:', reranked)
print('candidate recall:', recall(reranked, relevant))
print('RR before / after:', reciprocal_rank(candidates, relevant),
reciprocal_rank(reranked, relevant))
The output order is B, A, D, C. Candidate-pool recall stays at 0.5 while RR changes from approximately 0.3333 to 1. If only the first passage is retained, quality and coverage of that final context need separate measurement.
Choosing the candidate count
A larger pool may increase the chance of finding missing evidence, but requires more pair evaluations. The number of passages sent to generation is a separate parameter: reranking 50 passages and keeping 5 does not mean both stages use top-5.
Those numbers are trial settings, not universal defaults. Compare pool sizes on your own query set while recording candidate recall, final ranking quality, end-to-end latency and resource use. Batching may improve throughput; the gain depends on the model, hardware, text length and concurrency.
Check truncation for long documents. A retrieved candidate can still lose its critical evidence if that sentence falls outside the model input. Chunking, preserving titles and including necessary neighboring context make this failure easier to inspect than feeding an entire long document into a fixed-length input.
Inspect failures before deployment
Separate missing evidence, poorly ordered evidence, redundant selected passages, and generation failures despite adequate context. Only the second directly identifies an ordering problem. Reranking does not establish factual truth and cannot replace access-control filtering.
Use labeled queries that represent the actual languages and domain. Compare ordering before and after reranking on the same candidate pool, so retrieval changes do not masquerade as reranking gains. Track tail latency and failures, and decide how timeouts should fall back—for example, to the original candidate order. Measure fallback results separately.
If relevant evidence never enters the pool, inspect retrieval, chunking and query formulation first. If evidence is already present but irrelevant passages repeatedly displace it from the context, reranking is a more direct intervention. This isolates the bottleneck better than judging only the final answer.
References
Publication note: this article was published as a catch-up on October 8, 2026 (Beijing time), retaining the originally scheduled 14:00 article date. The teaching data is original; no real-model performance test was conducted.
Support
If this article helped you, you can support this site.
Click a QR code to enlarge. More options: support page。


