Cross-Encoder Reranking for RAG

黎 浩然/ 8 10 月, 2026/ 大语言模型/LARGELANGUAGEMODEL/LLM, 机器学习/MACHINELEARNING, 研究生/POSTGRADUATE, 计算机/COMPUTER/ 0 comments

RAG can retrieve a useful document and still produce an off-topic answer if the best evidence appears too far down the list. A Cross-Encoder scores the query together with each candidate and reorders the material before context selection. Its boundary is simple: reranking cannot recover evidence that never entered the candidate pool.

Contents
  1. Where the reranker belongs
  2. A score is not an accuracy estimate
  3. Better order, unchanged candidate recall
  4. Choosing the candidate count
  5. Inspect failures before deployment
  6. References

This follows RRF for hybrid search. Retrieval fusion assembles candidates; reranking examines their match to the current question. 中文版

Where the reranker belongs

A common pipeline is lexical and vector retrieval → merge and deduplicate → rerank → select context → generate. Retrieval searches a large collection. Reranking evaluates only the candidate pool, avoiding expensive pairwise scoring across the entire corpus.

RAG reranking changes candidates A C B D into B A D C; missing document E stays absent
Original diagram. Ordering and scores illustrate the mechanism, not measured model performance.

A standard text Bi-Encoder encodes queries and documents separately, allowing document embeddings to be stored in advance. A Cross-Encoder processes the query and candidate together so they can interact during computation. It produces a pair score; a query-independent document embedding cannot directly replace that pairwise computation.

A score is not an accuracy estimate

Write the score for query q and candidate d as s(q,d), then sort in descending order. The range depends on the model and output processing. The Sentence Transformers usage guide explicitly notes that some MS MARCO models return logits.

Applying sigmoid puts values between 0 and 1, but a value of 0.9 does not automatically mean the passage has a 90% probability of being correct. A monotonic transformation can preserve ordering; a probability interpretation needs calibration and evaluation. Raw scores from different models should not be mixed without justification.

Better order, unchanged candidate recall

Suppose B and E are relevant, but retrieval returns only A, C, B and D. Moving B to the first position exposes useful evidence earlier. E remains missing.

Metric Before After
Candidate-pool recall 1/2 1/2
First relevant rank 3 1
Reciprocal rank for this query 1/3 1

RR is 1 divided by the first relevant rank, or zero when no relevant result appears. MRR averages RR over multiple queries. This table contains one query and is not a system benchmark. Candidate-pool recall divides the number of retrieved relevant documents by the number of all labeled relevant documents.

The following Python example was executed and its assertions passed. The scores are invented; no model inference was performed. It checks the logic, not the effectiveness of a trained reranker.

def reciprocal_rank(order, relevant):
    return next((1 / rank for rank, doc in enumerate(order, 1)
                 if doc in relevant), 0.0)

def recall(order, relevant):
    if not relevant:
        raise ValueError('relevance labels are required')
    return len(set(order) & relevant) / len(relevant)

candidates = ['A', 'C', 'B', 'D']
relevant = {'B', 'E'}
# Invented scores illustrate ordering, not model predictions.
scores = {'A': 0.6, 'C': 0.1, 'B': 0.9, 'D': 0.3}
reranked = sorted(candidates, key=lambda d: (-scores[d], d))
assert set(reranked) == set(candidates)
assert reranked == ['B', 'A', 'D', 'C']
assert recall(candidates, relevant) == recall(reranked, relevant) == 0.5
assert reciprocal_rank(candidates, relevant) == 1 / 3
assert reciprocal_rank(reranked, relevant) == 1.0
assert reciprocal_rank([], relevant) == 0.0
print('order:', reranked)
print('candidate recall:', recall(reranked, relevant))
print('RR before / after:', reciprocal_rank(candidates, relevant),
      reciprocal_rank(reranked, relevant))

The output order is B, A, D, C. Candidate-pool recall stays at 0.5 while RR changes from approximately 0.3333 to 1. If only the first passage is retained, quality and coverage of that final context need separate measurement.

Choosing the candidate count

A larger pool may increase the chance of finding missing evidence, but requires more pair evaluations. The number of passages sent to generation is a separate parameter: reranking 50 passages and keeping 5 does not mean both stages use top-5.

Those numbers are trial settings, not universal defaults. Compare pool sizes on your own query set while recording candidate recall, final ranking quality, end-to-end latency and resource use. Batching may improve throughput; the gain depends on the model, hardware, text length and concurrency.

Check truncation for long documents. A retrieved candidate can still lose its critical evidence if that sentence falls outside the model input. Chunking, preserving titles and including necessary neighboring context make this failure easier to inspect than feeding an entire long document into a fixed-length input.

Inspect failures before deployment

Separate missing evidence, poorly ordered evidence, redundant selected passages, and generation failures despite adequate context. Only the second directly identifies an ordering problem. Reranking does not establish factual truth and cannot replace access-control filtering.

Use labeled queries that represent the actual languages and domain. Compare ordering before and after reranking on the same candidate pool, so retrieval changes do not masquerade as reranking gains. Track tail latency and failures, and decide how timeouts should fall back—for example, to the original candidate order. Measure fallback results separately.

If relevant evidence never enters the pool, inspect retrieval, chunking and query formulation first. If evidence is already present but irrelevant passages repeatedly displace it from the context, reranking is a more direct intervention. This isolates the bottleneck better than judging only the final answer.

References

Publication note: this article was published as a catch-up on October 8, 2026 (Beijing time), retaining the originally scheduled 14:00 article date. The teaching data is original; no real-model performance test was conducted.

Support

If this article helped you, you can support this site.

WeChat support QR code; click to enlarge
WeChat
Alipay support QR code; click to enlarge
Alipay
Buy Me a Coffee; support this site
Buy Me a Coffee

Click a QR code to enlarge. More options: support page。

Share this Post

Leave a Comment

您的邮箱地址不会被公开。 必填项已用 * 标注

*
*