Skip to content

What is Reranking?

Generative AI & LLMs, explained by the engineers who build it. Definition, how it works, use cases and common questions.

Reranking definition

Reranking is a second retrieval stage that takes the top candidates from an initial search and reorders them with a more accurate but slower model, usually a cross-encoder that reads the query and each document together. It improves precision at the top of the results, which matters most in RAG systems that pass only a few documents to an LLM.

Why search needs a second stage

First-stage retrieval must search millions of documents in milliseconds, so it uses fast methods: keyword indexes and vector similarity between precomputed embeddings. Those methods are good at finding a relevant set but imperfect at ordering it. The single best answer might sit at position 12, below documents that merely share vocabulary with the query.

Reranking fixes the order. It takes perhaps the top 50 or 100 candidates, scores each with a model that examines the query and document together, and returns the best few. Because it runs on a small set, it can afford a heavier model without slowing the overall search much.

In a RAG system this matters a great deal: the language model usually sees only the top five to ten chunks, so whatever ranks at position 11 might as well not exist. Better ordering means the answer is grounded in the right evidence.

Bi-encoders vs cross-encoders

Embedding models used for vector search are bi-encoders: they encode the query and each document separately into vectors, which is why document vectors can be computed in advance. That speed has a cost, since the model never sees the query and document side by side and cannot weigh how specific words relate to each other.

Cross-encoders take the query and a document as one input and output a relevance score, letting the model attend to interactions between them. They judge relevance much more accurately but are far too slow to run against a whole corpus, which is exactly why they rerank a shortlist. LLMs can also act as rerankers, at higher cost.

Reranking options

Teams can choose between hosted APIs, open models they run themselves and built-in features of search platforms. Common options include the following, and most can be swapped with little code change if evaluation favors another:

  • Hosted rerank APIs such as Cohere Rerank, Voyage AI and Jina AI rerankers
  • Open cross-encoder models, such as BGE rerankers and MiniLM models trained on MS MARCO, served with sentence-transformers
  • Built-in semantic ranking in platforms such as Azure AI Search, Elasticsearch and Vertex AI Search
  • LLM-based reranking, where a language model scores or orders candidates for complex relevance judgments

Using reranking well

Rerank a candidate set large enough to contain the right answers, often 25 to 100 items from hybrid search, then keep the top few. Measure with ranking metrics such as recall at k and NDCG on a labeled set of real queries, and watch latency, since reranking typically adds tens to hundreds of milliseconds depending on the model and candidate count.

Reranking cannot rescue poor first-stage retrieval: if the right document is not in the candidate set, no reranker will find it. Nexzem tunes both stages together in RAG development projects, adding a reranker when evaluation shows it improves answer quality enough to justify the extra latency and cost.

Reranking: common questions

Something else on your mind? Ask a consultant and get a reply within one business day.

What is a cross-encoder?

A cross-encoder is a transformer model that reads a query and a document together as one input and outputs a single relevance score. Because it sees both at once, it judges relevance more accurately than comparing separate embeddings, but it must run once per document, so it is used only on shortlists.

Does reranking improve RAG answers?

Often, yes. Placing the most relevant chunks at the top means the language model receives better evidence within its limited context, which reduces irrelevant or unsupported answers. The gain depends on how good first-stage retrieval already is, so measure answer quality with and without a reranker.

How many documents should be reranked?

A common starting point is 50 to 100 candidates, returning the top 3 to 10. Too few candidates and the right document may be missing; too many and latency and cost rise. Tune the number on your own queries by checking how often the correct document appears in the candidate set.

Keep exploring the generative ai & llms glossary

Need Reranking in your product?

A solutions consultant replies within one business day with next steps, a rough estimate and a suggested team.