JTjason.teixeira() Glossary
Home / Learn / Glossary / Reranker
RAG & Retrieval

Reranker

A reranker is a second pass over your retrieved chunks that re-sorts them by how useful they actually are for the question. Your first retrieval casts a wide net and ranks fast but roughly. The reranker reads each candidate against the query more carefully and pushes the genuinely relevant ones to the top, so the model sees the best few first.

Why it matters

Vector search is good at "roughly related" and weak at "actually answers this." So the real answer often sits at rank 7 while three near-misses hog the top slots. If you only feed the model the top 3, it never sees the good chunk and starts guessing. A reranker fixes the ordering before anything reaches the model, which lifts context precision without you touching the retriever.

How it works

You retrieve wide first, say the top 50 by embedding similarity. Then you hand those 50 plus the query to a cross-encoder that scores each pair directly. It reads the query and chunk together, which is slower but far sharper than the initial search. Keep the top 5 or so by that new score and pass only those to the model. It is a quality-for-latency trade: more candidates reranked means better ordering and a few hundred more milliseconds.

In practice

A support bot gets asked "can I expense a flight I already booked?" Plain vector search returns the general travel policy and two booking how-tos up top, with the actual reimbursement rule buried at rank 8. The reranker reads all eight against the question and promotes the reimbursement rule to first. Now the bot answers from the right paragraph instead of the closest-sounding one.

Want this checked on your own AI feature?
Get a free mini-eval — real findings on your live feature, no call required.
Free mini-eval →
© 2026 Jason Teixeira · Sage Ideas LLC · Glossary · Learn · privacy