Retrieval fetches the passages most relevant to a query, typically by comparing embeddings, and hands them to the model as context. Retrieval quality - precision and recall - sets the ceiling on RAG quality: a model cannot answer faithfully from context that never reached it.

Why it matters

A model cannot answer faithfully from context it never received. Retrieval quality is the hard ceiling on RAG accuracy.

How it works

The query is embedded and compared against stored vectors to fetch the top passages, which are then inserted into the prompt. Precision (no irrelevant passages) and recall (no missing relevant ones) both matter.

Example

If the right passage ranks eleventh and you retrieve ten, the model never sees the answer - and may confidently improvise one.

← Back to the full glossary

Put the platform behind the terms

Route, evaluate, and monitor every AI request from one OpenAI-compatible platform.

Start Free → Explore the Features