Retrieval fetches the passages most relevant to a query, typically by comparing embeddings, and hands them to the model as context. Retrieval quality - precision and recall - sets the ceiling on RAG quality: a model cannot answer faithfully from context that never reached it.
Why it matters
A model cannot answer faithfully from context it never received. Retrieval quality is the hard ceiling on RAG accuracy.
How it works
The query is embedded and compared against stored vectors to fetch the top passages, which are then inserted into the prompt. Precision (no irrelevant passages) and recall (no missing relevant ones) both matter.