RAG - retrieval-augmented generation - retrieves relevant content from your own data and adds it to the prompt so the model answers from that context rather than its training memory alone. It grounds answers in current, private sources, but only if retrieval is good and the model stays faithful to the retrieved context.

Why it matters

Models do not know your private or current data. RAG lets them answer from your own sources without fine-tuning, and keeps answers grounded in documents you can cite.

How it works

At query time, relevant passages are retrieved (usually by embedding similarity) and added to the prompt so the model answers from that context. Quality depends on the whole pipeline - chunking, retrieval, and the model's faithfulness - not just the model.

Example

A support assistant retrieves the three most relevant help articles and instructs the model to answer only from them, citing the article used.

← Back to the full glossary

Put the platform behind the terms

Route, evaluate, and monitor every AI request from one OpenAI-compatible platform.

Start Free → Explore the Features