RAG - retrieval-augmented generation - retrieves relevant content from your own data and adds it to the prompt so the model answers from that context rather than its training memory alone. It grounds answers in current, private sources, but only if retrieval is good and the model stays faithful to the retrieved context.
Why it matters
Models do not know your private or current data. RAG lets them answer from your own sources without fine-tuning, and keeps answers grounded in documents you can cite.
How it works
At query time, relevant passages are retrieved (usually by embedding similarity) and added to the prompt so the model answers from that context. Quality depends on the whole pipeline - chunking, retrieval, and the model's faithfulness - not just the model.