Berk Bayri

RAG

Retrieval-augmented generation: a pattern where a model is given relevant documents or data retrieved at question time, so its answers rest on current, specific sources rather than only its training.

Retrieval-augmented generation, usually shortened to RAG, is a way of giving a language model information it was not trained on. When a question arrives, the system first searches a body of documents or data, retrieves the most relevant passages, and places them in the model's context. The model then writes its answer with those sources in front of it.

Why it is used

  • It lets an LLM work with private, current or domain-specific knowledge without retraining.
  • It makes answers easier to cite and check against a source.
  • It can reduce, though not remove, hallucination.

What to watch

RAG quality depends on the retrieval step as much as on the model: chunking, search quality, permissions, and freshness of the index all decide what the model sees. If the right passage is not retrieved, the model may answer confidently from the wrong one. Retrieval is also part of the harness around the model, so it has to be evaluated as part of the whole system and not assumed to work.