RAG
Retrieval-augmented generation: a pattern where a model is given relevant documents or data retrieved at question time, so its answers rest on current, specific sources rather than only its training.
Retrieval-augmented generation, usually shortened to RAG, is a way of giving a language model information it was not trained on. When a question arrives, the system first searches a body of documents or data, retrieves the most relevant passages, and places them in the model's context. The model then writes its answer with those sources in front of it.
Why it is used
- It lets an LLM work with private, current or domain-specific knowledge without retraining.
- It makes answers easier to cite and check against a source.
- It can reduce, though not remove, hallucination.
What to watch
RAG quality depends on the retrieval step as much as on the model: chunking, search quality, permissions, and freshness of the index all decide what the model sees. If the right passage is not retrieved, the model may answer confidently from the wrong one. Retrieval is also part of the harness around the model, so it has to be evaluated as part of the whole system and not assumed to work.
Related terms
Hallucination
When an AI model produces fluent, confident output that is false or unsupported by its sources. What matters is which errors are unacceptable in a given workflow and how they are detected.
Harness
The software around a model that manages context, tool use, sub-agents and the execution environment. It shapes real-world capability, so it must be evaluated with the model, not ignored.
LLM
Large language model: an AI model trained on very large amounts of text to predict and generate language, which can be used to answer, summarise, write, reason about and act on text.