Definition
RAG (retrieval-augmented generation) is when an AI system first retrieves relevant documents, then has a language model write its answer from them.
RAG explained
Language models only know what was in their training data, which may be out of date and won't include your private documents. RAG solves this by adding a search step. When a question comes in, the system searches a knowledge source, such as your help center, policy documents or the web, picks the most relevant passages, and gives them to the model along with the question.
RAG is behind two things marketers care about:
- AI search: ChatGPT search, Perplexity and Google's AI features all retrieve web pages before answering, which is why crawlable, clearly written pages can be cited.
- Internal assistants: support bots and knowledge tools that answer from a company's own documents, with citations back to the source.
RAG quality depends mostly on the retrieval: well-organized, up-to-date source content, split into sensible chunks, and a search step that finds the right passages. It reduces hallucination but doesn't remove it, so good systems show their sources.
Example
Your support team builds an assistant that answers customer questions from your help articles. Using RAG, it retrieves the three most relevant articles for each question and answers with links to them, so answers stay current when the articles change.
Why it matters
RAG explains how AI search chooses what to cite, and it is the standard way to build AI tools that answer from your own data.
Related service
Agentic AI Tools
Custom AI agents and automations that connect to your data and tools, with a human in the loop.
Custom agentic AI toolsPublished by Vidern, founded and led by Malhar Shah. Updated .
Have a process an AI agent could run?
Describe the workflow and we'll reply within 1 business day with what we'd build.
Related terms
- EmbeddingsEmbeddings are lists of numbers that represent the meaning of text, so software can find passages with similar meaning even when they use different words.
- LLM (Large Language Model)An LLM (large language model) is an AI model trained on huge amounts of text to generate language, such as the models behind ChatGPT, Claude and Gemini.
- AI hallucinationAn AI hallucination is a confident but false statement from a language model, such as an invented fact, quote, citation or product detail.
- LLM citationAn LLM citation is a link or named source that an AI assistant shows alongside its answer to indicate where the information came from.
- Query fan-outQuery fan-out is a technique where an AI search system splits one question into several related searches, then combines what it finds into a single answer.