Contact, if you are interested in this website / domain name / Sponsorship / Advertisement / Partnership

RAG explained: retrieval-augmented generation in plain language

A clear walkthrough of how RAG lets AI answer from your own documents, where it breaks, and how to test it.

A language model only knows what was in its training data, and that knowledge has a cut-off date. It has never seen your product manuals, your internal policies or the contract you signed last week. Ask it about them and it will either admit it does not know or, worse, produce a confident answer that sounds right and is not.

Retrieval-augmented generation, usually shortened to RAG, fixes this by adding a research step. Before the model answers, the system searches your documents for the passages most relevant to the question and hands them to the model along with the question. The model then writes its answer from that material, ideally citing where each point came from.

It is the approach behind most "chat with your documents" tools and many internal help desks. This guide explains each part in plain language, compares RAG with fine-tuning, and shows you where these systems fail and how to test them.

How RAG works, step by step

A RAG system has two phases. Preparation happens once, and again whenever documents change. Answering happens every time someone asks a question.

Preparation

  1. Collect the documents: PDFs, web pages, help articles, spreadsheets, transcripts.
  2. Clean them: strip navigation menus, headers, footers and duplicate copies.
  3. Split them into chunks, usually a few paragraphs each.
  4. Convert each chunk into an embedding and store it in a searchable index, along with the original text and metadata such as title, date and source link.

Answering

  1. Convert the user's question into an embedding.
  2. Search the index for the chunks closest in meaning.
  3. Optionally re-rank those results to put the most useful ones first.
  4. Send the top chunks plus the question to the language model with instructions to answer only from the provided text and to cite sources.
  5. Return the answer with its citations.

Embeddings: turning meaning into numbers

An embedding is a long list of numbers that represents the meaning of a piece of text. A specialised model produces it. The key property is that texts with similar meanings get similar lists of numbers, even if they share no words.

So "How do I get my money back?" and "What is your refund policy?" end up close together, while "money market account" lands somewhere else entirely. You can picture each chunk as a point on a map with hundreds of dimensions, where nearby points are about similar things.

This is what makes RAG better than a basic keyword search for natural questions. People rarely phrase a question using the same words as the document that answers it. Embeddings bridge that gap. For more on the vocabulary, see our AI glossary.

Want this working in your business, not just on paper? Get a free, written AI starting plan.

Get my free AI plan

Chunking: cutting documents into useful pieces

You cannot embed a 200-page manual as one item; the meaning would be blurred and the result too large to pass to the model. So documents are cut into chunks. How you cut them has a large effect on quality.

  • Too small and a chunk loses context. A sentence that says "This does not apply to annual plans" is useless if the chunk does not say what "this" is.
  • Too large and a chunk mixes several topics, so its embedding matches many questions weakly and none strongly.
  • Respect structure. Split on headings, sections and paragraphs rather than at an arbitrary character count. Keep tables and lists intact.
  • Add overlap. Letting each chunk share a sentence or two with its neighbours reduces the chance of cutting an idea in half.
  • Attach context. Prefix each chunk with its document title and section heading so it still makes sense on its own.

Vector search and re-ranking

Embeddings are stored in a vector index, a database built to find the stored points nearest to a query point quickly, even across millions of chunks. Vector search is fast and good at meaning, but it has blind spots. It can miss exact terms such as product codes, part numbers, names and acronyms.

That is why many systems use hybrid search: run a traditional keyword search and a vector search in parallel, then merge the results. Keyword search catches "error E-4402"; vector search catches "the dishwasher won't drain".

Re-ranking is a second pass. The first search retrieves a generous set of candidates, say a few dozen. A re-ranking model then reads the question and each candidate together and scores how well that passage actually answers the question. Only the best few go to the language model. Re-ranking is slower per item than vector search, which is why it runs on a short list rather than the whole collection, but it often produces a noticeable lift in answer quality.

Metadata filters help too. If a user is asking about the current policy, filter out archived versions before searching. If they are in one region, filter to that region's documents.

Grounding and citations

Grounding means the answer is based on the retrieved text rather than the model's general knowledge. You encourage it through the instructions sent with the chunks, for example: answer only using the sources below; if the sources do not contain the answer, say so; cite the source number after each claim.

Citations matter for two reasons. They let users check the answer against the original, which builds justified trust. And they make errors visible: if a claim cites a source that does not support it, you have found a problem to fix. Good interfaces link each citation to the exact passage, not just the document.

Grounding reduces made-up answers but does not eliminate them. A model can still misread a passage, combine two sources incorrectly, or fill a small gap with a guess. Writing clear grounding instructions is a prompting skill; our prompt engineering guide covers the techniques.

When RAG beats fine-tuning

Fine-tuning means further training a model on your own examples so that its behaviour changes. People often assume fine-tuning is how you "teach a model your data". For factual knowledge, RAG is usually the better tool.

FactorRAGFine-tuning
Best atSupplying facts and specific informationShaping style, tone, format and narrow task behaviour
Keeping knowledge currentEasy: update the documents and re-indexHard: requires retraining
Citing sourcesNatural: answers point to retrieved passagesNot possible in a reliable way
Access controlCan filter documents per userAnything learned is available to every user
Removing informationDelete it from the indexVery difficult once trained in
Upfront effortModerate: pipeline, index, testingHigher: curated training data, training runs, evaluation
Per-question costHigher, because each prompt includes retrieved textLower prompts, but training and hosting costs

The two are not mutually exclusive. A team might fine-tune for a consistent house style and use RAG to supply the facts. But if your goal is "answer questions about our documents accurately", start with RAG. For help weighing models generally, see choosing an AI model.

Common failure points

  • The answer was never retrieved. The right passage exists but did not make the top results, due to poor chunking, vocabulary mismatch or missing keyword search. This is the most common failure.
  • Messy source data. Scanned PDFs with poor text extraction, tables flattened into gibberish, or boilerplate repeated on every page.
  • Conflicting or outdated documents. Three versions of the same policy, and the system retrieves the old one.
  • Questions that need the whole collection. "How many contracts mention penalty clauses?" requires counting across everything, which retrieving a handful of chunks cannot do.
  • Too much context. Stuffing in many marginally relevant chunks can distract the model from the one that matters.
  • Ignoring permissions. If the index contains HR files, someone may retrieve them. Apply access controls at retrieval time.

The last point connects to broader data handling; our AI privacy and security checklist covers it.

How to evaluate a RAG system

Judge a RAG system with evidence, not a handful of demo questions. A practical approach:

  1. Build a test set. Collect fifty to a few hundred real questions, ideally from actual users or support logs. For each, record the correct answer and which document contains it. Include questions the documents cannot answer.
  2. Test retrieval on its own. For each question, check whether the correct passage appears in the retrieved results. If it does not, no amount of prompt tuning will save the answer.
  3. Test the answers. Check whether each answer is correct, whether every claim is supported by its cited source, and whether the system correctly says "I don't know" for unanswerable questions.
  4. Change one thing at a time. Adjust chunk size, add hybrid search or add re-ranking, then rerun the same test set and compare.
  5. Monitor in production. Log questions, retrieved passages and user feedback. Review a sample regularly and add new failure cases to the test set.

A language model can help grade answers at scale, but spot-check its grading by hand. Automated graders make mistakes too.

To find ready-made RAG and document-chat tools, browse our AI tools directory. If you are planning a system over business documents and want a second opinion, our AI services team can help.

Frequently asked questions

Does RAG stop AI from making things up?

It reduces the problem considerably when retrieval works, because the model has the facts in front of it. It does not eliminate it. Citations and a clear "say you don't know" instruction make remaining errors easier to catch.

Do I need a dedicated vector database?

Not necessarily. Small collections can work with simple libraries or with vector features added to databases you already use. Dedicated vector databases become more useful as collections grow large or need advanced filtering.

How often should I re-index my documents?

Whenever the source content changes in a way users would care about. Many teams update the index automatically when a document is added or edited, and run a full rebuild periodically to remove stale entries.

Can RAG work with spreadsheets and databases?

Partly. Text-heavy rows can be embedded like documents, but questions that need calculations or filtering across many rows are better handled by letting the model query the data directly, sometimes alongside RAG.

This guide is general information, not professional advice. Spotted an error? Tell us.