Skip to content
Kentron Technologies

AIRAGBuying guide

RAG, explained for people buying it: making AI answer from your own documents

Retrieval is what stops a model inventing answers about your business. What it is, where it breaks, and the three questions that reveal whether a vendor has built it properly.

Author
Kentron Technologies
Published
Reading time
6 min read
A mobile app home screen

A language model knows what was in its training data. It does not know your price list, your refund policy or which tests your laboratory runs, and when asked it will often produce a confident, plausible, wrong answer. Retrieval-augmented generation, universally shortened to RAG, is the standard fix: before the model answers, the system finds the relevant passages in your own documents and puts them in front of it. This is a buyer's explanation of how that works and where it goes wrong.

What actually happens, in five steps

  1. Your documents are ingested and parsed. PDFs, which are most business documents, have to be turned back into text, and this step is messier than it sounds.
  2. The text is split into chunks, each a few hundred words, because a whole document is too big to hand to a model usefully.
  3. Each chunk is converted into an embedding, a list of numbers representing its meaning, and stored.
  4. When a question arrives, it is embedded the same way, and the system finds the chunks whose meaning is closest.
  5. Those chunks are placed in the prompt with the question, and the model is instructed to answer from them.

The quality of the answer is decided almost entirely in steps one to four. By step five, the model is just writing up whatever it was given. Vendors spend their demo time on step five.

Where it breaks

In our experience four failures account for nearly all disappointing RAG systems.

  • Bad parsing. A scanned PDF with no text layer yields nothing without OCR. A table parsed into a single run of words loses the relationship between a test name and its price, so the model reads the price of the row above.
  • Chunks that split the answer. If a policy's condition is in one chunk and its exception is in the next, retrieving one gives a confidently incomplete answer. This is the most common and least visible failure.
  • Retrieval that is lexically fooled. Ask about a 'renewal fee' when the document says 'annual maintenance charge' and a weak retriever finds nothing, so the model answers from its training data instead.
  • No instruction to refuse. If the prompt does not say what to do when the retrieved passages do not contain the answer, the model will improvise rather than say it does not know.

That last point is the cheapest to fix and the most often skipped. A system that says 'I do not have that information, here is the number to call' is more valuable than one that is right ninety percent of the time and silently wrong for the rest.

The useful question is not whether it can answer. It is whether it will admit when it cannot.

What we did on our own platform

In VoxAgent, a tenant uploads documents which are parsed, including PDFs, chunked and embedded so the agent answers from that tenant's own material rather than improvising. Two decisions mattered more than the retrieval algorithm itself.

First, strict tenant scoping: retrieval is filtered at the query, not in the interface, so one tenant's documents can never surface in another tenant's answer. In a multi-tenant RAG system this is the security property that matters most, and it is invisible in a demo. Second, transcripts: every conversation is stored, so when an answer is wrong you can see which passages were retrieved and fix the document, the chunking or the prompt rather than guessing. The detail is in the case study.

RAG versus fine-tuning, briefly

RAGFine-tuning
Teaches the modelFacts, at answer timeStyle, format and behaviour
Updating contentReplace the documentRetrain
Shows its sourceYes, you can cite the passageNo
Right forYour prices, policies, catalogueA consistent tone or output shape

Almost every business question that sounds like it needs fine-tuning actually needs retrieval. If your content changes, which it does, retrieval is the correct tool, because updating an answer should mean editing a document rather than running a training job.

Three questions that reveal a serious implementation

  1. Show me what it does with a question your documents do not answer. A good system declines. A weak one invents.
  2. Show me a table from one of my PDFs, and ask it something that requires reading two columns. This exposes parsing immediately.
  3. Ask it something using my words rather than the document's words. This tests whether retrieval understands meaning or is matching strings.

Run these against your own documents, not the vendor's sample set. Twenty minutes of this tells you more than a month of proposals.

What it costs to keep running

Embedding your documents is a one-time cost per document version, and it is small. The recurring cost is per question: every query embeds the question, retrieves chunks, and sends those chunks to the model as input tokens. Longer retrieved context means a higher per-question cost, so there is a real trade-off between stuffing in more passages and paying for them. Re-embedding happens whenever documents change, which for a price list is more often than people plan for.

Frequently asked questions

What does RAG mean in AI?

RAG stands for retrieval-augmented generation. Before the model answers, the system searches your own documents for the passages most relevant to the question and includes them in the prompt. The model then answers from that supplied material rather than from memory, which is what makes it possible to get accurate answers about your prices, policies or catalogue.

Does RAG stop AI from hallucinating?

It reduces hallucination substantially but does not eliminate it. The model can still misread a retrieved passage, and if retrieval returns nothing useful and the prompt does not instruct the model to decline, it will improvise. The two things that matter most are good parsing and chunking of your documents, and an explicit instruction to say it does not know when the passages do not contain the answer.

Should I use RAG or fine-tune a model on my data?

Use RAG for facts and fine-tuning for behaviour. If the content changes, such as prices, policies or a catalogue, retrieval is correct because an update means editing a document instead of retraining. Fine-tuning is appropriate when you need a consistent tone or a specific output format, and the two are often combined.

Can a RAG system keep each customer's documents separate?

It must, and this is a question worth pressing in a multi-tenant system. Retrieval has to be scoped at the query itself, so one tenant's chunks are never candidates for another tenant's answer. Filtering in the interface is not sufficient, and the failure is invisible in a demo, so ask specifically where the tenant filter is applied.

Kentron Technologies

Editorial team

Builds and runs Kentron Technologies’s products. Writes here when a decision was hard enough to be worth explaining.

Next step

Tell us what you are running, and what is slow.

A demo of any product, or a conversation about something that does not exist yet. Either way, you will talk to someone who builds the software.

CallWhatsAppTalk to us