So, each of us has encountered a “smart AI assistant” at least once in our lives when calling a bank, a mobile operator, or writing to tech support on some website.

This is exactly where the RAG approach (Retrieval-Augmented Generation) comes into play.

To put it simply: RAG is a way to make a neural network respond not from its head, but based on real data.

Unfortunately, the neural networks you’ve ever used (ChatGPT, Alisa GPT, DeepSeek, etc.) are not trained to provide you with answers about the operating hours of a Sberbank branch or, for example, how to block a SIM card. But they can do that. How?

How it works

The logic is actually quite simple:

  1. Imagine you have a knowledge base of your organization (documents, articles, lectures, company database)
  2. A user asks a question via a call to support or in a chat with an operator
  3. The algorithm compares the user's question with the answers available in the knowledge base
  4. Then the algorithm sends the matched pieces to the neural network and tells it:
    “Here is the information from our knowledge base ‘’. And here is the user's question ‘’. Answer the user based on the context of the information from the knowledge base.”
  5. The neural network responds based on the found information

Thus, you get the impression that there is some super-intelligent trained neural network on the other end of the line/chat that knows everything about the bank/operator/store.

But in reality, the neural network simply responds according to the context it has.

Example:
You are going to an exam, and you will be asked about a specific subject. To answer the questions, you need to read several books. The logic in RAG is absolutely the same.

Why data is important

The most important thing is that you can take any neural network — smart and expensive or not very smart but cheap, but the quality of the data in the base must constantly strive for perfection.

After all, if “garbage” gets into the neural network regarding your question, it won’t be able to provide any coherent answer.

That’s why some bots/AI operators are so annoying and can’t say anything substantial. They are simply poorly trained.

Despite the seemingly simple architecture of such systems, the quality of the data (its labeling) often suffers greatly. Therefore, instead of a smart AI assistant, you often get an annoying, clueless bot that doesn’t let you reach an operator — a real person for a truly important question.

How it works under the hood

Now for the techies, for those interested in how it works under the hood:

  • First, the data needs to be prepared. That is, “clean” our lectures/instructions/manuals from unnecessary junk (chapters, paragraphs, irrelevant symbols). Then break all the data into chunks.

  • Next, we need to convert these chunks into vectors (this is called embeddings). Each chunk is translated into a numerical representation - a vector. This is done by a special neural network trained to convert texts into numbers. The vectors are stored in a special vector database. Why do we need vectors? The answer is coming.

  • Finding the answer to the user's question. Here’s where it gets interesting: the question that came from the user must also be converted into a vector (through a special neural network). Then we find the nearest chunks using a special vector comparison method (cosine similarity). Yes, we compare the user’s question vector with all the vectors of our chunks in the vector database. Even if there are a million, such comparisons happen quickly in the vector database. After that, we get the top N most relevant chunks (i.e., those chunks in the knowledge base that are similar to the user’s question).

  • Generating the answer by the neural network (LLM): after we find the top N similar chunks, we take these pieces + the user’s question and send them to the LLM with a prompt (instruction):
    “Here’s the user’s question, here’s the information from our database, give an answer based on the context.”

Conclusion

So, in practice — 80% of RAG quality depends not on the model, but on the data.

To put it very briefly: RAG = information retrieval + answer generation based on it.

Where RAG is used

Typical cases:

  • chatbots with documentation
  • knowledge base search
  • support assistants
  • phone assistants
  • educational services

A simple example

And finally, a simple example from life:

Suppose we have a database of lectures on Machine Learning.
We create a chatbot where the user asks: “What is a derivative?”

Now, applying RAG, we can use our lectures on Machine Learning to have the model explain something like a teacher in a lecture.


I actually created such a system to show how it works, visit my website: https://johngear.dev/n-gpt_rag_llm_service/