The tool itself is here: https://johngear.dev/n-gpt_rag_llm_service/

1. Preprocessor (video -> text).

Let's imagine we have a large amount of information in the form of video lectures. For example, these are videos from YouTube that we carefully selected during our study of machine learning.

To extract textual information from these videos, we need to pull out the subtitles. I chose the most suitable OSS tool in terms of resources/quality — Faster Whisper.

Faster Whisper works well even on a CPU, doesn’t require any fine-tuning, and within 5 minutes of deployment in your virtual environment, it’s ready to “fight” with the videos.

The result at this stage: second-by-second transcription of all our videos in the form of txt documents.

2. Data normalization.

Now that we have data in the form of txt documents, we need to “groom” it and bring it to a uniform format. This is necessary for correctly splitting such data into chunks later.

At this stage, I confess, I didn’t overthink it: timestamps were removed, text was joined into a single line (line breaks were removed), and some technical symbols were deleted (commas and periods were left alone).

3. Chunking.

We need to take all our texts and split them into chunks of a certain size (I wrote more about this in the article: https://johngear.dev/articles/chankovanie-v-rag.html)

I ended up with 1667 chunks. I cut the text to CHUNK_SIZE = 1600 characters (minimum of 1000) by finding the first period < CHUNK_SIZE (since a period in my texts indicates the end of a sentence). I used an OVERLAP of 250 characters. For a baseline, I consider this a successful solution at the start.

4. Converting text to embeddings.

This is where things got really interesting. First, I dived into the open-source bge-m3, which I installed locally. I converted the chunks into embeddings and saved them in the database. But then I realized that bge-m3 works slowly on a CPU, and since we will also need to convert user questions into embeddings in production, the model would simply lag on my simple Linux server. That’s an architectural mistake.

But I quickly found a solution: I used the API in OpenRouter (you can use OpenAI) to get the embedding model text-embedding-3-small from OpenAI. Yes, even though it costs money, it actually solves two serious problems: speed — text-embedding-3-small via API works dozens of times faster than bge-m3 did on my local machine, and cost — $0.00002 per 1000 tokens, and even my 1.5 million characters cost me 0.75 cents in embedding vectors.

And since my final inference is on a budget server (2 CPUs, 4 GB RAM), that architecture just flies.

5. Vector DB.

Everything here is quite standard, and I didn’t reinvent the wheel. Qdrant, and that’s it.

6. Retriever.

Overall, this part is also simple. distance=Distance.COSINE. I wanted to create a maximally simple baseline without any frills to understand the quality of the “out of the box” solution. And overall, I’m satisfied with that quality.

Cosine similarity

7. Generative model.

Following the logic with the embedding model, I decided to use the LLM via the OpenRouter API, specifically gpt-4o-mini.
The advantages are the same. The model is high quality, and the API is quite stable, plus it’s cheap: Input $0.15 per 1 million tokens, Output $0.60 per 1 million tokens. Additionally, OpenRouter always provides access to two providers, Azure or Fireworks, and if one of the providers is down, sometimes the token prices drop lower than in the OpenAI API (use this life hack).

I experimented a bit with prompting the model. However, you can check it out yourself on GitHub at N_gpt/src/rag_service.py.

8. And of course, the API.

I couldn’t do without the good old grandpa Flask. I wrapped everything in an API. Standard: @app.get(\"/health\"), @app.post(\"/ask\") for questions, and I also added @app.get(\"/logger\") for logging, so you could “play around” and show the user how the “response” traveled on its way to them.

In the end, I shoved everything into Docker and deployed it on my server.


So the final scheme turned out to be quite simple:
embedding → Qdrant → top_k chunks → prompt → LLM → answer


What I consciously didn’t do was Reranker. Yes, a cross-encoder reranker significantly improves the final output. But in my project, it’s unnecessary. My information structure is arranged in such a way that it flows sequentially from topic to topic, and there are minimal repetitions. Plus, there aren’t that many chunks — only 1667. During testing, I realized that I didn’t need a reranker, and the top-N candidates already look decent at the cosine similarity search stage.