I’m not going to disclose the equipment or the organization, because that is not the point. From here on, I’ll keep it simple: there is one complex machine, and I needed to build a local AI agent for searching and analyzing its technical documentation.
The task sounds trivial: build a chat interface, upload the documents, ask a question, get an answer.
In practice, it is much more interesting. You need to design the architecture properly, avoid mixing unrelated data, deploy local models, deal with weird document parsing issues, and solve a lot of other small problems along the way.

First, lock down the architecture
The first decision was not technical but disciplinary: define the MVP and stop adding things to it.
It was very tempting to immediately build a custom web UI, a knowledge graph, automatic folder synchronization, a nice model switcher, separate pipelines, and everything else that usually turns a working prototype into a construction project that never ends.
Instead, I decided to implement a standard architecture:
User
↓
RAGFlow Web UI
↓
a separate knowledge base for one machine
↓
hybrid retrieval
↓
Qwen3-Embedding-8B
↓
Qwen3-Reranker-4B
↓
local LLM or an external research model via API
One important detail: for the first version, I created two chat profiles on top of the same knowledge base.
They use exactly the same documents but different LLM APIs: one is local, deployed on an NVIDIA Jetson, while the second uses an external research model through the OpenRouter API, which gives access to powerful models such as GPT-6.
It is a simple solution, but it significantly reduces the risk of hallucinations.
The knowledge base is the same.
Retrieval is the same.
Only the answer generator changes.
Why RAG instead of simply fine-tuning a model?
With factory documentation, you cannot rely on the model’s memory.
Accuracy matters.
RAG gives us that accuracy because we always know the original source.
If a fine-tuned model confidently invents the purpose of a signal, a module channel, or a connection to an electrical schematic, we may not even notice the mistake immediately.
So I fixed the main principle as follows:
The LLM does not decide whether it should search the documents.
Retrieval is mandatory.
Reranking is mandatory.
Only then does the LLM receive the retrieved context + my prompt.
In other words, the model acts as a careful translator between the human and the documentation fragments that have already been retrieved.
For search, I chose hybrid retrieval.
Semantic search is needed for queries like finding the description of a machine unit by meaning.
Text search is needed for actual factory reality:
X161
D5708
W1482
5A25
CH187
If you search only by meaning, identifiers like these can easily be lost.
If you search only by exact matches, the system becomes bad at understanding normal human questions.
Local embeddings: without them, there is no knowledge base
The first mandatory model was the embedding model.
Without embeddings, documents do not become a vector space, which means there is no proper semantic retrieval.
I deployed a local Qwen3-Embedding-8B service in BF16.
It runs through vLLM in Docker and exposes an OpenAI-compatible /v1/embeddings API.
The requirements were practical:
Russian
English
Chinese
technical terminology
long descriptions
preferably an OpenAI-compatible API
Reranker: the component I always insist on
By default, I consider a reranker highly desirable.
Without embeddings, there is no knowledge base at all.
Without a reranker, search quality is worse, but the system can already be tested.
In the end, I deployed Qwen3-Reranker-4B in BF16.
It also runs locally, on the same Docker network as RAGFlow.
I separately configured a persistent model cache so the model weights would not be downloaded or rebuilt every time the container was recreated.
Why 4B instead of 8B?
According to benchmarks, the difference in reranking quality is small, and in some cases the 4B model even looks better.
Meanwhile, the 8B model is more useful for embeddings because it builds the initial search space.
I tested it not on abstract sentences but on a multilingual industrial example: a Russian query, a relevant Russian fragment, a Chinese description, and irrelevant documents.
The relevant fragments ended up at the top, while the irrelevant ones were pushed down.

Local LLM through an OpenAI-compatible API
For the answer generator, I deliberately added a compatibility layer.
RAGFlow should not need to know whether the backend is specifically Qwen, Ollama, or something else.
For RAGFlow, it is just an OpenAI-compatible endpoint:
/v1/models
/v1/chat/completions
model = qwen3:32b
Bearer auth
The local Qwen3:32B model was connected through my own Flask proxy to Ollama.
This turned out to be the right architectural decision.
If the model starts behaving badly tomorrow, I can replace the backend while keeping the same API contract.
That makes it possible to switch models quickly without changing the architecture of the entire project.
There was also an interesting issue.
After starting the embedding model and the reranker, the old memory guard on the NVIDIA Jetson began blocking the loading of all three models at once, including Qwen3:32B.
I had to investigate the cause, raise the RAM guard limit to 85%, separately adjust the GPU memory guard, and leave a reserve for the model.
It turns out that even a powerful NVIDIA Jetson does not always have enough resources for local AI workloads.
The most important part: the knowledge base
Then came the less glamorous but most useful part: building the document corpus.
For a single machine, the documents were scattered across different sources:
official PDFs, electrical schematics, PLC/HMI/Motion exports, register tables, comments, alarm maps, service materials, research files, and ML artifacts.
The initial inventory looked like this:
617 files total
386 RAG-friendly
231 binary, graphical, or project files
Then I checked for duplicates:
82 exact duplicates by SHA256
43 files were already in RAGFlow
261 new unique RAG-friendly files
Around 280 documents went into processing.
After all repairs and additions, the final knowledge base looked like this:
320 documents fully processed
4 produced zero chunks (I had to investigate them manually)
40,049 chunks
approximately 19 million tokens
Where everything broke
The CSV files were not UTF-8
Some PLC and HMI exports failed during chunking because of their encoding.
Some files were UTF-16, others were cp1251 or cp1250.
The solution was simple: automatically convert them to UTF-8-SIG and submit them for indexing again.
After that, the problematic CSV files were processed correctly and ended up in repair_ok.
Large documents exceeded the embedding model context
Some Markdown/PDF chunks exceeded the embedding model’s context limit.
I had to physically split large Markdown files into smaller self-contained documents.
After that, retrieval became much more predictable.
JSON is not always RAG-friendly
Several structured JSON files produced zero chunks.
I converted them to Markdown while preserving the field structure in this format:
object[index].field: value
After that, the documents indexed normally.
The UI mode mattered more than expected
The most unpleasant bug was not in the models at all, but in RAGFlow’s operating mode.
In Low / Medium / High / Ultra modes, the system switched to an agentic workflow and tried to make the LLM call the rag tool by itself, because that is how it is configured out of the box.
The local Qwen model in this configuration did not properly support that workflow and produced the expected failure.
The logs showed that the actual chunks never reached the model.
Sometimes the model answered strangely.
Sometimes it simply started hallucinating from its own weights.
Switching to Native mode fixed the problem.
After that, the logic became exactly what I needed:
retrieval is mandatory
↓
reranking is mandatory
↓
context is passed to the LLM
↓
the answer is generated from the retrieved sources
For debugging, I also added an audit trail to the LLM API.
It stores the original request, the payload sent to the backend, streaming chunks, the final response, usage, duration, status, and request ID.
That quickly showed that the proxy was not modifying the requests and that the incorrect prompt was coming from higher up in the pipeline.
Final validation test
For the final test, I used a precise industrial question about a PLC signal.
The system had to find a specific identifier, cabinet, module/channel, signal purpose, and neighboring signals.
The system passed the full pipeline:
question
↓
embedding
↓
hybrid retrieval
↓
reranking
↓
context
↓
local LLM
↓
answer with a source reference
The correct table containing the required signal range consistently appeared among the top search results and was passed to the answer generator.
That was an important point.
It was not simply a chatbot producing a nice-looking answer.
The retrieval system actually found the correct technical fragment.
Practical rules I would take away from this project
If I were building a similar documentation system, I would follow these rules:
- Do not mix documentation from different machines into a single knowledge base. For example, identical PLC addresses on different machines can quickly turn search results into a mess.
- Use hybrid retrieval because technical documentation exists simultaneously in semantic meaning and exact identifiers.
- Do not allow the LLM to decide whether search is necessary. In industrial documentation, retrieval should be mandatory.
- A reranker is useful and can be treated as a second quality layer.
- Test the system with real PLC/HMI/electrical identifiers, not demo questions.
- Build an audit trail from the beginning. When the pipeline becomes long, debugging without logs quickly becomes painful.
Conclusion
In the end, I got a working local MVP for analyzing factory documentation: a separate knowledge base for one machine, local embedding and reranker models, a local LLM behind a compatible API, the option to connect an external research model, and a proper end-to-end pipeline from a question to a source-grounded answer.
The most valuable part is not that there is now a chat interface next to the documentation.
The real value is that there is now an engineering-controlled layer between a specialist’s real-world question and a large technical archive: with data isolation, understandable retrieval, logs, a document repair process, and the ability to switch models without rebuilding the entire system.