Before getting to the model, you have to navigate through documentation, old databases, strange exports, security restrictions, and the question of “can we even trust this data?”.
I can’t talk about specific equipment and organizations, so I’ll keep it generic. But the story is typical for industrial automation: there’s a machine, a PLC, an HMI, registers, no data collection at all, but there’s a new storage scheme. And all of this needs to be carefully connected so that nothing gets lost or broken.

First — Understanding Sources, Not Code
The most unpleasant thing about factory data is that it rarely resides in one neat place. Some information is in manuals, some in electrical diagrams, some in the HMI, some provided by Chinese/European partners, and some needs to be checked directly through the PLC (bombing the PLC with shaking knees).
So the first step wasn’t “let’s quickly train the model,” but rather gathering and structuring what we actually have.
A colleague compiled an archive of manuals and electrical diagrams. From this, we separately extracted information on D-registers and general information about the machine. We didn’t conduct a full analysis of the entire archive right away — there wasn’t time for that. But just having the export in a structured format, for example in md and json, significantly changes the situation.
When registers are only in PDFs or diagrams, they exist as “documentation.” When they turn into a proper catalog, you can start working with them as data: searching, comparing, verifying, linking to future collector signals.
Why a Signal Catalog is More Important Than It Seems
In ML, you usually want to quickly get a table: timestamp, features, target. But in production, you can’t just pull addresses out of thin air, otherwise you won’t understand what to search for and, more importantly, where.
You need to understand:
- what type of device or register it is;
- how to read it safely;
- whether it can be read in groups;
- what the signal means at least roughly;
- how confident we are in that meaning;
- from which document the information came.
So a separate task became to think through the structure of the “books” of signals. In my case, these are YAML books, from which we can then build a reading plan. This approach is convenient because the same catalog can be used for both verification and the future collector.
For example, the smoke test should take one YAML book, extract addresses from it, read them from the machine via MC/SLMP, and save the result. Not guessing the meaning of the signal, not making conclusions like “the machine is currently in this mode,” but simply answering basic questions:
- is the address readable or not;
- what is the current value;
- are there any errors for the group of addresses;
- what is the reading delay;
- can this type of signal be included in the future collector (where we will accumulate them for the ML model).
This is boring work, but it saves a lot of time later. If an address isn’t readable, it’s better to find that out during the smoke test than during the dataset collection for the model.
Here, the principle of trust but verify applies.
Safety: Only Read-Only and No Heroics
In industrial automation, there’s an important rule: if you’re not sure — don’t write.
For the smoke test, I set strict limitations:
no write commands
no force commands
no reset commands
no batch writes
read only
one read thread
one active MC/SLMP session
a short pause between read groups
And as you can imagine, first, I had to figure this out (ACS TP - hello).
For PLC outputs — only reading the state. For D-registers related to parameters — only reading as a word, without writing to recipes, servo, or other settings.
This may sound overly cautious, but in a factory, caution is not bureaucracy; it’s a normal engineering instinct because if the machine “burps” at $1.5 million, it won’t seem trivial to anyone. The smoke test should safely show which addresses are available for reading.

SQLite as the First Version
Next began a separate story about data storage.
It’s important to clarify: there was no old inherited database that someone had maintained for years before me. SQLite came about because, in the first phase, I needed to quickly and safely start writing data. I didn’t fully understand how to properly lay out the architecture yet, so the first collector was made through SQLite: locally, clearly, without unnecessary infrastructure. What could possibly go wrong?
For a first version, this is a normal path. It’s better to start collecting data in a simple way than to spend months drawing the perfect scheme and end up with nothing at all (System Design - hello).
But then the simple scheme started hitting reality.
The collector was constantly writing to SQLite. The exporter also wanted to work with this same SQLite: reading the queue, checking rows, confirming transfers, removing already sent data. As a result, due to race conditions around writing and exporting, the collector was constantly occupying the database, the queue was piling up, and normal transfers to ClickHouse began to interfere with the actual data writing.
So the problem wasn’t that SQLite is bad in itself. The problem was that it became simultaneously:
- a local storage;
- a queue;
- a source for transfer;
- a place where the collector actively writes.
At a small volume, this is still tolerable. But when data is flowing constantly, such a scheme quickly turns into a bottleneck.
Next, I came to the solution of getting rid of SQL and writing data in batches directly into the RAM of the factory PC, and from there sending it via API to the server with ClickHouse. I won’t describe the detailed scheme here, as it’s the subject of a separate article.
Why ClickHouse Fits Well Here
ClickHouse is convenient not because it’s a trendy database, but because factory data quickly becomes a history of measurements.
There are many rows, many timestamps, many repeated queries by periods, signals, and identifiers. For such a scenario, ClickHouse fits more naturally than local SQLite. And for a large time series — it’s a must-have.
SQLite was a good first step for me: to start quickly (THE MAIN THING), check the approach, and not complicate the start. But when a proper history, analytics, and future ML pipeline appear, you want not just to “write somewhere,” but to have a database where it’s convenient to verify data and quickly retrieve the necessary slices.
Speaking of ClickHouse, who else can handle that much data in a week of collection: 2676886030 (yes, that’s the number of metrics collected from the machine in a week, not a ticket number).
Practical Benefits: What I Would Take to Any Similar Project
In short, my checklist after this work looks like this.
1. First the Catalog, Then the Collector
Don’t start with reading all addresses in a row. It’s better to gather signal books: address, type, source, approximate meaning, confidence, category. This will help later during the ML model training phase.
2. Smoke Test Should Be Read-Only
Its task is to check reading availability, not to “poke the machine.” No writes, forces, resets. One thread, one session, pauses between groups. The documentation for the specific PLC needs to be studied separately (whether it’s Mitsubishi or Siemens).
3. A Simple First Architecture is Okay
SQLite as a first version wasn’t a mistake. The mistake would be to think that a temporary scheme must forever remain the main one.
Sometimes it’s right to first collect data in a simple way and then calmly move the history to a more suitable storage rather than spending months developing the perfect architecture.
4. A Queue Doesn’t Save Itself
If the collector writes to the same database from which the exporter is trying to unload and confirm rows, you need to think ahead about locks, races, and the growth of the pending queue.
If the queue accumulates and starts interfering with writing — that’s no longer a “technical detail,” but a signal that it’s time to separate the architecture.
Conclusion
In this work, there was no beautiful moment of “press the button — and the signal collection for ML started.” But there was a more useful result: data began to transform from a set of disparate pieces into a manageable system and, most importantly, an UNDERSTANDABLE system.
Documentation became a catalog. A safe read-only approach to the smoke test was established for the PLC. The first version of the collector provided real data, and then this data was carefully moved from SQLite to ClickHouse.
For an ML engineer at a factory, this seems to be one of the most important parts of the job. It’s not just about training a model, but ensuring that the data reaches that model carefully, verifiably, and without unnecessary risk to the equipment.