The more data I collected from the machine, the more important the question became not “what else can I add?”, but rather “what of this actually makes sense to show the model?”

In one of my experiments, I had 7230 PLC signals and 666,208 time slices from the machine.

At first glance, it seems like a huge dataset.

But the PLC stores not only the physical parameters of the process. It has everything: internal program states, errors, service flags, intermediate calculations, HMI signals, duplicates, signs of already occurred defects, and a huge number of registers that haven’t changed even once during the entire study period.

If you just feed all this to the ML, the model will definitely find something.

The question is only — what exactly will it find and how will we understand it later?

First Surprise: Most PLC Signals Are Almost Static

In one of the first large runs, I analyzed 6299 signals (we had several collector options).

Out of these, 5650 turned out to be static during the study period.

That is, only 649 actually changed.

This is quite sobering.

You look at the PLC map and see thousands of addresses:

X...
Y...
M...
D...
W...

Then you open the time series and realize that most of this wealth looks something like this:

0
0
0
0
0
0
0
...

For a specific training dataset, such features are useless.

But there is one caveat.

A static signal cannot be automatically declared garbage at the machine level.

Some error bit could have been zero for two months simply because no error occurred, and that’s normal.

For the current ML experiment, I can exclude it.

But I cannot exclude it from the PLC catalog. That’s the paradox.

And now the main rule:

The ML dataset and the machine signal map are not the same.

The dataset can be aggressively cleaned, but the equipment map cannot.

Second Surprise: The Model Loves to Peek at the Answer (Experienced ML Folks Will Understand)

The next problem turned out to be more interesting.

Let’s say we want to predict defects.

Somewhere in the PLC, there is already a chain that looks something like this:

process
  ↓
a problem occurs
  ↓
PLC detects it
  ↓
sets a reject / alarm / defect bit
  ↓
the product is rejected

If you take a signal from the end of this chain and add it to the features, the ML will be happy.

The metrics will be fantastic.

But there’s no prediction here.

The model simply learned to read the answer from the PLC.

That’s why I separately removed features related to the already known reject chain: ready defect features, diagnostic flags, and signals that appear almost simultaneously with the target event.

Because this is not ML; it’s just the logic of the PLC operation.

In one of the runs, there were 321 such potentially dangerous features.

After filtering, I was left with 1322 dynamic non-leakage features.

And only after that did the experiment become interesting.

Because now the model was forced to look for not:

PLC has already reported "reject"

but something like:

a few seconds before the reject
in this part of the process
a specific combination of signals begins to change

Now that looks like ML.

But Here Comes Another Trap

After cleaning, I started looking at what machine states were actually present in the data.

Because the PLC time series is not always production.

It’s a mix of:

  • normal operation;
  • stops;
  • startups;
  • transitional states;
  • retooling;
  • technical operations.

For a person near the equipment, these are obvious things.

For ML, they are just different combinations of several thousand numbers.

I decided to look at the data without any defect information and ran unsupervised analysis.

PCA + KMeans confidently separated the time series into several machine states.

In the final experiment, there were three major states:

state 0 — 609 978 snapshots — 91.57%
state 1 — 26 342 snapshots — 3.95%
state 2 — 29 810 snapshots — 4.48%

The largest cluster resembled ordinary stable production.

The others resembled stops and transitional states.

And then an unpleasant thought emerged.

What If the Model Is Not Predicting Defects?

Let’s imagine a situation.

Before a defect, the machine often slows down or transitions into some other state.

The model sees this and shows a good ROC-AUC.

We rejoice:

We found a precursor to defects!

But in reality, it could have learned something entirely different:

the machine is running
        ↓
something happens
        ↓
the machine stops

That is, instead of a physical cause of defects, we just got a very good machine mode detector.

For production ML, this is a dangerous trap.

Especially when there are thousands of features.

So, we always need to exclude transitional states of the equipment from ML experiments; otherwise, we will find completely different patterns.

Another Trap — PLC Duplicates

At the same time, signals started popping up that looked like two independent good predictors.

Then it turned out that they were almost synchronous.

For example, in one of the experiments, pairs:

M5136 ↔ M5402
M5137 ↔ M5403

turned out to be practically duplicate representations of one event in the PLC.

If you don’t check such things, it leads to a funny situation.

The model says:

The most important features are M5136 and M5402.

You think:

Great, two independent reasons confirm each other.

But physically, it could turn out to be the same state, just recorded in two different places in the PLC program.

So, correlation between signals doesn’t prove anything by itself.

You need to constantly return from ML back to the logic of the machine:

feature
  ↓
PLC address
  ↓
what this signal represents
  ↓
which equipment is behind it
  ↓
what is physically happening at that moment

Without this last step, it’s not ML; it’s machine learning on mysterious register numbers. Or just pretty numbers for a pretty report and nothing more.

In the End, I Stopped Viewing PLC as a Regular Feature Table

This is the main takeaway from this stage.

In regular tabular ML, you can represent data like this:

feature_1
feature_2
feature_3
...
target

After all, this is how we learned in training notebooks and on Kaggle.

You can’t work with PLC like that.

For example, you have:

D5708
D7808
M5965
X161
...

Why not just make X features from them and send it to the model?

Because behind each such column is a piece of a program or a physical process.

One signal means the position of a mechanism.

Another — a calculated parameter.

The third — a command.

The fourth — confirmation of command execution.

The fifth — already prepared fault diagnosis.

The sixth — the operator interface state.

And if you mix all this indiscriminately, the model easily gets an excellent metric in a completely wrong way.

How My PLC Preparation Pipeline for ML Looks Now

In simplified form, I came to about this sequence:

raw PLC signals
        ↓
time-series quality checks
        ↓
removal of static features
        ↓
identification of duplicates and service signals
        ↓
removal of the reject chain / leakage
        ↓
identification of machine operating modes
        ↓
model validation during production
        ↓
identification of temporal precursors
        ↓
physical interpretation of the identified signals

And only after this can you take the found patterns seriously.

The most interesting thing is that the actual ML takes very little time (surprising, right?).

Most of the work is data work: understanding what exactly the model saw.

Main Insight

I used to think that a large number of PLC signals was a rich base for ML.

Now I look at it differently.

7230 signals are not 7230 useful features.

They are 7230 candidates, among which are:

  • good physical parameters;
  • internal PLC states;
  • duplicates;
  • constants;
  • diagnostics;
  • leakage;
  • rare events;
  • machine modes;
  • and somewhere among all this — real precursors of the process.

The job of an ML engineer is precisely to separate useful signals from useless ones.

Offline ML Model ≠ ML Model on Online Data

You can run historical datasets as much as you want offline, change windows, and look at metrics.

But for the experiment, the model needs to be placed next to the real machine and see:

what it will say when production is happening in real time?

And I will write about this next week.