Because an ML model will find a good pattern in the data (if there is one), but how do we explain that in the real world later on?

In previous articles in this series, I’ve already written about how I dealt with PLC signals, organized the data, stored it in ClickHouse, and displayed it in Grafana. That was the foundation. Now this foundation is gradually turning into a more applied task: understanding why some products are rejected and whether we can notice signs of a problem in advance.

Let me clarify right away: I won’t disclose the exact names of the equipment, manufacturer, line, product, or organization. What matters here is not the machine's passport, but the approach.

Why Auditing Defects is Harder Than Finding the Right Bit

From the outside, it seems like everything should be simple: there’s a product, there’s a defect, there’s a signal in the PLC. We take this signal as the target variable, train the model, and rejoice.

In practice, it turned out to be a bit different.

The first question we ran into was: what types of defects actually exist and how do they physically manifest? It’s one thing if a product is clearly rejected by the control system. It’s another if the operator turned something off on the panel, the line is in setup mode, or part of the detection logic is temporarily ignored.

And this is where it gets interesting. If we don’t sort out these nuances, the ML model can easily turn into a beautiful but useless thing. It will predict not defects, but the specifics of how the equipment was set up on a particular day.

For example, we can take a signal that appears only after rejection. The model will show excellent accuracy on it. But the usefulness is minimal: it will learn to recognize an event that has already occurred. What we want is to understand what might lead to a defect in advance.

First, Agree on Terms

Before diving into the data, we broke down the audit into understandable steps.

We needed to establish:

  • what types of defects exist;
  • where they occur during the process;
  • who or what detects them;
  • what constitutes a single unit of production;
  • what mechanisms are involved in forming a specific defect;
  • what sensors, axes, drives, temperatures, pressures, and setpoints might be related to this location.

I agree it sounds boring, but it’s essential.

In a factory, there are many things that may look like ordinary columns in a table but have very different meanings in reality. One signal can be a cause. Another can be an early warning. A third can simply be context. A fourth can be the fact that a defect has already been found. A fifth can be a consequence of rejection. A sixth may not change at all over a long data collection period (i.e., it remains just a constant).

So, we had to mentally categorize all signals into groups:

  • possible causes;
  • early warnings;
  • technological context;
  • signals detecting an already occurred defect;
  • signals after rejection that cannot be used for prediction.

And this separation for industrial ML is sometimes more important than choosing between a conditional CatBoost and a neural network.

Look with Your Eyes, Not Just SQL Queries

For the audit, we needed to combine two realities.

The first is physical. What happens on the equipment: where the defect appears, in which zone, under what scenario, and what happened just before that.

The second is digital. What PLC signals can be extracted from ClickHouse later, what they are called, how frequently they are recorded, and how to tie them to real events.

At some point, it became clear that simply observing the process was no longer enough. We needed not just general impressions but precisely marked moments of rejection.

The most practical option turned out to be very simple: place a camera at the rejection point and then reconstruct the time and quantity of rejected products from the video.

But there’s an important detail without which everything falls apart: the video needs to be synchronized with real time. For example, once capturing the screen of a phone with the time in hours and seconds or verbally stating the exact time. Then each rejection can be tied to seconds in ClickHouse.

To put it even simpler - do manual labeling and compare it with the labeling from the PLC.

For the first exploratory analysis, we aimed for 20-30 reliably marked events. This doesn’t mean that 20-30 cases will prove any hypothesis. There are nuances:

  • if all cases are the same — that’s already strong material;
  • if there are several different types of defects — there will be few examples of each;
  • some rejections may occur for one reason, while others for another;
  • externally similar events may be triggered by different scenarios within the equipment.

So, the goal was not to prove a pattern at any cost but to collect an honest set of events and see what’s actually there.

What a Dataset for Such a Task Looks Like

After decoding the video, the dataset should be as grounded as possible. Not an abstract table with a million rows, but a list of events:

event_001 10:02:xx 2 products
event_002 10:05:xx 2 products
event_003 10:13:xx 1 product
...

Next, for each event, we can automatically extract a data window from ClickHouse. For example, 60 seconds before the rejection and 10 seconds after.

Why like this? Because we’re interested not only in the moment of rejection. It’s more important for us to understand what happened before it:

  • did the speed change;
  • were there any discrepancies;
  • did the loads increase;
  • did temperatures or pressures spike;
  • did setpoints change;
  • were there any accidents, blockages, or non-standard modes;
  • what was the recipe or type of product;
  • what was the operator doing on the panel.

And control is essential: similar time windows without defects. Otherwise, we might find a signal that occurs before every normal cycle and mistakenly label it as the cause of the defect.

This is a typical trap. You can always find something in the data. The question is whether it relates to the defect.

Why We Need Grafana and ClickHouse

ClickHouse in this story is a storage solution where it’s convenient to store fast time data from PLCs. From it, we can extract windows around events, compare defects with normal operation, and test hypotheses based on signals.

Grafana is the eyes for people.

For an engineer, it’s convenient to quickly open graphs and see what happened around an event. For management, it’s even more important: there’s no need to look at raw registers, incomprehensible addresses, and tables with thousands of rows. You can show a timeline, signals, rejections, modes, and explain in human language: here’s the moment of the defect, here’s what happened before it, here’s what repeats.

The combination of Grafana and ClickHouse in industrial monitoring looks very successful: ClickHouse quickly delivers data, and Grafana presents it beautifully and understandably. And it’s not just pretty dashboards for the sake of dashboards. It’s a way to discuss product quality not at the level of feelings but at the level of facts.

In my previous article about monitoring PLCs, I already wrote that it’s not enough to just collect data. You need to be able to present it in a way that allows for decision-making. In defect auditing, this became even more apparent: when there are marked events, the graph turns into an investigation tool.

Documentation Still Catches Up

A separate part of the work is to understand what signals actually mean.

We already had an engineering map of the process: from material feeding and main technological nodes to control, defect removal, and further product transfer. It includes an I/O registry, PLC accidents, production counters, product model bits, and servo drive parameters.

This helps a lot, but it doesn’t completely solve the problem. Documentation can have discrepancies, empty descriptions, unclear scales, and non-obvious HMI parameters. Sometimes a signal exists, but it’s unclear in what units it arrives. Sometimes the name is clear, but it’s unclear when exactly the bit is set.

Therefore, a request was prepared for the supplier: not just to send the manual, but to fill in variable tables, specify addresses, data types, units of measurement, scaling, and export variables from the development environment.

Parameters that the operator changes from the panel are especially important: sizes, speeds, temperatures, setpoints, modes. Without them, you can only see the consequences but not understand the context.

For example, if the equipment was operating in a non-standard mode, the model needs to know this. Otherwise, it might conclude that the cause of the defect is some sensor, while in reality, the line was in setup mode or the operator temporarily changed an important parameter. Here’s an interesting hack: operating mode ≠ defect.

Where’s the Machine Learning Here

ML appears closer to the end.

The correct sequence looks like this:

  1. Understand the physical process.
  2. Sort out the types of defects.
  3. Find the point and method for marking events.
  4. Synchronize video, logs, and PLC data.
  5. Extract windows from ClickHouse.
  6. Compare defects with normal operation.
  7. Separate causes from effects.
  8. Only then build the model.

The model should have a clear target variable, unit of observation, historical window, warning horizon, and validation on new days, batches, and recipes.

If this is missing, the model may be formally accurate, but it will be of little use to production.

I think this is one of the main lessons of industrial ML. In a factory, you can’t just take a table, hit fit, and wait for magic. First, you need to understand what the rows and columns mean. Otherwise, the model will confidently answer a poorly posed question.

Conclusion

Auditing defects in manufacturing is a story about engineering preparation: collecting events, tying them to time, understanding signals, checking documentation, and only then looking for patterns.

We’re not trying to guess the cause from a single graph, nor are we building a model on unclear labeling (as is often the case). We’re gradually turning chaotic factory data into a system that can be worked with.

And when the data becomes understandable, machine learning stops being a buzzword and starts being an ordinary useful tool for assessing equipment quality.