InsightFab
Knowledge Base/Quality Warning System: Designing Early Warning Indicators
Quality Assurance6 min read

Quality Warning System: Designing Early Warning Indicators

This article details a factory's costly experience where a reactive alarm system led to significant losses due to delayed quality issue detection. It proposes designing a proactive "early warning" quality system that utilizes leading indicators to identify potential problems when they are still minor, thus preventing major disruptions and financial impact.

That Day, When the Cpk Report Came Out, the Entire Room Was Silent for Three Seconds

I remember several years ago, we had a batch of new products in our factory that ran into trouble right after introduction. That day, the production line machines blared alarms, and the entire MES system was glowing red. The supervisor rushed over, his face ashen, asking, "What on earth happened?" Upon investigation, we found that the Cpk of a critical parameter had plummeted to 1.08, collapsing from its original 1.33. The key point was that this parameter had always been very stable; no one expected anything to go wrong. That time, just finding the root cause and clarifying the scope of impact took nearly two days, halting the entire production line. The losses were no laughing matter.

Where Was the Problem?

To put it plainly, our alarm system back then was simply "locking the stable door after the horse has bolted." By the time the Cpk was bad enough to trigger an alarm, the disaster had already occurred. It's like driving a car; what good is it if the "overheating" warning only shows up when the engine is already smoking? A quality warning system needs "early warning," not "post-event announcement." In simpler terms, it's about letting you see the warning signs when the problem is still small, triggering an alert before the Cpk drops to 1.08.

Therefore, the key is that we need to design "leading indicators" that can detect signals before quality deteriorates. These indicators won't directly tell you that the product is broken, but rather that it's "about to break."

How to Do It in Practice?

The simplest and most common approach is to observe the "trend changes" of process parameters. For example, if the tolerance for a certain etching time is +/- 5 seconds, you might set an alarm only when the value exceeds the tolerance. But early warning doesn't work that way.

  1. Observe Changes in Standard Deviation (StDev):
When monitoring quality, we typically check if the mean value has drifted. However, many times, the mean remains unchanged while the standard deviation slowly expands. Suppose the normal range for the StDev of a critical process parameter is 0.5. If I find that this StDev value falls between 0.65 and 0.70 for ten consecutive batches, even though it hasn't exceeded the tolerance, it's already a warning sign. This indicates that the process stability is deteriorating, and the variability of product quality is increasing.

Therefore, the key is that in addition to monitoring the mean, the trend changes in standard deviation should also be included as an early warning indicator.

  1. Establish Criteria for "Consecutive Anomaly Points":
Suppose you've set a warning interval; for instance, a Cpk between 1.10 and 1.20 is a "yellow light zone." You cannot trigger an alarm just because a single batch falls into the yellow light zone, as that might just be random fluctuation. However, if three, or even five, consecutive batches fall into this yellow light zone, then there's a problem!

For example, we found that DPMO (Defects Per Million Opportunities) is usually around 2000, but one day it suddenly surged to 6210. This is, of course, a red light. But if the DPMO slowly climbed from 2000 to 2500, 3000, 3500 for five consecutive batches, even though it hasn't reached the red light, this is definitely a potential risk.

Therefore, the key is to define how many "consecutive points" of data falling into a certain warning interval should trigger an alarm, rather than relying solely on single data points.

The Most Common Pitfalls

To be honest, many people make early warning systems too sensitive, leading to continuous "cry wolf" scenarios. Alarms keep sounding, but every time they're investigated, nothing is wrong. Over time, everyone becomes desensitized. Eventually, when a real problem occurs, no one pays attention. We encountered this previously: engineers set the warning interval too narrowly, triggering an alarm for even slight fluctuations, causing everyone to run around chasing alarms daily, only to find nothing each time. Later, we reviewed the data and realized that most of the "early warnings" were simply normal fluctuations, merely because the interval was set too rigidly.

Another pitfall is setting up a myriad of early warning indicators that no one actually monitors or knows how to act upon if they do. If you configure a dozen or more warning indicators, each with its own trigger conditions, and then hundreds of warning messages are generated daily, who has the time to process them all? The truth is, more indicators are not necessarily better; the focus should be on selecting truly critical and meaningful ones.

One Thing You Can Do Today

Go back and check your current quality alarm system to see if it incorporates "standard deviation trends" into its monitoring.

Want to try it yourself?

Every tool mentioned in this article is available on InsightFab — just upload a CSV to analyze.

Go to Tools