Scenario
"The SPC chart for this batch of wafer film thickness looks all green, so why has the yield dropped by another 0.5% compared to last month?" You stare at the control chart on the screen, every point obediently staying within the control limits, but you can't shake the feeling that something is off. During the night shift handover, the senior engineer didn't say anything unusual. Am I just overthinking? This situation of "data not exceeding limits, but feeling wrong" is certainly familiar in Taiwan's manufacturing sites.
Explained Simply
When we look at SPC control charts in the factory, the most intuitive thing is to see if any points run outside the Control Limits. This is like a traffic light; once a point is out of bounds, an alarm sounds, indicating that a Special Cause Variation has occurred and needs immediate attention. However, sometimes the process hasn't "run a red light" yet but has started to "swerve". In such cases, merely looking at the control limits is not enough.
This is where Run Rules come into play. They don't look at whether a single point is out of bounds, but rather whether continuous data points show a certain "unusual pattern". You can imagine them as a more sensitive radar that not only detects individual targets but also analyzes their flight paths and behavioral patterns.
The purpose of run rules is to identify latent, slowly accumulating process anomalies—early signals of "special Cause Variation"—from seemingly normal data.
Let's break down a few core concepts:
- Control Limits: The most basic judgment criterion, like a traffic light; an alarm sounds if a single point is out of bounds. They are set based on the process's Common Cause Variation, and exceeding them indicates a special event has occurred.
- Run Rules: More advanced judgment criteria, like radar detection, which look not just at individual points but also at the "behavioral patterns" of the data. For example, several consecutive points moving in the same direction, or several consecutive points on the same side of the center line.
- Special Cause Variation: Refers to non-random, assignable variation in a process, such as equipment malfunction, operator error, abnormal raw material batches, etc. Run rules are designed to detect these early.
- False Alarm Rate / Alpha Risk: This is a very important concept. The more sensitive we set an alarm, the easier it is to detect real problems, but also the more likely it is to generate "false alarms". A high false alarm rate will exhaust on-site engineers, eventually leading to a "boy who cried wolf" numbness towards alarms, and potentially missing genuine warnings. Therefore, finding a balance is crucial.
Practical Application
In practical applications of run rules, we typically refer to standard judgment criteria, such as the eight rules of the Western Electric Company. However, not all rules should be activated simultaneously; selection must be based on process characteristics and risks.
Below are some common and practical run rules:
| Rule Name | Judgment Condition (Example) | Potential Cause | False Alarm Rate Impact (Used Alone) | Practical Recommendation |
|---|---|---|---|---|
| Out of Control | 1 point outside 3-sigma control limits | Dramatic process shift, measurement error, parameter misconfiguration | 0.27% (most basic) | This is the most basic alarm and requires immediate investigation. |
| Run Above/Below Center Line | 9 consecutive points on the same side of the center line | Process mean shift, equipment calibration drift, raw material batch difference | 0.002% (lower) | This is a frequently overlooked but important warning sign, indicating the process is beginning to drift. |
| Trend | 6 consecutive points steadily increasing or decreasing | Tool wear, reactant depletion, slow environmental temperature change | 0.001% (even lower) | A good time for preventive maintenance; allows for predictable adjustments. |
| Alternating | 14 consecutive points alternating up and down | Measurement system issues, multiple processes interacting, over-adjustment | 0.0001% (very low) | Check measurement instrument stability or process control logic. |
Practical Operation Recommendations:
- Start gradually, don't overdo it: When first introducing run rules, do not enable all of them at once. While enabling more rules makes it easier to catch problems, the False Alarm Rate (Alpha Risk) will also accumulate. Too many alarms will exhaust on-site engineers, leading to "alarm fatigue".
- Prioritize 'Run Above/Below Center Line' and 'Trend': These two rules most frequently reflect slow process shifts or drifts and are important bases for preventive maintenance and adjustment. Their false alarm rates are relatively low, and they provide valuable early warnings.
- Regular review and adjustment: Processes evolve, and the settings for run rules should be reviewed regularly. If a rule frequently generates false alarms, its applicability may need to be re-evaluated or its parameters adjusted; if a rule has never triggered but a process problem has occurred, more sensitive rules may need to be enabled.
- Combine with process knowledge: No statistical rule should be detached from actual process knowledge. When an alarm is triggered, combining on-site experience with process principles for analysis is essential to find the true root cause.
How InsightFab Helps
InsightFab has built-in automatic detection functions for various run rules, not only helping you monitor process data in real-time but also visually presenting the patterns that trigger alarms, allowing you to identify problems at a glance. More importantly, it helps you adjust the sensitivity of the rules to effectively manage the false alarm rate, ensuring alarms are both sensitive and reliable.
Key Takeaway
"The 'behavioral pattern' of data is more important than individual point values. Learning to understand it allows you to act proactively before a problem escalates into a major issue."