The Day My Boss Asked: "Are You Sure There's a Difference?" I Was Speechless
That afternoon, a batch of wafers suddenly went awry on the production line; a critical process parameter had drifted significantly. We immediately investigated and found an anomaly in a gas flowmeter on a specific machine. Once adjustments were finally made, we quickly ran a few experimental wafers, aiming to prove that the adjustments had indeed brought improvement. The data came back: CPK had increased from the original 1.08 to 1.25. I excitedly rushed to report to my boss: "Boss, look! The CPK is up, proving our adjustments were effective!" My boss squinted at the report and calmly asked, "Are you sure there's a difference? Could it just be a coincidence?" I was instantly speechless, my inner monologue screaming: "What else could it be? The data is right there!"
What Was the Problem? It's Not Enough for Data to Just Show a Difference
Simply put, what my boss was asking was, "Is your conclusion robust enough?" As engineers, we often see a slight difference in data and rush to draw conclusions. But have you ever considered whether this difference might just be random fluctuation during the experiment? Or perhaps your sample size was simply insufficient, making the conclusion "not robust enough" even if a difference was observed? This is where Statistical Power comes into play. In layman's terms, it's your ability to "successfully detect a true difference."
In other words, if a process truly has improved, can your experiment correctly tell you "there's an improvement"? If your statistical power is too low, it's like driving in thick fog: even if there's a traffic light ahead, you might miss it because you can't see clearly. In a factory setting, if power is too low, you might spend a lot of time and money on improvements, only to find that due to poor experimental design, you cannot prove its effectiveness, making all your efforts for naught.
How to Apply It in Practice? Let the Numbers Speak
Frankly speaking, you can't calculate power after an experiment is done; that's too late. It should be considered during the "experiment design" phase. Its most common application is to determine your required "sample size."
For example, imagine you want to compare two new coating formulations, aiming to reduce the thin film thickness variation on wafers. You anticipate formulation B will reduce the standard deviation by 10% compared to formulation A. At this point, you need to set:
- Significance Level (Alpha): Typically set at 0.05, meaning you're willing to accept a 5% risk of incorrectly concluding there's a difference (when there isn't one).
- Desired Power: Usually set at 0.8 or 0.9. This means you want an 80% or 90% chance of correctly detecting this 10% difference.
- Expected Effect Size: As mentioned, formulation B reduces the standard deviation by 10% compared to formulation A.
With these three values, you can use statistical software (JMP and Minitab both have this function) to calculate how many wafers you need for the experiment to achieve your desired power. For instance, if it calculates 30 wafers are needed, then you should conscientiously run 30 wafers. If you only run 5, even if you see a tiny difference, when your boss asks, "Are you sure there's a difference?", you will genuinely feel uncertain.
The Most Common Pitfall: Saving Money Only to Create Problems
I've fallen into this trap before. Once, we needed to validate a new material that theoretically could reduce DPMO from 6210 to 5000. But the material was expensive, and the experimental cost was high, so my boss said, "Just run 10 wafers to start; if there's a trend, that's fine." After 10 wafers, the DPMO indeed dropped a little, but it was not statistically significant. My boss shook his head, concluded the new material was useless, and dismissed it.
Later, I recalculated using power analysis and found that to detect a DPMO reduction from 6210 to 5000 with 80% power, at least 50 wafers were needed! We only ran 10 wafers; it was akin to searching for a needle in a dark cave with a flashlight—of course, we couldn't find it. To be honest, it wasn't that the new material was useless; it was that our experimental design lacked power, and we needlessly squandered the opportunity to validate the new material. This lesson taught me that spending more time on experimental design absolutely pays off more than doing fruitless work later.
One Thing You Can Do Today
Before your next experiment, consider how "certain" you need to be before drawing conclusions.