InsightFab
Knowledge Base/DOE Sample Size Calculation: Effect Size and Statistical Power
DOE6 min read

DOE Sample Size Calculation: Effect Size and Statistical Power

This article addresses the common workplace challenge where perceived differences are not statistically significant, illustrating how inadequate sample size in experimental design can lead to insufficient statistical power, failing to detect true effects. It aims to deepen understanding of experimental design and statistical testing, providing guidance for future problem-solving.

That day, DPMO soared to 6210, and the Section Chief was visibly upset

That afternoon, the production line suddenly reported an urgent situation, saying that a batch of new materials was put into production, and the DPMO for optical inspection of the produced wafers actually soared to 6210. The section chief was visibly upset on the spot, immediately called everyone for a meeting, and demanded we find the cause within a limited time. At that time, I happened to be working on an improvement project, wanting to evaluate the impact of several new process parameters on yield. After my designed experiment was run, and the data came out, the statistical analysis said "no significant difference." I thought to myself, how is that possible? It clearly looked different to the naked eye! At that time, I muttered to myself, could it be that the sample size was too small, obscuring the real difference?

Frankly, it's about whether your data is powerful enough

Have you also encountered this situation? You clearly feel there's an effect, but the statistical report says it's not significant. To be honest, this is likely because your experimental design "missed the point," most commonly due to insufficient sample size. The purpose of doing DOE (Design of Experiments) is to use the fewest experiments to identify key factors affecting product quality. But if your sample size is too small, and the data volume is not large enough, even if there is indeed an effect, your statistical test might fail to detect it due to "lack of power." This is what's known as "insufficient statistical power."

In other words, sample size calculation is to ensure that your experimental data has enough "convincing power" to prove your hypothesis. Otherwise, you spend a lot of time and resources doing experiments, only for it to be a wasted effort – who can stand that?

In practice, you should judge this way

So, how many samples are enough? This involves several key factors:

  1. The "Effect Size" you want: How big of a difference do you wish to detect? For example, you hope the new process can increase the yield from 99.5% to 99.8%. This 0.3% difference is your "effect size." The smaller the effect size, the more samples you need.
  2. Your Risk Tolerance (Alpha and Beta):
* Significance Level (Alpha, α): We usually set it at 0.05 (5%), which means you are willing to accept a 5% chance of incorrectly rejecting a true null hypothesis (claiming an ineffective process is effective).

* Statistical Power (Power, 1-β): Usually set at 0.8 or 0.9 (80% or 90%), meaning you want an 80% or 90% chance that when an effect truly exists, your experiment can successfully detect it. In other words, you only have a 10% or 20% chance of missing a real effect. This Beta (β) is the probability of committing a Type II error, which is "claiming an effective effect is ineffective."

Frankly, the setting of these parameters directly affects your sample size. If you want to detect very small differences and also want very high power, then the sample size will certainly explode. In my experience, if process fluctuation is large, or if you want to detect very small improvements, the power is usually raised to 0.9.

The most common pitfall: haphazardly cutting corners to save money

The most common pitfall I've encountered is drastically reducing the sample size to save money or meet deadlines. As a result, the Cpk values obtained are 1.08 and 1.12, which are numerically different, but statistically, there's no significant difference, and the boss, of course, won't buy it. At this point, you have to go back and redo the experiment, wasting time and materials for nothing.

There's another type, which is not calculating at all, and guessing the sample size by feel. Usually, they just follow the "custom" of their predecessors, but the objectives and variability of each experiment are different, and blindly copying can easily lead to errors. For example, in a previous experiment on the impact of a parameter on line width, you might only need 30 wafers for the sample size. But now you need to evaluate the impact of a new material on wafer warpage, where variability might be much larger. If you still use 30 wafers, you might very likely fail to detect the difference.

One thing you can do today

Next time before doing DOE, spend 10 minutes to run a sample size calculation using software.

Want to try it yourself?

Every tool mentioned in this article is available on InsightFab — just upload a CSV to analyze.

Go to Tools