That Day, a Customer Asked How Long My Product Could Live, and I Was Stunned.
I still remember a few years ago, when I first transitioned from process engineering to reliability and had my first meeting with a customer. The customer, a foreigner, directly asked me: "So, your chip, how long can it last in my device?" I was completely dumbfounded at that moment, my mind filled only with numbers like Cpk 1.08 and DPMO 6210, but I couldn't articulate a "lifespan." After that meeting, I realized that reliability isn't just about yield; it's about how long your product can "live" in the customer's hands. Frankly, what customers want isn't how high your yield is, but how long your product can operate stably without causing complaints in the market.
What's the Problem? Failure Rate, MTBF, and Reliability Function – Do You Understand Them?
Speaking of reliability, many people immediately think of "Failure Rate," and that's correct. Failure rate is the probability of a product failing within a specific period. For instance, if you buy a batch of new light bulbs, and within the first hour, 1 out of 1000 bulbs fails, then its failure rate is 1/1000 = 0.001 failures/hour. That's intuitive, right?
But just looking at the failure rate isn't enough. Often, you'll hear about MTBF, or "Mean Time Between Failures." MTBF is essentially the reciprocal of the failure rate. If your failure rate is 0.001 failures/hour, then your MTBF is 1/0.001 = 1000 hours. So, the key point is: the higher the MTBF value, the more durable the product, and the lower the frequency of failures.
So, what is the Reliability Function? This is interesting. It describes the probability that a product can still operate normally "before a specific point in time." For example, if a product's reliability function at 1000 hours is 80%, it means that 80% of the products can live beyond 1000 hours. Therefore, the failure rate tells you "how often it fails," MTBF tells you "on average how long it takes to fail once," and the reliability function directly tells you "what the probability of living for how long is." These three are essentially two sides of the same coin, interconnected.
How Is It Done in Practice? Understanding the Bathtub Curve Is Key.
Frankly, to understand these concepts, the simplest way is to look at the "Bathtub Curve." This curve divides product life into three stages:
- Infant Mortality Period: When products are new from the factory, some batches may fail quickly due to manufacturing defects or poor design. At this point, the failure rate is high but decreases rapidly with time. It's like a newborn baby; the probability of illness is high, but once they pass the critical period, they're fine. So the key is that the failures in this stage are usually screened out by "Burn-in."
- Useful Life Period: The product enters a stable phase, where the failure rate becomes steady and low. The failure rate at this point is the product's "intrinsic" failure rate, and it's the stage we most commonly use to calculate MTBF. For example, our chips can achieve a DPMO below 100 under normal operation, which reflects the performance during this stage.
- Wear-out Period: After a long period of use, components begin to age and wear, and the failure rate gradually increases. This is like humans getting older, where body functions decline. So the key is that failures in this stage are primarily influenced by factors such as material aging and fatigue.
Therefore, when a customer asks how long your product can live, you cannot just give one number. You need to know which stage your product is currently in and what the failure rate is for that particular stage.
The Most Common Pitfall: Using the Wrong Failure Rate to Calculate MTBF
The most ridiculous thing I've encountered is someone applying the failure rate from the infant mortality period to the MTBF of the useful life period. The resulting MTBF was incredibly high, which, of course, delighted the customer. However, it wasn't long before a flood of customer complaints came in due to the excessively high early failure rate. Frankly, you cannot include the defective products screened out during the burn-in phase and then claim your product is very durable.
Another pitfall is providing only an MTBF value without specifying the environmental conditions under which it was measured. The MTBF measured at room temperature is absolutely miles apart from the MTBF measured in a high-humidity environment at 85 degrees Celsius. This is like saying a car is fuel-efficient without specifying whether it's driven in the city or on the highway.
One Thing You Can Do Today
Go back and check your product specification sheet to see if the failure rate is clearly defined for which life stage and under what environmental conditions it was measured.