That day, when the CPK report came out, the whole room fell silent for three seconds, and then I decided to talk about reliability
I remember several years ago, shortly after we shipped a new product, the customer reported sporadic board failures. Though the quantity wasn't large, the quality department was deeply troubled. At that time, everyone was wondering, all our products had passed every test before leaving the factory, and Cpk was maintained above 1.08; by all accounts, such issues shouldn't have occurred. After investigating for a long time, we discovered that certain environmental stresses only triggered the product's weaknesses after shipment. At this point, you realize that reliability testing isn't just about 'doing it'; it's about 'how it's done' to uncover those potential issues.
Where exactly was the problem? 'Not broken' before product shipment doesn't mean 'won't break'!
To be frank, many times we treat pre-shipment product testing as the entirety of quality. But this is like buying a new car; just because it's fine when you drive it out of the showroom doesn't mean it won't break down on the highway after five years. Electronic products are the same. Conventional tests can, at most, only ensure your product is good 'at this very moment,' but whether it can withstand temperature changes, vibrations, and humidity impacts over the next few years — that's the problem reliability testing aims to solve.
Therefore, HALT, HASS, and ESS, which we often hear about, are all about 'early detection, early treatment' of potential product issues. Frankly, the purpose of these tests is to try and break the product while it's still in your hands. You heard that right, to break it, and to break it earlier and faster than its normal lifespan.
So, what exactly are the differences between HALT, HASS, and ESS?
Simply put, they are like three different levels of 'product torture techniques'.
- HALT (Highly Accelerated Life Test)
- HASS (Highly Accelerated Stress Screen)
- ESS (Environmental Stress Screening)
The most common pitfall: Using HALT data as HASS thresholds
The most absurd thing I've encountered is some newcomers directly using the 'destruction boundary' determined by HALT as the screening threshold for HASS. Do you know? HALT is meant to break the product, while HASS is meant to screen out defective products. If you break all the products, what are you screening for? Doing this will only make your product yield zero, and then your boss will ask if you've misunderstood something. The correct approach is for HALT to identify the boundaries, and then HASS stress levels should be set within a range that can 'screen out defective products without damaging normal products'. This is typically around 60%~80% of the HALT destruction boundary.
One thing you can do today
Go ask if your products undergo HALT.