That morning shift, the machine had another unexpected shutdown
Hey, do you remember the scene last time when our veteran senior engineer stormed into the office, face dark, slamming the table and yelling, "Which genius messed with the PM schedule?! The machine just had another unexpected shutdown on me!"? I tell you, that day was a Monday morning, and before I even finished my large iced milk tea, I heard alarms constantly blaring in the Fab. A bunch of people rushed around, only to find that one of the critical machines had suddenly crashed, bringing the entire production line to a halt. Later, after much investigation, everyone realized, damn, some consumables hadn't even reached their end-of-life but inexplicably failed; some things were PM'd just the day before, only to malfunction again today. At that moment, I wondered, do these machine failures have any patterns? Do we have to treat it like a lottery guess every time?
Where's the problem? A machine doesn't just break down because you want it to.
To put it simply, machine failures generally fall into three modes, which engineers often refer to as the "bathtub curve."
- Early Failure: This is like buying a new car and having it break down shortly after driving it off the lot. It's usually caused by poor design, manufacturing defects, improper installation, or inadequate PM execution. New machines just coming online or new parts just replaced are particularly prone to this. If the Cpk is very low, for example, only 0.8, it indicates that your product or process variability is too high, making it very susceptible to issues.
- Random Failure: This type is truly reliant on luck, with no discernible pattern. It's like a small rock suddenly hitting your windshield on the road—unpreventable. It typically occurs after the machine has been operating stably for some time, possibly due to external environmental factors, voltage instability, or microscopic material defects. This type can have a DPMO as high as 6210 ppm, indicating issues with your process stability that need to be addressed at a system level.
- Wear-out Failure: This is the easiest to understand: things naturally break down after prolonged use. Worn bearings, burnt-out light bulbs, aging cables—these are all predictable. It usually occurs in the later stages of a machine's service life.
So here's the point: if you don't even understand the failure modes, how can you prevent them? How can you schedule PMs?
How to actually do it? Let data speak.
To determine which mode it is, frankly, it's about "looking at the data."
- Analyze failure time points: If most failures are concentrated in the first few days after "parts have just been replaced" or "the machine has just come online," then it's most likely early failure. You'll need to check the quality of parts from suppliers and ensure that installation SOPs are strictly followed.
- Observe failure distribution: If failure times are widely dispersed, with no clear peak, and occur during the middle stage of the machine's life, then it's likely random failure. In this scenario, your focus should be on improving process stability, environmental control, and even considering the implementation of preventive maintenance.
- Monitor part lifespan: If you find that a certain part always fails after a fixed number of operating hours, and there's a clear trend, then congratulations, this is the easiest wear-out failure to manage. You just need to set the PM cycle based on the data, before the part's expected lifespan, which can significantly reduce unexpected shutdowns. For example, we had a motor bearing that historically failed after an average of 5000 hours of use, so we set the PM cycle at 4500 hours, directly preventing many shutdowns.
In other words, you need to lay out past failure records and not just treat them as reports to be submitted.
The most common pitfall: laziness in classification, leading to worse repairs
To be honest, I've fallen into this trap before. When I first started as an equipment engineer, every time a machine broke down, I'd rush to fix it, and then it was considered done. I simply didn't have the time, nor did I think about analyzing which failure mode it was. The result was that early failure issues kept recurring because I only replaced parts without addressing the root cause; wear-out parts were replaced only after they completely failed, leading to production line bottlenecks every time; random failures were even worse, as they couldn't be predicted, only reactively dealt with. Once, one of our pumps frequently leaked, and each time we just replaced the O-ring, but it still leaked after three replacements. We later discovered that the root cause wasn't the O-ring itself, but excessive assembly tolerance in the pump, causing uneven force on the O-ring—a classic early failure where merely replacing consumables was useless!
Frankly, many times it's not that we don't know we should classify, but rather that we feel we "don't have the luxury of time." But the more this happens, the more you'll fall into a vicious cycle of "endless repairs."
One thing you can do today
Take out the last three failure records for your machines and try to determine if they were early, random, or wear-out failures.