InsightFab
Knowledge Base/Equipment Maintenance KPI Design: What Metrics Truly Matter
Equipment Engineering6 min read

Equipment Maintenance KPI Design: What Metrics Truly Matter

A common issue in equipment maintenance is tracking 'what was done' rather than 'how well it was done'. This article discusses how focusing on superficial KPIs like PM completion rates can lead to poor performance, such as low CPK, and highlights the importance of shifting KPIs to measure the actual effectiveness and output of maintenance activities to avoid resource waste and escalating problems.

That Day, the CPK Report Came Out, and the Room Was Silent for Three Seconds

I still remember several years ago, our machines suddenly started showing a lot of strange defects, and the output rate kept plummeting. The factory manager was so furious he slammed the report on the table during the meeting, pointing at the CPK figure of 1.08 and demanding, "What exactly are you maintenance guys doing? You're doing maintenance every day, how could it turn out like this?" Everyone was stunned at the time, because the reports showed that all equipment's preventive maintenance (PM) had been done on time, and man-hours targets were also met. As a result, the entire team was harshly criticized and had to spend several weeks getting to the bottom of the issue.

Isn't it strange? Maintenance was clearly performed, so why were the results so terrible?

Where the Problem Lies

To put it bluntly, many times when we design equipment maintenance KPIs, we only look at "what was done," rather than "how well it was done." You tick off all 100 points on the PM checklist, fill in all the man-hours, and when the supervisor sees the report, "Wow, PM completion rate 100%! Excellent!" But in reality? Machines that were destined to fail still failed, and those that were drifting still drifted. This is like going to the gym every day to check in, but always just playing on your phone on the treadmill, and then complaining that your body hasn't changed.

Frankly speaking, if your KPIs are merely "PM completion rate" or "PM man-hours," you are essentially encouraging superficial work. Of course, you can tick off all items, but were the things that needed cleaning actually cleaned thoroughly? Were things that needed replacing actually replaced? Were instruments properly calibrated? These are what truly affect machine performance.

What to Actually Do

So, what's the key? You need to shift your KPIs from "input" to "output," or at least, link input to output.

  1. Machine Stability After PM: This is the most direct approach. After PM is completed, are the machine's Key Process Parameters (KPP) stable within the allowed range for a subsequent period? For example, you can look at the Out of Control (OOC) event rate after PM. If the machine goes OOC two days after PM, then there's definitely a problem with that PM. We previously introduced a metric requiring the OOC event rate for that machine to be < 0.1 within 7 days post-PM.
  2. Defect Trend: This is the most indicative of problems. Is there a significant improvement in the product defect rate (DPMO) after PM? Or at least, has it not worsened? If your DPMO consistently hovers around high points like 6210, then even a 100% PM completion rate is useless. You must link PM to certain specific defect types.
  3. MTBF/MTTR Improvement: Isn't the goal of PM to reduce unexpected machine downtime and shorten repair times? If your MTBF (Mean Time Between Failures) does not lengthen with an increase in PM frequency, but instead gets shorter, then your PM is simply ineffective. Conversely, if MTTR (Mean Time To Repair) decreases due to proper PM execution, this is also a good sign.

In other words, you need to make the "quality" of PM a measurable metric.

The Most Common Pitfall

The most common pitfall we used to fall into was treating all machines' PM KPIs equally. For example, new and old machines, critical and non-critical machines, all used the same standards. The result was that those old, problem-prone critical machines, despite having PM performed, simply didn't show good results.

To be honest, some machines are high-maintenance, while others are workhorses. For high-maintenance machines, you might need to look more strictly at KPP stability after PM; for workhorse machines, perhaps you need to pay more attention to how long it can last before failing (MTBF). If you put all your eggs in one basket, then when the basket tips over, you'll have nothing left. Customizing your KPIs will save you a lot of detours.

One Thing You Can Do Today

Pick one machine that frequently causes problems, and start tracking its KPP stability after PM.

Want to try it yourself?

Every tool mentioned in this article is available on InsightFab — just upload a CSV to analyze.

Go to Tools