Anyone who works in maintenance hears these three acronyms every day, but in practice they often end up being calculated wrong, interpreted superficially, or, worse, used only to fill out a report without leading to any real decisions. Let’s look at how to make them genuinely useful in a production plant.
MTBF – Mean Time Between Failures
The average time between one failure and the next. It’s calculated by dividing total operating time by the number of failures in that period:
MTBF = Total operating time / Number of failures
In everyday practice, the most common mistake is calculating it over periods that are too short, or across machines with very different failure patterns, which produces a number that doesn’t really say anything. A useful MTBF should be calculated per individual machine or, even better, per critical component (a motor, a drive, a valve), and tracked over time: if it’s dropping, that’s an early sign that something in the machine, or in the way it’s being managed, is getting worse — even before it becomes a serious problem.
A low MTBF on a specific component is often the best indicator for deciding where to direct a revamp or an improvement maintenance intervention, before downtime costs force the issue.
MTTR – Mean Time To Repair
The average time needed to repair a failure and bring the machine back into production, from the moment it stops to the moment it restarts:
MTTR = Total repair time / Number of interventions
Here the interesting part isn’t so much the number itself, but how it breaks down. A high MTTR can stem from very different causes: diagnosis time (often the most critical part, especially on plants with complex automation), waiting for spare parts, availability of the right person at the right time, or difficulty accessing the components that need repairing. Separating out these times when tracking interventions is what turns MTTR from a passive number into an operational tool: if diagnosis time weighs too heavily, you need to invest in technical documentation or diagnostic tools; if the wait for spare parts weighs too heavily, the problem lies in warehouse management.
OEE – Overall Equipment Effectiveness
This is the most complete indicator, because it brings together three factors:
OEE = Availability × Performance × Quality
- Availability: how much time the machine actually produced compared to planned time (this is where both failures — so MTBF and MTTR — and scheduled downtime and changeovers come into play).
- Performance: how much the machine produced compared to its rated speed (slowdowns, micro-stops, suboptimal cycles).
- Quality: how much of the output produced is actually compliant, without scrap or rework.
The value of OEE lies in the fact that it forces you to look beyond maintenance alone: a machine can have excellent MTBF and a very low MTTR, yet a mediocre OEE due to performance bottlenecks or quality issues linked to settings, materials, or wear that hasn’t turned into a failure yet. That’s why OEE is the right indicator to bring into discussions with production and management: it speaks the same language as budget and the plant’s overall efficiency, not just maintenance.
Making them operational, not just report numbers
Some practical criteria that make the difference:
- Consistency in data collection. If machine downtime is logged inconsistently by shift staff, every KPI built on top of it will be unreliable. It’s worth investing time in standardizing downtime reason codes even before calculating the indicators.
- Reference thresholds per machine, not generic ones. An OEE of 75% can be excellent on a complex plant and poor on a simple line: benchmarks should be built on the specific history of the plant, not taken from generic industry tables.
- Link the KPIs to budget decisions. A declining MTBF and a rising MTTR on a critical asset are the strongest argument for justifying an investment in revamping or replacement to management: they turn a spending request into a data-based case.
- Periodically review what you’re measuring. A KPI that no longer leads to any concrete action should be revised or dropped: the risk is ending up with dashboards full of numbers no one uses to make decisions anymore.
Used well, MTBF, MTTR and OEE aren’t just indicators to bring to a monthly meeting: they’re the tool that lets you direct budget, staff and modernization efforts where they’re really needed, before downtime ends up deciding for you.
