Learn · Industrial Maintenance
MTBF, MTTR and Maintenance Metrics
Part of Maintenance Tech to CMRP · step 28 of 30 · next: The Millwright Apprenticeship
In learning paths: Maintenance Tech to CMRP
Assumes you know: Preventive Maintenance
Every maintenance metric answers one question and conceals another. MTBF speaks about reliability and says nothing about how long repairs take. MTTR speaks about repair and says nothing about how often you are repairing. Availability can look excellent while both of the others are getting worse. Reading them as a set, and knowing which decision each one supports, is the skill.
*Learn the Trades is a free study resource. We are not a licensing body, an authorized training provider, or an exam administrator. Reading this page does not award any card, license, or certification. Always verify requirements with the issuing authority linked in the sources.*Why it matters on the job
Metrics decide where the money goes. A plant that believes it has a reliability problem buys condition monitoring, alignment tools and analysis time. A plant that believes it has a maintainability problem buys spares, access improvements, better procedures and training. Both are expensive, and picking the wrong one costs a year.
They also decide how your department is judged. Knowing what a number can and cannot support is what lets a lead tech push back on a target that would make the plant worse.
What each measure is about
MTBF, mean time between failures, is about reliability. It expresses how much running the asset delivers between interruptions. When MTBF falls, the equipment is failing more often, and the work in front of you is analysis: what is causing the failures, and what removes the cause. Note the scope: MTBF describes repairable items. For items that are discarded rather than repaired, the corresponding measure is MTTF, mean time to failure.
MTTR, mean time to repair, is about maintainability. It expresses how quickly the asset comes back. When MTTR rises, the equipment is not necessarily failing more; getting it back is taking longer, and the causes tend to be logistical (spares not on site, no procedure, access requires dismantling something else, one person in the plant knows how). Be careful with this one: some plants measure time to repair, some time to recover, some time to restore, and those are not the same denominator. Comparing an MTTR against another site’s is comparing definitions.
Availability is about outcome. It expresses the share of the scheduled time the asset was actually available to run. It is the number production cares about, and it is a blend of the other two: a machine can hold high availability by failing constantly and being fixed quickly.
PM compliance is about the schedule. It expresses whether the planned tasks were done in their windows. It says nothing about whether they were the right tasks, and it is the easiest maintenance number in the plant to make look good.
Worked example: identical availability, opposite problems
Two machines, one month, 720 scheduled hours each.
Machine A stopped 6 times, each stoppage costing 2 hours: 6 × 2 = 12 hours of downtime, so 720 − 12 = 708 hours of uptime.
Machine B stopped once, for 12 hours: also 12 hours of downtime, and also 708 hours of uptime.
Availability is identical. Both delivered 708 of 720 hours, which is 708 / 720 = 0.9833, or 98.3 percent. A production report that carries availability alone shows these two machines as the same machine.

Same uptime, same availability, and two problems that need opposite responses
Now look at what the record actually says. Machine A ran an average of 708 / 6 = 118 hours between stoppages, and each stoppage was cleared in 2 hours. Machine B ran the whole month on a single interruption that took 12 hours to clear.
Machine A has a reliability problem: it is failing six times as often, and the department is very good at getting it back. Spare-parts investment will not help it, and neither will faster repairs, because the repairs are already fast. It needs root cause analysis.
Machine B has a maintainability problem: it hardly fails, and when it does the plant loses most of a shift. Analysis of the failure is worth less here than the questions of why the repair took 12 hours, what was waiting on a part, and what would have made access quicker.
Same availability, opposite work orders. That is why availability is never reported alone.
Turning numbers into decisions
- MTBF falling, MTTR steady: reliability is degrading. Go to analysis and to the condition data.
- MTBF steady, MTTR rising: repairs are getting harder. Look at spares availability, procedures, tooling, access and skills coverage.
- Availability holding while MTBF falls: the department is absorbing a worsening asset with speed. This is the state that looks healthy on a dashboard right up to the week it does not.
- PM compliance at 100 percent with failures rising: the schedule is being met and the tasks are not preventing anything. Review the task content and the interval basis before adding more PMs.
Where it bites
- Availability is not reliability. They move independently, and a single number that blends them can hide either one. Report them together or you have not reported anything.
- A metric with no decision attached is overhead. Before adopting a measure, name the action a bad value would trigger. If nobody can name one, collecting it costs time and buys nothing.
- Definitions have to be written down and left alone. Whether a planned outage counts as downtime, whether scheduled time means calendar time or production time, when the repair clock starts: change any of those and the trend breaks, in a way that looks like a real improvement.
- Small samples make wild averages. A mean built on two failures moves violently when the third arrives. Judge the direction over a meaningful window rather than the value this month.
- Targets get met the cheapest way available. A PM compliance target is met by doing the tasks or by rescheduling the window. Any metric used to judge people will be optimized as a metric, so pair it with one that moves the other way.
Verified requirements
| Where | Expires | Renewal | Continuing education |
|---|---|---|---|
| United States (federal) | Yes | 3 years | 50 course hours per 3-year cycle, drawn from two or more of the recertification activity categories; recertification application due within 90 days of the expiration date or the exam must be retaken |
Verified against the issuing authority; see sources below. Always confirm current rules with the authority before acting.