FMEA (Failure Mode and Effects Analysis) is a structured method for identifying how an asset can fail, rating each failure's severity and likelihood, and prioritizing corrective actions.
What is FMEA?
FMEA, or Failure Mode and Effects Analysis, is a step-by-step risk assessment technique that examines every realistic way a component, machine or process can fail, then scores how severe each failure would be and how likely it is to be detected before it causes harm. The output is a ranked list of failure risks paired with the actions that reduce them.
Maintenance teams use FMEA to move away from reactive repairs and toward planned, evidence-based work. Instead of waiting for a pump to seize or a conveyor to jam, a reliability engineer lists the possible failure modes of each component — bearing wear, seal leakage, misalignment, electrical overload, contamination — and rates them. Because the method forces a structured look at cause, effect and detection, it surfaces single points of failure that generic inspection checklists miss.
The technique originated in the United States military in the late 1940s and was later adopted by aerospace, automotive and process industries before spreading into manufacturing, energy, facilities management and healthcare. It differs from a simple risk register because it works bottom-up at the component level and produces a numeric priority score rather than a subjective ranking.
FMEA is not the same as Root Cause Analysis. FMEA is proactive — it runs before failures happen and asks "what could go wrong?". Root cause analysis is reactive — it runs after an incident and asks "why did this go wrong?". Mature maintenance programmes use both.
How FMEA Works in Maintenance Management
A maintenance FMEA follows a repeatable sequence. A cross-functional team of maintenance technicians, operators and reliability engineers works through these steps for one asset or system at a time.
- 1Define the scope. Select one asset, system or process and set clear boundaries so the analysis stays manageable.
- 2Break the asset into functions. State what each component is supposed to do, in measurable terms such as "deliver 40 litres per minute at 3 bar".
- 3Identify failure modes. List every realistic way each function can fail — fails completely, fails partially, fails intermittently or degrades over time.
- 4Describe the effects. Record what actually happens when each failure mode occurs — reduced output, safety risk, environmental release, unplanned downtime, quality loss.
- 5Assign scores. Rate severity, occurrence and detection on a scale of 1 to 10 using a written scale so scores stay consistent between analysts.
- 6Calculate the Risk Priority Number. Multiply severity × occurrence × detection to produce an RPN between 1 and 1,000.
- 7Set actions and owners. For every high-ranking risk, agree a corrective task, name a responsible person and set a due date. Actions might be a design change, a new inspection route, a spare-part stocking policy or a lubrication change.
- 8Re-score and review. Once an action is complete, re-rate the failure mode and confirm the RPN has dropped. Repeat on a scheduled cycle.
Understanding the Risk Priority Number
Severity measures how bad the effect is, with 10 reserved for safety or regulatory breaches. Occurrence measures how often the cause is likely to happen, based on failure history pulled from your CMMS. Detection is inverted: a score of 10 means the failure will almost certainly go unnoticed until it stops the asset, while a 1 means existing controls would catch it immediately. Teams then set an RPN threshold — often 100 to 200 — above which action is mandatory. Modern automotive and manufacturing programmes increasingly use the Action Priority tables from the AIAG-VDA FMEA handbook instead of a raw RPN cut-off, because a severity-10 safety risk should never be buried behind a high-occurrence nuisance fault.
Key Characteristics of FMEA
- Systematic and bottom-up. It works from individual component failure modes upward to asset-level consequences, so nothing important is assumed away.
- Quantified, not subjective. Every risk receives numeric severity, occurrence and detection scores against a documented scale, making comparisons between assets defensible.
- Team-based. It relies on a facilitator plus technicians, operators and engineers, which is why it captures tacit knowledge that documentation alone does not hold.
- Proactive by design. It is performed before failures occur, so it feeds directly into preventive and predictive maintenance planning rather than repair work.
- A living document. An FMEA is reviewed and re-scored after modifications, incidents and changes in operating conditions, not filed away after the first workshop.
- Traceable. Each entry links a failure mode to a cause, an effect, an existing control and an assigned action, creating an audit trail for safety and quality systems.
Types of FMEA
FMEA is not a single document. The version you use depends on whether you are analysing a design, a process or an operating asset.
| FMEA Type | Primary Focus | Who Runs It |
|---|---|---|
| Design FMEA (DFMEA) | Failure modes built into a product or equipment design before it is installed | Design and product engineers |
| Process FMEA (PFMEA) | Failure modes in manufacturing, assembly or service steps | Process and quality engineers |
| Asset or Maintenance FMEA | Equipment failure modes on the plant floor and their effect on production | Reliability and maintenance teams |
| FMECA | FMEA plus a criticality ranking that weights severity against failure rate | Defence, aerospace, nuclear, rail |
| Functional FMEA | High-level system functions, used early in a project before components are chosen | Project and systems engineers |
FMEA Examples and Use Cases
FMEA is most valuable when it is applied to a specific asset with real failure data behind it. These three examples show how the output changes maintenance decisions.
Example 1: Centrifugal pump in a water treatment plant
A reliability team analyses a critical transfer pump. One failure mode — mechanical seal leakage — scores severity 7 (partial loss of supply plus a housekeeping hazard), occurrence 6 (based on three seal failures in two years) and detection 4 (drip trays and daily rounds catch most leaks early), giving an RPN of 168. The team responds with a vibration and temperature route, a seal-face inspection at every planned shutdown, and a stocked spare seal kit. A second mode — motor bearing seizure — scores severity 9, occurrence 3 and detection 8, an RPN of 216. Even though it happens less often, its high severity and poor detectability push it above the seal leak in priority.
Example 2: Conveyor drive on a packaging line
On a high-speed packaging line, FMEA highlights belt mistracking as a low-severity but high-occurrence failure mode that repeatedly stops production for short periods. Because the aggregate downtime is significant, the team adds a guide-roller inspection to the weekly schedule and installs a belt-alignment sensor. Meanwhile, gearbox failure scores high on severity and detection, so the maintenance plan shifts that gearbox from reactive repair to oil analysis every quarter. FMEA has, in effect, split one asset into two different maintenance strategies.
Example 3: Rooftop HVAC unit in a commercial building
A facilities team runs an FMEA across twelve rooftop units serving a data-adjacent office. Compressor failure scores severity 8 because it affects occupant comfort and tenant agreements. Dirty condenser coils score severity 3 but occurrence 9, since the building sits near a construction corridor. The RPN comparison shows that a cheap quarterly coil-cleaning task reduces a very common failure mode, while a refrigerant-leak detection programme addresses the rare but expensive one. Both are funded because the scoring makes the trade-off visible to the budget holder.
Related Terms
Root Cause Analysis is the reactive counterpart to FMEA — it investigates failures that have already happened, while FMEA predicts them in advance.
Reliability-Centered Maintenance often uses FMEA output as its starting evidence when deciding which tasks belong in a maintenance programme.
Risk Priority Number is the score FMEA produces from severity, occurrence and detection, and it drives the ranking of corrective actions.
Preventive Maintenance is usually the first place FMEA findings land, because new inspection and servicing tasks are the cheapest way to lower a risk score.
Condition-Based Maintenance is often the answer when FMEA shows a failure mode with high severity and poor detectability, since monitoring improves detection scores.
Criticality Analysis ranks whole assets rather than individual failure modes, so it is typically done first to decide which assets deserve a full FMEA.
Frequently Asked Questions
FMEA (Failure Mode and Effects Analysis) is a structured risk assessment method used in maintenance to identify every way an asset can fail, describe the effect of each failure, score its severity, likelihood and detectability, and rank the risks so maintenance teams can act on the most critical ones first.
An FMEA team breaks an asset into functions and components, lists possible failure modes, and records the effect of each one. Every failure mode receives severity, occurrence and detection scores from 1 to 10. Multiplying the three produces a Risk Priority Number that ranks the risks and directs corrective actions.
A Risk Priority Number (RPN) is the product of three ratings from 1 to 10: severity of the failure effect, occurrence of the cause, and detection. Detection is inverted, so a score of 10 means the failure will probably go unnoticed. RPN values run from 1 to 1,000, and higher numbers signal higher priority.
FMECA adds a criticality analysis to standard FMEA. Each failure mode gets a second dimension that combines severity with how often the failure occurs, producing a criticality number rather than a simple risk ranking. FMECA is common in defence, aerospace, nuclear and rail, where quantified failure consequences are mandatory.
FMEA is proactive: it runs before failures occur and predicts what could go wrong across an entire asset or process. Root cause analysis is reactive: it investigates a failure that has already happened and traces it back to its origin. Most mature maintenance programmes run both, using FMEA to prevent and root cause analysis to learn.
Treat the FMEA as a living document. Review it after every significant failure, when an asset is modified or relocated, when operating conditions change, or when new failure data appears in your CMMS. A full review at least once every two years keeps risk scores honest and actions relevant.