Root cause analysis (RCA) is a structured way of working back from a problem, such as a scrapped lot, a stopped line or a late shipment, to the conditions that allowed it to happen, so that the fix removes those conditions instead of the symptom. A root cause is the deepest cause you can act on. NASA’s procedural requirements define it as a condition “primarily associated with organizational factors” that existed before the nearer causes. Most plant investigations stop one layer short of it.
In this guide
What is root cause analysis, in one definition?
What is the difference between a root cause and a symptom?
Why do plants do root cause analysis?
What does a problem cost when the cause is never found?
What are the steps in a root cause analysis?
Which root cause analysis method should a plant use?
What does a documented root cause analysis look like?
What does root cause analysis look like on a discrete line?
Why do most root cause analyses fail?
How do you verify that a corrective action worked?
What should a plant record from every root cause analysis?
Where does Morsa fit in root cause analysis?
What changed at a plant running Morsa?
FAQ
Sources, Changelog, Related pages
What is root cause analysis, in one definition?
Root cause analysis is the discipline of separating what happened from why it was possible to happen, then acting on the second. The US Agency for Healthcare Research and Quality, which has studied the method longer than any manufacturing body, defines it as “a structured method used to analyze serious adverse events” that identifies “how the event occurred (through identification of active errors) and why the event occurred (through systematic identification and analysis of latent errors).”
The active error is the fuse that blew. The latent error is whatever made the overload inevitable. A plant that replaces the fuse has corrected the problem. It has not done root cause analysis.
The word “root” carries a warning. Peerally and colleagues wrote in BMJ Quality & Safety in 2017 that “by implying, even inadvertently, that a single root cause (or a small number of causes) can be found, the term ‘root cause analysis’ promotes a flawed reductionist view.” Real failures usually have several causes at different depths, which is why the next section matters more than the definition.
What is the difference between a root cause and a symptom?
A symptom is what you noticed. A root cause is the earliest condition you can change so the symptom cannot come back. Between them sit two more layers, and NASA’s procedural requirements for mishap investigation, effective 6 July 2020, are the clearest public definitions of all four.
Layer | NASA’s definition (NPR 8621.1D, 2020) | In Taiichi Ohno’s machine-stop example |
|---|---|---|
Symptom | The undesired outcome | The machine stopped |
Proximate cause | “The event that occurred, including any conditions existing immediately before the undesired outcome, directly resulted in its occurrence” | There was an overload and the fuse blew |
Intermediate cause | “An event or condition that existed before the proximate cause, directly resulted in its occurrence, and if eliminated or modified, would have prevented the proximate cause from occurring” | The bearing was not lubricated; the pump was not pumping; the pump shaft was worn |
Root cause | “An event or condition, primarily associated with organizational factors, which existed before the intermediate cause and directly resulted in its occurrence” | There was no strainer on the pump, so metal scraps got in |
Contributing factor | “An event or condition that may have contributed to the occurrence of an undesired outcome, but if eliminated or modified, would not on its own have prevented the occurrence” | Nobody had a standard for what a lubrication pump needs |
FOUR CAUSAL LAYERS
Where most plant investigations stop, and where a root cause actually sits
SYMPTOM
The machine stopped
What the operator reported. Replacing the fuse fixes this for one shift.
PROXIMATE CAUSE
The fuse blew on an overload
The event immediately before the outcome. Most shop-floor fixes end here.
INTERMEDIATE CAUSE
The bearing ran dry because the pump shaft was worn
Conditions that made the proximate cause certain. A maintenance work order ends here.
ROOT CAUSE
ACT HERE
No strainer on the pump, and no standard that said there should be one
The organizational condition. Fix this and the failure mode is gone.
Definitions: NASA NPR 8621.1D, Appendix A, 2020. Example: Taiichi Ohno, Toyota Production System, 1988, as reproduced by the Lean Enterprise Institute.
The Ohno chain is the canonical five whys, reproduced by the Lean Enterprise Institute from his 1988 book: “Why did the machine stop? There was an overload and the fuse blew. Why was there an overload? The bearing was not sufficiently lubricated. Why was it not lubricated? The lubrication pump was not pumping sufficiently. Why was it not pumping sufficiently? The shaft of the pump was worn and rattling. Why was the shaft worn out? There was no strainer attached and metal scraps got in.” The Institute’s point: “without repeatedly asking why, managers would simply replace the fuse or pump and the failure would recur.”
Why do plants do root cause analysis?
Plants do it because a customer, a regulator or a standard requires it, and because a repeat failure costs more than the first one. The regulatory form is explicit. Under US law for medical device makers, 21 CFR 820.100 (as of the 1 April 2024 edition) requires procedures for “investigating the cause of nonconformities relating to product, processes, and the quality system” and for “verifying or validating the corrective and preventive action to ensure that such action is effective.” Every activity “shall be documented.”
ISO 9001 asks for the same thing in the clause on nonconformity and corrective action. The certification body NQA puts the intent plainly: “An important part of corrective action is to carry out a root cause analysis in relation to the issue that occurred. If you don’t get to the bottom of why or how it happened, then it’s likely whatever fix you implement will not be fully effective.” Automotive customers add their own layer: AIAG’s CQI-20 Effective Problem Solving guideline (version 2, August 2018, $70 for members and $211 for non-members) exists so that suppliers “ensure that steps are not skipped.”
The US Department of Energy has required root cause analysis on reportable occurrences since its 1992 guidance document, with one rule worth borrowing: “The level of effort expended should be based on the significance attached to the occurrence.”
What does a problem cost when the cause is never found?
It costs five times more downstream than at the machine. NIST’s survey of US discrete manufacturers (AMS 100-34, June 2020, 2016 data) put preventable maintenance losses at “$119.1 billion: $18.1 billion due to downtime, $0.8 billion due to defects, and $100.2 billion due to lost sales from delays and defects.” The same survey found that the quarter of plants most reliant on reactive maintenance “was associated with 3.3 times more downtime” and “16.0 times more defects” than the least reliant quarter, and that plants investing in preventive or predictive maintenance had “44% less downtime” and a “54% lower defect rate.” NIST’s own reading: “reactive maintenance reduces quality and increases uncertainty in production time.”
WHAT AN UNFOUND CAUSE COSTS
What does a problem cost when the cause is never found?
$100.2B
of the $119.1 billion US discrete manufacturers lost to preventable maintenance issues in 2016 was lost sales from delays and defects, not downtime
NIST AMS 100-34, June 2020
16.0x
more defects at the quarter of plants most reliant on reactive maintenance than at the least reliant quarter
NIST AMS 100-34, June 2020
81 min
average time to recover from an unplanned stop in 2024, up from 49 minutes five years earlier
Siemens, True Cost of Downtime 2024, 181 interviews
$2.3M
cost of one hour of unplanned downtime in a large automotive plant
Siemens, True Cost of Downtime 2024
READ TOGETHER
Failures are getting rarer and harder to diagnose at the same time. Siemens attributes the slower recovery to a skills and knowledge gap. That gap is what a written root cause analysis replaces.
NIST is a government survey of 85 US establishments. Siemens is vendor research (Senseye) across automotive, FMCG, heavy industry and oil and gas.
Diagnosis is getting slower even as failures get rarer. Siemens’ True Cost of Downtime 2024, 181 interviews at large industrial firms, found plants “now suffer an average of 25 downtime incidents a month per facility, down from 42 in 2019,” still losing “27 hours a month.” But “five years ago, it took an average of 49 minutes to get production back up and running following downtime. Now, it takes 81 minutes,” which Siemens attributes to a “skills and knowledge gap” after skilled maintenance people left. In automotive, “unplanned downtime now costs $2.3 million an hour.” Siemens owns a predictive maintenance vendor, so treat the totals as vendor research with a stated sample.
What are the steps in a root cause analysis?
Five steps, in order, and the first and last are the ones plants skip. AHRQ’s protocol starts the same way: “data collection and reconstruction of the event in question through record review and participant interviews.”
Define the problem in numbers. What, where, when, how many, against what standard. “Scrap on Line 2” is not a problem statement. “38 of 500 pieces on work order 4471 failed the edge inspection between 14:10 and 15:40 on 9 September” is.
Contain and preserve. Quarantine the lot, hold the tool, save the machine log, photograph the setup before anyone “tidies up.” Evidence disappears in the first hour.
Reconstruct the timeline. Records first, then the people who were there. Write down what happened in sequence before anyone says why.
Work down the causal layers. Proximate, intermediate, root, contributing, using a method from the next section. Stop when the next “why” would be outside the plant’s control.
Act and verify. Assign the countermeasure to a named owner with a date, then check the failure mode later with data, not with a signature. This is the step US device regulation spells out as “verifying or validating the corrective and preventive action to ensure that such action is effective.”
FIVE STEPS
A root cause analysis that holds
STEP 01
Define in numbers
What, where, when, how many, against which standard.
THE EXAMPLE
38 of 500 pieces on work order 4471 failed edge inspection, 14:10 to 15:40, 9 September
STEP 02
Contain and preserve
Quarantine the lot, hold the tool, save the log, photograph the setup.
THEN
The 38 pieces and the tooling insert go to the quality cage, tagged
STEP 03
Reconstruct
Records first, then interviews. Sequence before reasons.
THEN
Insert changed at 13:50 by the relief operator; no first-piece check recorded
STEP 04
Work down the layers
Proximate, intermediate, root, contributing.
THEN
Root: the changeover standard does not require a first-piece check on relief shifts
STEP 05
Act and verify
Named owner, date, and a data check on the failure mode weeks later.
THEN
Standard revised by 16 September; edge rejects on Line 2 reviewed at 30 days
Steps follow the AHRQ protocol and 21 CFR 820.100. The example is illustrative.
Scale the effort to the loss. The DOE’s rule from 1992 still holds: most routine occurrences “need only a scaled-down effort,” and the formal models are for the serious ones.
Which root cause analysis method should a plant use?
Use the simplest method that reaches the organizational layer for the problem in front of you. The six a plant actually needs are below; the full tool-by-tool guide with a decision table is root cause analysis tools.
5 whys
Use it when: One clear failure with a single chain of causes
What it produces: A chain from symptom to root, as in Ohno’s example
Where it falls short: One chain only; a second contributing chain gets missed
Fishbone (Ishikawa) diagram
Use it when: Several possible causes across people, machine, material, method, measurement, environment
What it produces: A sorted list of candidate causes to test
Where it falls short: It lists causes; it does not rank or prove them
Pareto chart
Use it when: Many recurring defects and you need to pick which to attack
What it produces: The few defect types that make most of the loss
Where it falls short: Needs coded data; narratives cannot be charted
FMEA
Use it when: A new part, process or line, before failures happen
What it produces: Ranked failure modes with actions
Where it falls short: Heavy; a living document or a dead one
Fault tree analysis
Use it when: Safety-critical or multi-condition failures
What it produces: A logic tree of the conditions that must combine
Where it falls short: Needs an analyst and time
8D
Use it when: A customer complaint that requires a formal, signed response
What it produces: Eight disciplines, with root cause and escape point at D4
Where it falls short: The form gets filled; the analysis gets skipped
8D began at Ford in the 1980s as Team Oriented Problem Solving, per Quality-One, and its fourth discipline is “Root Cause Analysis (RCA) and Escape Point,” which asks not only why the defect was made but why it was not caught. The layered view of failure, where several defenses each have holes and a loss happens when the holes line up, is James Reason’s Swiss cheese model, published in 1990, and it is why a good analysis looks for more than one cause.
What does a documented root cause analysis look like?
The clearest recent example is a federal one. On 11 August 2025 an explosion at U.S. Steel’s Clairton coke works killed two people, and the US Chemical Safety Board published its investigation.pdf?17370) in August 2026. Read it as a layered analysis.
Symptom. “Two fatalities, five serious injuries, six other injuries, and an estimated $52.5 million in property damage,” from “approximately 19 pounds of coke oven gas.”
Proximate cause. “The overpressurization of a double disc gate valve constructed of cast iron, which resulted in the release of flammable coke oven gas that ignited and exploded.”
Intermediate causes. “U.S. Steel’s and MPW’s lack of a procedure detailing how to safely perform the valve washing operation, along with U.S. Steel’s and MPW’s failure to identify or adequately address the potential hazards of the operation.”
Root cause. “U.S. Steel lacked effective process safety management systems that could have prevented or reduced the severity of this incident.”
The recurrence. The plant had a comparable event before: “U.S. Steel missed a key opportunity after the July 2010 incident, which injured 20 people, to recognize that a coke oven gas explosion in a coke battery basement was not only possible but could be extensive enough to cause severe consequences.” On applying a management system that would have covered the operation, “U.S. Steel affirmatively chose not to do so, based on the company’s belief that it was not required to do so.”
The CSB states the lesson in general terms: “Limited investigations may only provide feedback concerning a specific incident scenario, rather than discovering lessons that might apply to broad categories of incidents.” The 2010 investigation found a cause. It did not find the root. Fifteen years later the same category of failure returned with a different valve.
What does root cause analysis look like on a discrete line?
On a discrete line the same layers apply, but the failure is usually a stopped job or a scrapped lot rather than an explosion, and the root cause is usually a missing standard or a missing owner. Bill Bourgeous, Plant Operations Manager at Intertape Polymer Group’s Tremonton plant, told IndustryWeek why the discrete case is the lucky one: “Discrete processes stop when there’s a problem. Continuous processes, they’ll keep running and keep running and keep making bad product.”
Take a coordination failure from a Morsa customer plant (Morsa customer data). A manager set an 8,000-part night-shift target. Seven hours later, dispatch posted photos in the group showing parts unavailable. Nobody connected the two.
Symptom. The night target was blocked.
Proximate cause. Parts were not available at dispatch.
Intermediate cause. The schedule said 4,131 units for the shift and dispatch had 4,035, a 96-piece gap that sat in two separate records for seven hours.
Root cause. No one owned the comparison of schedule against dispatch during the shift, so a gap had no path to the person who could close it.
Countermeasure. Compare the two continuously and route the gap to an owner. Morsa found the 96-piece shortage, linked it to the target, told dispatch to arrange the parts, told production the target was blocked by supply, and opened a high-priority dependency, in under a minute.
The countermeasure that holds is the one aimed at the root: a standing owner for the comparison, not a reminder to “check dispatch more often.” Which kind a plant picks is the subject of the next section.
Why do most root cause analyses fail?
Most fail because the investigator stops at a person and the countermeasure is a reminder. The evidence on this is uncomfortable and mostly from healthcare, where the method has been studied hardest. AHRQ’s primer states that “studies have shown that RCAs often fail to result in the implementation of sustainable systems-level solutions,” naming “overreliance on weak solutions (such as educational interventions and enforcing existing policies)” as a common reason. Peerally and colleagues found “the endemic tendency of investigators to settle for administrative and perhaps ‘weaker’ solutions (such as reminders) rather than those that address the latent causes,” and concluded the method “has consistently failed to deliver benefits on the scale or quality needed.”
Plant practitioners describe the same pattern. Mark Kaganov wrote in Quality Magazine on 18 September 2026: “Most companies don’t struggle because they use the wrong RCA technique; they struggle for two completely different reasons,” and the untrained conclusions he lists are “Operator failed to follow instructions” and “Training was incomplete.” Michael Holloway, President of 5th Order Industry, wrote in Plant Services that “the bearing failed” is “one of the most consistently inaccurate conclusions drawn in the field,” because “many bearing failures are not purely mechanical events but are the result of decisions that were not made, were delayed, or were based on incomplete or inaccurate information.”
W. Edwards Deming’s estimate, from his own experience rather than a study, was that “94% belongs to the system (responsibility of management) 6% special.” An analysis that ends at an operator has, on that estimate, a one in sixteen chance of being right.
How do you verify that a corrective action worked?
You verify with the failure mode’s own data, weeks later, not with a closed ticket. Mani Chandra Raparla put the test in Quality Digest on 17 September 2026: “Fix the case, and the defect returns tomorrow. Fix the attribute, and the failure mode is gone.” The standard he holds a floor to is “out-of-control signals investigated; root causes found; corrective actions closed.”
Three checks separate a verified action from a filed one.
The measure existed before the fix. If the defect rate was never charted, there is nothing to compare against.
The check is dated and owned. A named person looks at the failure mode 30, 60 or 90 days out. If the check has no owner, it does not happen.
The fix changed a standard, not a person. A revised setup sheet, a poka-yoke, a strainer on the pump. Rajesh Sharma, Quality and Operating Systems Manager at MSM, Magna Powertrain, described the target state to IndustryWeek: “We build to not allow you to make a bad part.”
That plant, 361 employees, reported a first-pass quality yield of 99.8% in its 2025 IndustryWeek Best Plants audit; its sister plant Pullmatic, 152 employees, reported 99% and 100% on-time delivery. Bourgeous at Intertape describes what holding that takes. “It’s hand over hand. You pull,” he says. “But if you let go, you ease up at all, you fall really fast.”
What should a plant record from every root cause analysis?
Record the cause as a code, not a story, and record the verification as a date with an owner. Kaganov’s test for a quality manager is whether they can name the five most common root causes of the last year: “Could they answer in five minutes?” Most cannot, because causes live in narrative fields nobody can count. A minimal record that can be counted:
Field | Example |
|---|---|
Problem statement | 38 of 500 edge rejects, WO 4471, Line 2, 9 Sep, 14:10 to 15:40 |
Cost | 38 pieces scrapped, 1.5 hours rework, one customer date at risk |
Proximate cause | Insert changed without first-piece check |
Root cause code | STD-03: changeover standard incomplete for relief shifts |
Countermeasure | First-piece check added to relief changeover standard |
Owner and date | Line 2 supervisor, 16 September |
Verification | Edge rejects on Line 2 reviewed at 30 days by the quality lead |
Recurrence | None in 30 days, or the record reopens |
The regulation’s version of the last line is 21 CFR 820.100(b): “All activities required under this section, and their results, shall be documented.” The manufacturing version is simpler: an analysis nobody can find is an analysis that did not happen. Where the counts come from once the plant has codes, and how they feed the daily review, is covered in manufacturing management software.
Where does Morsa fit in root cause analysis?
Morsa, the AI that operates the factory for you, does not run the analysis. It does the two things the evidence says plants skip: it makes the countermeasure a commitment with an owner and a date, and it verifies on proof.
What it reads. The ERP, the MES, the QMS and CMMS where they exist, plus the places people report problems, whatever they are: WhatsApp, Microsoft Teams, email, SMS or whatever the plant runs on. A rejection rising on a press since the morning shift is a signal the same way a scrap entry in the MES is.
What it decides. Which work order, customer order, machine and owner the problem touches, what it affects downstream, and within approved rules what happens next: hold the lot, route to quality, escalate, or ask for approval.
What it does. Opens the corrective action as a job with an owner and a due date, chases before the date, escalates when it goes quiet, and closes it only on evidence: a photo, a document, a system entry. It also keeps the counts that tell a plant what to investigate next, such as manpower short on Shift B four times this month, or the rework backlog growing whenever escalation waits a day.
What changed at a plant running Morsa?
J4S, a 120-person glass plant onboarded in two days, took on-time completion of operational commitments from about 30% to about 75% in the first four weeks, across about 900 commitments, with no new software (how the glass plant did it). Those commitments include the ones a root cause analysis produces: a changed standard, a staged lot, a maintenance window. Each stays open until the evidence shows up. Production Head Anil Kohli: “Our people don’t have to learn any new software. People just message the way they always have. Morsa coordinates all the messages in the background.” At JRG Automotive, running live, Morsa identified 5 of 8 production-stopping material shortages early enough to act, and saved the plant $2 million.
How the loop runs, from signal to verification, is in AI copilot for manufacturing. The machine-side causes, and what happens after a prediction, are in predictive maintenance software.
Sources
NASA, NPR 8621.1D, NASA Procedural Requirements for Mishap and Close Call Reporting, Investigating, and Recordkeeping, Appendix A definitions, effective 6 July 2020, https://nodis3.gsfc.nasa.gov/displayDir.cfm?t=NPR&c=8621&s=1D&page_name=AppendixA
Agency for Healthcare Research and Quality, PSNet, Root Cause Analysis primer, last updated 15 June 2024, https://psnet.ahrq.gov/primer/root-cause-analysis
Peerally, Carr, Waring and Dixon-Woods, The problem with root cause analysis, BMJ Quality & Safety 26(5), 2017, open access, https://pmc.ncbi.nlm.nih.gov/articles/PMC5530340/
Lean Enterprise Institute, 5 Whys, Lean Lexicon, quoting Taiichi Ohno, Toyota Production System, 1988, p. 17, https://www.lean.org/lexicon-terms/5-whys/
US Government Publishing Office, 21 CFR 820.100 Corrective and preventive action, CFR edition of 1 April 2024, https://www.govinfo.gov/content/pkg/CFR-2024-title21-vol8/xml/CFR-2024-title21-vol8-sec820-100.xml
NQA, ISO 9001 Implementation Guide, undated PDF (certification body guidance, paraphrases ISO 9001), https://www.nqa.com/medialibraries/NQA/NQA-Media-Library/PDFs/NQA-ISO-9001-Implementation-Guide.pdf
Automotive Industry Action Group, CQI-20 Effective Problem Solving Guideline, version 2, August 2018, store listing with prices, https://www.aiag.org/store/publications?category=Quality
US Department of Energy, Root Cause Analysis Guidance Document, DOE-NE-STD-1004-92, February 1992, OSTI record 5696132 (abstract), https://www.osti.gov/biblio/5696132
NIST, Thomas and Weiss, Economics of Manufacturing Machinery Maintenance, AMS 100-34, June 2020, 2016 data, 85 survey responses, https://nvlpubs.nist.gov/nistpubs/ams/NIST.AMS.100-34.pdf
Siemens, The True Cost of Downtime 2024, 181 interviews, April 2019 to March 2023 (vendor research), https://assets.new.siemens.com/siemens/assets/api/uuid:1b43afb5-2d07-47f7-9eb7-893fe7d0bc59/TCOD-2024_original.pdf
US Chemical Safety and Hazard Investigation Board, U.S. Steel Clairton Coke Works investigation report, Report No. 2025-03-I-PA, published August 2026, https://www.csb.gov/assets/1/20/us_steel_clairton_investigation_report_publication_copy_(1).pdf?17370
Quality-One International, Eight Disciplines of Problem Solving (8D), https://quality-one.com/8d/
Wikipedia, Swiss cheese model (James Reason, 1990; secondary source), https://en.wikipedia.org/wiki/Swiss_cheese_model
The W. Edwards Deming Institute, Dr. Deming discussing the 14 points, quoting Out of the Crisis, p. 315, https://deming.org/deming-library-video-with-dr-deming-discussing-the-14-points/
IndustryWeek, Robert Schoenberger, Intertape Polymer Group, Tremonton, Utah: A Best Plants Three-Peat, 29 October 2025 (Bill Bourgeous quotes), https://www.industryweek.com/members/article/55326176/intertape-polymer-group-tremonton-utah-a-best-plants-three-peat
IndustryWeek, Jill Jusko, MSM, Magna Powertrain Takes a Quality Approach to Operational Excellence, 21 November 2025 (Rajesh Sharma quote, 99.8% first-pass yield), https://www.industryweek.com/resources/industryweek-best-plants-awards/article/55331914/msm-magna-powertrain-takes-a-quality-approach-to-operational-excellence-industryweek-best-plants
IndustryWeek, Jill Jusko, 2025 IndustryWeek’s Best Plants: A Model of Culture-Driven Manufacturing Excellence, 27 October 2025 (Pullmatic figures), https://www.industryweek.com/resources/industryweek-best-plants-awards/article/55325939/2025-industryweeks-best-plants-a-model-of-culture-driven-manufacturing-excellence
Quality Magazine, Mark Kaganov, The Two Reasons Most Corrective Actions Fail, 18 September 2026, https://www.qualitymag.com/articles/99924-the-two-reasons-most-corrective-actions-fail
Plant Services, Michael D. Holloway, Maintenance Mindset: The bearing didn’t fail, the system did, 29 April 2026, https://www.plantservices.com/monitoring/machinery-lubrication/article/55374140/maintenance-mindset-the-bearing-didnt-fail-the-system-did
Quality Digest, Mani Chandra Raparla, Your Transactions Are a Process Too: SPC Beyond the Shop Floor, 17 September 2026, https://www.qualitydigest.com/inside/improvement-tools-article/your-transactions-are-process-too-spc-beyond-shop-floor-091726
Morsa’s own figures (J4S, JRG Automotive, the 96-piece night shift) are Morsa customer data. Every external source above was fetched on 20 September 2026 and the sentence carrying each figure was quoted; see the sources listed on this page.
Changelog
21 September 2026: first draft. Definition guide built on NASA’s four causal layers, the CSB Clairton report of August 2026 as the documented case, NIST AMS 100-34 and Siemens 2024 for cost, and the AHRQ and BMJ evidence on why corrective actions fail.

