Active and Latent Failures Explained: 4 Evidence Tests That Keep Incident Reviews Systemic
Active failures appear close to the event, while latent failures remain in decisions, design, supervision, and work conditions. These four evidence tests help investigators read both without stopping at the operator.
Key takeaways
- 01An active failure is the action or omission close to the event, while a latent failure is a condition that made the event more likely or less recoverable.
- 02James Reason’s model is useful because it connects frontline behavior with decisions, design, supervision, and organizational conditions.
- 03A complete review tests what the person faced, what the system expected, what the organization knew, and what barrier should have interrupted the path to harm.
- 04Andreza Araujo’s work on safety culture distinguishes a complete record from a control that changes the way work is actually performed.
- 05Retraining the person involved may be necessary, but it cannot close a latent failure when the work system still creates the same exposure.
Active failures are actions or omissions close to an incident, while latent failures are weaknesses in design, planning, supervision, resources, or management decisions that make the event more likely. A sound investigation reads both layers together, because the visible mistake rarely explains why the system allowed one moment to become harmful.
What do active and latent failures mean?
An active failure occurs near the point where work is performed. A person selects the wrong isolation point, skips a verification, misreads a signal, or responds too late. The action matters because it belongs to the event sequence, yet it is only one part of that sequence.
A latent failure may have been present for weeks or years before the incident. It can appear as an equipment design that invites bypassing, a production plan that leaves no recovery time, a procedure that does not match the task, or supervision that treats repeated deviations as normal. Because it sits farther from the event, it is easier to miss during a hurried review.
James Reason’s work gives investigators a disciplined way to connect these layers. The point is not to excuse an unsafe decision or to erase personal responsibility. The point is to identify the conditions that shaped the decision and the barriers that should have stopped the sequence.
Why does the distinction matter after an incident?
An investigation that stops at the active failure usually produces a familiar action, namely retraining the person involved. Training can be appropriate when knowledge or skill was genuinely missing, although it cannot repair a control that was unavailable, impractical, or contradicted by the way the operation was managed.
Andreza Araujo makes this distinction central to the difference between formal compliance and operating culture in Safety Culture: From Theory to Practice. A report may identify the correct rule and still fail as a safety intervention if the next shift inherits the same weak barrier.
The four evidence tests below help an investigator move from the visible act to the conditions that made it possible.
What are the four evidence tests?
Test 1: What did the person actually face?
Start with the work as it existed at the time, not with the ideal sequence copied from a procedure. Check the layout, tools, access, lighting, time pressure, staffing, information available, and interruptions that shaped the decision. Evidence includes photographs, records, interviews, equipment condition, and the sequence of work before the event.
This test separates a deliberate disregard from a task that was ambiguous or difficult to execute as written. The distinction should be evidence-based. If the same workaround appears across shifts, the review should examine the work design before assigning the problem to one person.
Test 2: What did the system expect?
Compare the written expectation with the decisions that managers, planners, and supervisors rewarded in practice. A procedure may require a full isolation check while the schedule assumes a restart in ten minutes. A permit may require a field verification while the form is approved remotely without a site visit.
The gap between stated and operated expectations is a latent condition. It shows that the system sends competing signals, which means a person can follow the pressure that is measured rather than the rule that is displayed.
Test 3: What did the organization already know?
Search for earlier warnings, repeat defects, near-miss reports, maintenance requests, audit findings, worker concerns, and previous recommendations. Do not treat every old record as proof that the organization understood the risk. Establish whether the signal reached someone with authority, whether the response was adequate, and whether the condition was verified after the response.
A latent failure becomes especially important when the organization had credible evidence and still left the exposure unchanged. That evidence may reveal a weak escalation path, unclear ownership, or a review process that closes paperwork without closing risk.
Test 4: Which barrier should have interrupted the sequence?
Map the barriers that were supposed to prevent the event, detect the deviation, or reduce the consequence. Then test each barrier against evidence rather than its presence in a document. Was it physically available? Was it understood? Did it work under the conditions of the task? Did anyone have the authority and time to act when it showed weakness?
This is where the investigation becomes useful to leaders. A failed barrier points toward a decision, an owner, and a verification method. If the conclusion says only that a worker failed to follow the rule, the barrier analysis has not yet answered why the rule was the last fragile defense.
How should investigators record the findings?
Separate the findings into three connected records. First, describe the active action or omission without loaded language. Second, list the latent conditions that influenced the work, including design, planning, supervision, resources, and prior signals. Third, assign corrective actions to owners who can change those conditions, with a field verification date.
| Evidence layer | Investigator asks | Useful response |
|---|---|---|
| Active failure | What happened close to the event? | Clarify the decision sequence and immediate control gap. |
| Latent failure | What made the decision likely or the recovery weak? | Change the work condition, design, planning, supervision, or resource. |
| Barrier performance | What should have prevented or limited harm? | Verify that the barrier is present, functional, understood, and owned. |
The distinction also improves communication. Leaders receive a clearer explanation of what must change, while workers can see that the review examined the conditions around the task rather than searching for a convenient name to place at the end of the report.
When is retraining an incomplete action?
Retraining is incomplete when the person already knew the requirement, the procedure did not fit the task, the control was difficult to access, or supervisors had accepted the same workaround repeatedly. It may restore knowledge, but it does not remove a design weakness or resolve a conflict between production pressure and safe execution.
Use retraining when evidence shows a genuine capability gap, then pair it with a change that makes the expected action possible and verifiable. A competent operator working inside a weak system can reproduce the same event pattern as a newly trained operator, because the latent condition remains.
How can leaders test whether the review is complete?
A leader should be able to answer four questions after reading the report. What happened? Why did the decision make sense in that context? Which barrier failed or was missing? What will be different for the next person doing comparable work?
If the report cannot answer the second or fourth question, the investigation has probably described the event without explaining the system. In more than 250 cultural transformation projects, Andreza Araujo’s practical lesson is consistent: safety evidence earns value when it changes an operating decision, not when it simply creates a more polished record.
FAQ
What is an active failure? An active failure is an action, omission, or decision close to the incident that directly contributes to the event.
What is a latent failure? A latent failure is a pre-existing weakness in design, planning, supervision, resources, or management decisions that creates conditions for harm.
Does identifying an active failure mean blaming the operator? No. The review remains incomplete until it examines the conditions that shaped the decision and the barriers that failed.
Should every incident investigation use this distinction? It is a useful lens when work conditions, decisions, or controls may have influenced the event, although the method should fit the evidence and potential severity.
What should happen after a latent failure is found? Assign an owner with authority to change the condition, verify the control in the field, and check comparable work for the same exposure.
An incident review becomes systemic when it explains both the visible action and the conditions around it. Active and latent failures are not competing explanations. Together, they show what happened, why the sequence was possible, and what must change before another person meets the same risk.
Frequently asked questions
What is an active failure?
What is a latent failure?
Does identifying an active failure mean blaming the operator?
Should every incident investigation use the active and latent failure distinction?
What should happen after a latent failure is found?
About the author
Andreza Araújo
Safety Culture Expert | Senior EHS Executive
Andreza Araújo is a safety culture expert and senior EHS executive with more than 25 years of experience in environment, health and safety. She is a Civil Engineer and Occupational Safety Engineer from Unicamp, holds a Master's degree in Environmental Diplomacy from the University of Geneva, and completed sustainability studies at IMD Switzerland. Andreza has served in Global Head of EHS roles in Fortune 500 environments, leading cultural transformation programs across multinational operations. She has represented Brazil as a speaker at the United Nations in Paris and has spoken at the International Labour Organization in Turin. She is the author of more than 16 books on safety culture in Portuguese, Spanish, English and German. Her work has earned more than 10 EHS awards, including two recognitions from Indra Nooyi, former PepsiCo CEO.
- Civil & Safety Engineer (Unicamp)
- M.A. Environmental Diplomacy (University of Geneva)
- Sustainability Cert (IMD Switzerland)
- People Management & Coaching (Ohio University)
- UN Paris speaker representative for Brazil
- ILO Turin speaker
- LinkedIn Top Voice
- Indra Nooyi PepsiCo CEO recognition (2x)
Documentaries
Watch Andreza's documentaries
Three productions on safety culture, organizational failure and the human lessons behind major disasters.
Podcasts
Listen to Andreza's podcasts
She hosts three shows on safety leadership, EHS and organizational culture, in English and Portuguese.