Safety Data Quality Explained: 5 Decision Checks
Safety data quality means that a reported number is traceable, comparable, timely, challengeable, and useful for a decision. These 5 checks expose metric noise.

Key takeaways
- 01Define each safety metric with a numerator, denominator, observation period, source record, and accountable owner before comparing results.
- 02Compare equivalent populations and display at least 2 context fields so leaders do not mistake arithmetic consistency for fair interpretation.
- 03Match review cadence to risk speed, using shift, weekly, and monthly levels when a delayed dashboard could hide a same-day control failure.
- 04Connect every indicator to 4 decisions about review, escalation, evidence verification, and corrective-action ownership instead of treating thresholds as decoration.
- 05Audit your safety culture and measurement discipline with Andreza Araújo’s Safety Culture Diagnosis: Learn how to do your own before expanding the dashboard.
Safety data quality is the degree to which a safety number can be traced to a defined event, compared across equivalent populations, reviewed within a useful time window, challenged with field evidence, and connected to a decision. A dashboard can contain 20 metrics and still provide weak evidence if those five conditions are missing.
A monthly safety meeting often begins with a polished chart, yet the difficult question arrives immediately afterward: what should the leadership team do differently because this number changed? If nobody can answer without reopening spreadsheets, the problem is not visual design. The problem is data quality.
ISO 45001:2018 requires monitoring, measurement, analysis, and evaluation of OH&S performance, but the standard does not turn every recorded value into a decision-grade indicator. The organization still has to define the event, protect the denominator, preserve the source, and assign an owner.
Safety data quality means that a reported safety value remains understandable from source to decision. The value has a clear definition, a known population, a reliable collection method, a stated review period, and an action that follows when the signal crosses an agreed threshold.
What does safety data quality mean?
Quality is not the same as volume. A plant that records 500 observations may have less useful evidence than a plant that records 50, because the larger dataset may mix coaching conversations, hazard reports, duplicate entries, and closed actions under one label.
Andreza Araújo’s work on safety culture places measurement inside the operating system of the organization. In Safety Culture: From Theory to Practice, measurement becomes meaningful when it helps reveal how the culture is actually operating, rather than when it simply increases the number of reports.
Five checks make that distinction practical. Each check asks whether the number can support a decision without requiring the team to guess what it represents.
Test 1: Can the team trace the number?
A traceable metric has one definition, one source, and one owner. If two analysts calculate the same indicator from different event lists, the organization does not have a measurement problem alone. It has a control problem.
Write the metric definition in 3 parts: numerator, denominator, and observation period. For example, a near-miss rate may use accepted near-miss reports as the numerator, hours worked as the denominator, and a calendar month as the period. The exact formula can change by organization, but the terms cannot remain implicit.
Traceability also requires a source record. Keep the original report, inspection, permit, or system entry for at least the retention period that applies to the organization. A result that cannot be reconstructed after 30 days is difficult to defend during an audit or an incident review.
Test 2: Is the comparison fair?
A number becomes misleading when the comparison changes the population. Comparing a day shift with a night shift may be useful, but only if the team states the differences in hours, work type, staffing, contractor exposure, and reporting access.
Use like-for-like comparisons before ranking teams. A 2% rate in a maintenance crew that performs high-energy isolation work does not carry the same meaning as a 2% rate in an office population, even though the arithmetic looks identical.
Fair comparison does not mean hiding variation. It means showing the context that changes interpretation. A dashboard should display at least 2 contextual fields beside the metric, such as exposure hours and work type, so that leaders can challenge a trend without confusing context with excuse.
Test 3: Does the timing support action?
Data can be accurate and still arrive too late. A quarterly review may be appropriate for a strategic trend, while a recurring control failure may need a 24-hour escalation path. The useful cadence depends on how quickly the signal can become harm.
Set 3 review levels. Supervisors can review point-of-work signals during the shift, EHS can examine patterns each week, and executives can review systemic exposure each month. The levels should connect, because a monthly dashboard cannot repair a control that required same-day intervention.
When a metric is delayed, label it honestly. “Closed actions last month” is not the same as “current control reliability.” The first describes completed administration, while the second requires evidence that the control still works in the field.
Test 4: Can the field challenge the number?
A dashboard should remain open to contradiction. If operators, supervisors, or maintenance teams report a changing condition that the metric does not show, the field signal deserves investigation rather than automatic dismissal.
Ask whether the collection method can capture missing reports, delayed entries, duplicate records, and work that sits outside the formal system. A value can be internally consistent while the process excludes the people or tasks that carry the greatest exposure.
Use a short sample of field records each month and compare it with the dashboard. A 10-record sample is not a statistical estimate of performance, but it can reveal a broken definition, a missing work group, or a reporting habit that needs correction.
Test 5: Does the number change a decision?
A metric earns its place when a defined change in the number leads to a defined response. If the team only discusses whether the chart is green or red, the threshold has become decoration.
For every indicator, record 4 decisions: who reviews it, what threshold triggers escalation, what evidence must be checked, and who owns the corrective action. The response does not need to be dramatic. It may be a field verification, a supervisor conversation, a control test, or a review of work allocation.
One useful test is to hide the title of the dashboard and ask a manager to explain what action follows from the trend. If the answer depends on a second meeting, the indicator is not yet decision-ready.
How to differentiate reliable data from metric noise
Use the following distinction before adding another KPI. The table separates the evidence property from the failure that weakens it.
| Evidence property | Reliable signal | Metric noise |
|---|---|---|
| Definition | One formula with named numerator, denominator, and period | Terms change between reports or teams |
| Source | Original records can be retrieved and sampled | Values are copied from an unowned spreadsheet |
| Context | Exposure, work type, and population are visible | Different populations are ranked as if equivalent |
| Timing | Review cadence matches the speed of the risk | Old data is presented as current control performance |
| Decision | Thresholds identify an owner and next verification | The chart produces discussion without a control change |
The distinction is also useful when reviewing TRIR. The rate may be correctly calculated and still fail to describe exposure to high-consequence events. Leaders should therefore read it beside leading evidence and critical-control verification, rather than asking one lagging number to carry the entire safety argument. The related guide on TRIR questions for leaders provides a useful starting point.
When should leaders use a safety data quality review?
Run a focused review before launching a new dashboard, after a reporting-system change, when two sites publish conflicting figures, or when a metric has remained green while field evidence is deteriorating. A 10-minute source check can prevent a month of confident discussion around the wrong number.
Begin with 5 questions: What event does this value represent? Who can reproduce it? Which population does it cover? How old is the evidence? What decision changes when the value moves? If the team cannot answer one of them, mark the indicator as provisional instead of presenting it as established performance.
Executives who need to connect metrics with financial decisions can also compare prevention cost, incident cost, and exposure cost in the CFO safety-cost guide. The central discipline is the same: a number should clarify responsibility, not conceal uncertainty.
What to remember about safety data quality
Safety data quality is not a software feature. It is a management discipline that protects the path from observation to decision. Use the 5 checks, document the definition, preserve the source, test the field signal, show the comparison context, and connect the result to an owner.
A smaller set of trusted indicators is often more useful than a large dashboard that nobody can interpret. When the metric is traceable, comparable, timely, and actionable, it can help leaders see risk before the next recordable event forces attention.
Andreza Araújo’s Safety Culture Diagnosis: Learn how to do your own offers a broader method for examining whether formal systems match lived practice. The same principle applies to metrics: measure what the organization can explain, then act on what the evidence shows.
Frequently asked questions
What are the 5 checks of safety data quality?
How can a safety team validate a metric definition?
How often should safety metrics be reviewed?
What is the difference between a safety KPI and a leading indicator?
Can safety data quality improve safety culture?
About the author
Andreza Araújo
Safety Culture Expert | Senior EHS Executive
Andreza Araújo is a safety culture expert and senior EHS executive with more than 25 years of experience in environment, health and safety. She is a Civil Engineer and Occupational Safety Engineer from Unicamp, holds a Master's degree in Environmental Diplomacy from the University of Geneva, and completed sustainability studies at IMD Switzerland. Andreza has served in Global Head of EHS roles in Fortune 500 environments, leading cultural transformation programs across multinational operations. She has represented Brazil as a speaker at the United Nations in Paris and has spoken at the International Labour Organization in Turin. She is the author of more than 16 books on safety culture in Portuguese, Spanish, English and German. Her work has earned more than 10 EHS awards, including two recognitions from Indra Nooyi, former PepsiCo CEO.
- Civil & Safety Engineer (Unicamp)
- M.A. Environmental Diplomacy (University of Geneva)
- Sustainability Cert (IMD Switzerland)
- People Management & Coaching (Ohio University)
- UN Paris speaker representative for Brazil
- ILO Turin speaker
- LinkedIn Top Voice
- Indra Nooyi PepsiCo CEO recognition (2x)
Documentaries
Watch Andreza's documentaries
Three productions on safety culture, organizational failure and the human lessons behind major disasters.
Podcasts
Listen to Andreza's podcasts
She hosts three shows on safety leadership, EHS and organizational culture, in English and Portuguese.