Safety Culture

Safety Culture Maturity Score: 6 Decisions That Expose a False Diagnosis

A safety culture maturity score is useful only when it changes decisions. This diagnostic guide shows leaders how to separate perception from control evidence, test bad-news response, and tie maturity claims to serious risk.

By 7 min read
corporate environment depicting safety culture maturity score 6 decisions that expose a false diagnosis — Safety Culture Matu

Key takeaways

  1. 01Define exactly what a culture maturity score can claim, because a survey result does not prove that a critical control works.
  2. 02Separate perception evidence from observed control evidence so a favorable climate does not hide weak execution.
  3. 03Trace what happens after bad news, including whether the person who raised the concern receives an answer and whether the exposure changes.
  4. 04Segment results by work group and decision environment before an enterprise average hides a local vulnerability.
  5. 05Tie maturity claims to the organization’s serious exposures and define the evidence that can lower the score.

F1 critical diagnostic for safety leaders, EHS directors, and plant managers

A culture maturity score can look precise while hiding the decision that matters most: whether people can rely on the organization to recognize exposure and change the work. A high score does not prove that a supervisor will stop a task, that a manager will fund a control, or that a worker can challenge a production plan without losing credibility.

This is why a culture diagnosis should be treated as a management decision system rather than a polished survey result. ISO 45001:2018 asks organizations to evaluate performance and improve the occupational health and safety management system, while ISO 45003:2021 extends the conversation to psychosocial conditions. Neither standard turns a single score into proof of control reliability.

A culture maturity score is useful when it helps leaders decide what must change, who owns that change, and what evidence will show that the change is holding. It becomes misleading when the score rewards agreement, averages away local weakness, or is detached from the serious exposures that the operation must control.

Why a culture maturity score can mislead leaders

Maturity models simplify a complex organization so that leaders can compare patterns, set priorities, and discuss progress. The simplification is valuable, although it creates a measurement risk. The number can become the object of management attention, while the conditions behind it disappear from view.

Patrick Hudson’s maturity model is often used to describe movement from reactive behavior toward more proactive management. Andreza Araujo’s Safety Culture: From Theory to Practice makes the same practical issue visible from another angle, because a label only matters when it changes leadership habits, operational choices, and the quality of prevention.

A plant can therefore improve its declared maturity without improving its most important barrier. If leaders celebrate a higher score before asking what happened to overdue actions, weak permit decisions, or repeated workarounds, the diagnosis has become a reputation exercise.

Decision 1: Define what the score is allowed to claim

Before collecting responses, write one sentence that limits the score’s meaning. It may describe how employees perceive leadership commitment, how consistently supervisors apply a management process, or how credible the speak-up response feels. It should not claim that the organization is safe in general.

This boundary matters because a culture score is not an incident rate, a control-verification result, or a substitute for a legal audit. The score can illuminate conditions that influence decisions, but it cannot prove that a guard is installed correctly or that an emergency plan works under pressure.

Leaders should also state what the score will not be used for. If it will not rank individual supervisors, determine bonuses, or close corrective actions, say so before the survey opens. People answer more honestly when the use of their information is visible, which makes the diagnosis more useful.

James Reason’s work on organizational accidents provides a sound test here. Human decisions occur inside systems that contain latent conditions, so the score should point leaders toward the conditions shaping behavior rather than invite a simple judgment about whether people have the right attitude.

Decision 2: Separate perception from observed control

Perception is important because people experience the management system through daily choices, conversations, and consequences. It is not the same as observed control. A team may report that supervisors care about safety while a field review shows that temporary barriers are rarely checked after a shift change.

Use two evidence streams and keep them visibly separate. The first can include survey responses, interviews, and structured conversations. The second should include samples of control verification, action closure, permit quality, maintenance backlog, change management, and stop-work decisions. Combining them too early creates a blended score that no one can interpret.

The safety climate tests that distinguish declared commitment from field reality can support this comparison. The practical question is not whether the two streams produce the same number. The question is where they disagree and what that disagreement reveals about the work system.

When perception is high and observed control is weak, leaders should investigate whether communication is stronger than execution. When perception is low and control evidence is strong, they should examine whether the controls are invisible, poorly explained, or applied inconsistently.

Decision 3: Test what happens after bad news

A mature culture is not demonstrated by positive answers to “management listens.” It is demonstrated by what happens after someone reports an inconvenient condition. The response may include a technical correction, a change in the work plan, a clear explanation, or an escalation when the first owner lacks authority.

Sample recent concerns and trace them from first report to decision. Check whether the original information survived the handoffs, whether the person who raised it received an answer, and whether the final action changed the exposure. A closed record without a changed condition is weak evidence, even when the system shows perfect completion.

This test also protects psychological safety from becoming a slogan. Amy Edmondson’s research describes psychological safety as a condition that supports interpersonal risk-taking, not unrestricted acceptance of every action. In safety management, the relevant question is whether people can raise a concern and still be held to clear technical and behavioral standards.

A score that rises while bad-news records become shorter, less specific, or harder to trace should not be celebrated. The organization may be measuring agreement with the survey rather than trust in the response.

Decision 4: Compare work groups before averaging the result

An enterprise average can hide the area where serious exposure is concentrated. A corporate result of “proactive” may combine a well-supported engineering group with a contractor interface where production pressure, language barriers, or unstable staffing make control decisions less reliable.

Segment the diagnosis by work design and decision environment, not only by department. Compare shifts, contractor interfaces, maintenance, logistics, remote operations, and teams undergoing major change. The segmentation should be large enough to protect confidentiality and specific enough to reveal a management pattern.

Do not treat the lowest score as proof that a group is the problem. It may be the group with the clearest view of an unresolved exposure. A low score becomes a leadership question when it identifies where workers encounter contradictory instructions, delayed resources, or weak follow-through.

During 25+ years leading EHS in multinational environments, Andreza Araujo has built her authority around the distinction between a system that looks aligned and one that changes decisions across different operating contexts. That distinction is especially important when an average score is being presented to a global leadership team.

Decision 5: Tie maturity claims to the risk profile

Culture maturity should be interpreted against the hazards that the organization must control. A score that is acceptable for routine office work may be inadequate for a site managing high-energy isolation, mobile equipment interaction, confined space entry, or major process hazards.

This does not mean that leaders need a dramatic new index. It means they should ask whether the diagnosis examines the behaviors and decisions that matter for the credible consequence. For a high-consequence exposure, the assessment should test authority to stop work, quality of critical control checks, management of change, contractor coordination, and response to weak signals.

ISO 45001:2018 requires organizations to consider hazards, risks, opportunities, and operational controls within their management system. A maturity label that ignores those links is too broad to guide a serious decision.

The SIF exposure review offers a useful companion lens because it keeps consequence visibility separate from the general climate conversation. Leaders should not let a favorable culture score reduce attention to a rare but credible fatality pathway.

Decision 6: Decide what evidence can change the score

A diagnosis earns credibility when people know how evidence can move the result. If the score changes only after another survey, the organization has created a measurement cycle instead of a learning cycle. Define evidence that can confirm, challenge, or revise the initial interpretation.

Useful evidence may include repeated field observations, control-verification samples, action quality, escalation records, worker follow-up, and management decisions made under operational pressure. The evidence should be reviewed by people who can challenge the first conclusion, because the team that designed the score may be too invested in defending it.

Andreza Araujo’s Safety Culture Diagnosis: Learn How to Do Your Own supports a practical discipline. Diagnosis is not a ceremony that produces a maturity badge. It is a structured way to identify patterns, test assumptions, and choose interventions that fit the organization’s actual conditions.

Set a review date and define what would count as disconfirming evidence. If leaders cannot name the observation that would make them lower the score, the instrument is probably measuring confidence rather than culture.

What a credible diagnosis should produce

A credible diagnosis should leave leaders with fewer vague statements and more owned decisions. It should identify which conditions help people make safe choices, which conditions push them toward workarounds, and which decisions require authority above the local supervisor.

Each priority should have an accountable owner, a change to the work or management process, a verification method, and a date for review. “Improve communication” is not an intervention. The organization must specify what information travels, who acts on it, and how the field will show that the response changed.

The distortions that make compliance look like control are useful to revisit before the leadership review. A completed training record, a signed procedure, or a higher survey score may show activity, but none is sufficient evidence that the exposure has been reduced.

Leaders can use the result as a priority map, not as a trophy. That shift preserves the value of the score while preventing the number from becoming a substitute for control decisions.

How leaders should present the result to the board

Board and executive audiences need a concise explanation of what the diagnosis can support and where uncertainty remains. Present the maturity result beside the most important disagreement between perception and observed control, the work group with the clearest vulnerability, and the decisions that need executive ownership.

Include one example of evidence that changed the interpretation. That example shows that the process is capable of correction rather than designed to confirm a preferred story. It also gives directors a sharper question than “Are we mature?” They can ask which exposure remains difficult to control and what authority is missing.

Andreza Araujo’s Make The Difference: Be a Leader in Health & Safety frames leadership as operational responsibility, which means the executive conversation should end with a decision, not a maturity adjective.

Turn diagnosis into decisions that hold. Explore Andreza Araujo’s books and Safety School resources for practical methods that connect culture, leadership, and prevention.

Topics safety culture culture maturity safety culture diagnosis leadership decisions control verification

Frequently asked questions

What is a safety culture maturity score?
A safety culture maturity score is an assessment result that describes selected patterns in leadership, decision-making, worker voice, and management practice. It is useful when it directs owned changes and verification. It does not prove that a specific safeguard works or that the workplace is safe in general.
Can a high culture maturity score still hide serious risk?
Yes. A high score can hide serious risk when the assessment averages different work groups, measures agreement instead of response quality, or remains separate from control verification, management of change, and high-consequence exposure reviews.
How should leaders validate a culture diagnosis?
Leaders should compare survey and interview results with field observations, control-verification records, action quality, escalation decisions, and follow-up after reported concerns. They should also define what evidence would challenge the original score.
What is the difference between safety climate and safety culture?
Safety climate describes how people currently perceive safety priorities and management behavior. Safety culture is broader and concerns the patterns of decisions, assumptions, practices, and responses that shape how work is organized and controlled over time.
How often should a safety culture maturity score be reviewed?
Review timing should follow the decision purpose and the pace of organizational change. A repeat survey without evidence from the work will only reproduce the same uncertainty, so leaders should review priority actions and field evidence between formal assessments.

About the author

Andreza Araújo

Safety Culture Expert | Senior EHS Executive

Andreza Araújo is a safety culture expert and senior EHS executive with more than 25 years of experience in environment, health and safety. She is a Civil Engineer and Occupational Safety Engineer from Unicamp, holds a Master's degree in Environmental Diplomacy from the University of Geneva, and completed sustainability studies at IMD Switzerland. Andreza has served in Global Head of EHS roles in Fortune 500 environments, leading cultural transformation programs across multinational operations. She has represented Brazil as a speaker at the United Nations in Paris and has spoken at the International Labour Organization in Turin. She is the author of more than 16 books on safety culture in Portuguese, Spanish, English and German. Her work has earned more than 10 EHS awards, including two recognitions from Indra Nooyi, former PepsiCo CEO.

  • Civil & Safety Engineer (Unicamp)
  • M.A. Environmental Diplomacy (University of Geneva)
  • Sustainability Cert (IMD Switzerland)
  • People Management & Coaching (Ohio University)
  • UN Paris speaker representative for Brazil
  • ILO Turin speaker
  • LinkedIn Top Voice
  • Indra Nooyi PepsiCo CEO recognition (2x)

Documentaries

Watch Andreza's documentaries

Three productions on safety culture, organizational failure and the human lessons behind major disasters.

Podcasts

Listen to Andreza's podcasts

She hosts three shows on safety leadership, EHS and organizational culture, in English and Portuguese.

Summarize with AI