Safety Culture

Safety Culture: 5 Distortions That Make Survey Scores Look Better Than the Worksite

A safety culture survey can show perception, but not control quality. Leaders need field evidence, not averages, to judge the worksite honestly.

By 7 min read
corporate environment depicting safety culture survey distortions — Safety Culture: 5 Distortions That Make Survey Scores Loo

Key takeaways

  1. 01A survey average can flatter the site while weak pockets stay hidden in shifts, contractors, or non-routine work.
  2. 02Question wording, timing, and sampling choices can measure mood more than culture.
  3. 03The right test is not the score alone, but the score beside field evidence, correction closure, and supervisor decisions.
  4. 04Andreza Araujo's safety culture books help leaders move from perception to evidence without turning the process into theater.
  5. 05When the survey and the worksite disagree, leaders should trust the field first and then fix the loop that hid the gap.

A safety culture survey can tell leaders how people say the site feels. It cannot tell them whether the worksite actually changed when pressure rose.

That gap matters because averages often improve before the field changes, and the weakest pockets are usually in the shift, the contractor group, or the task where leaders are not looking. Across 25+ years in executive EHS and more than 250 cultural transformation projects, Andreza Araujo has seen the same pattern repeat: survey scores rise after a campaign, while field evidence stays flat until a harder event forces the truth back into view.

This article is for leaders who need evidence, not comfort. If you want the diagnostic method behind this angle, Safety Culture Diagnosis: 250-Company Case shows how the field test works. If you want the faster leadership layer, Safety Culture: From Theory to Practice explains why the score is only the first layer of the discussion.

Why the average flatters the site

The average, which compresses different shifts into one number, can hide the exact places where risk still lives. A site may look healthy on paper because the main office, day shift, and salaried staff answer with confidence, while contractors, night crews, or maintenance teams answer with caution that gets diluted in the final figure.

That is why the score is not the same thing as control. James Reason's work on latent conditions matters here because the weak layer often sits upstream of the visible result, where routine decisions, staffing pressure, and poor follow-up shape the field long before an incident appears. The score can soften the picture, but it does not remove the condition that created it.

Andreza Araujo's Safety Culture: From Theory to Practice takes the same position. Culture is not the mood of the moment. It is the pattern of decisions that people can predict when production pressure rises, which is why a survey should support a diagnosis instead of replacing it.

Distortion 1: sampling that favors comfort over exposure

A survey sample can flatter the site when it overrepresents people who are easier to reach and underrepresents the people whose work carries the most exposure. That usually means office staff answer fast, while field crews, contractors, and lower-trust groups answer late or not at all.

The result is not random noise. It is a biased snapshot. A survey, which is supposed to widen voice, becomes a mirror for the parts of the organization that already feel safe enough to speak. The crew whose work changes fastest under pressure may be the crew whose voice is least visible in the result.

Leaders should therefore break the result by shift, site, function, contractor status, and language before they make a claim. If a subgroup falls away from the average, that subgroup is not a footnote. It is the signal that the sample hid the risk.

Distortion 2: question wording that measures mood instead of culture

Question design can create a flattering answer when it asks people to report a feeling instead of a testable reality. A question, which asks whether leadership cares, measures sentiment. A question that asks whether people received the right supervision, the right stop-work support, or the right follow-up checks a control or condition that can be tested.

That difference matters because culture is behavioral. If the wording is too soft, the answer rewards politeness. If the wording is too abstract, the answer rewards interpretation. The final score then looks precise while the underlying meaning stays fuzzy.

Andreza Araujo's Safety Culture Diagnosis: Learn how to do your own works better here because it pushes the reader toward evidence that can be inspected, not just felt. In a strong diagnostic loop, the question should lead to a field check, and the field check should either confirm or challenge the survey.

Distortion 3: timing that captures the campaign, not the culture

Timing can distort the result when the survey lands right after a campaign, a leadership visit, or a serious event. People answer in the emotional weather of the moment, which means the questionnaire can measure exposure to communication instead of stability in the operating system.

That is why a score taken immediately after a visible push can look healthier than the worksite really is. The same is true after a bad event, because fear can temporarily suppress honest answers. In both cases, the survey is not wrong. It is simply reading a temporary state whose meaning leaders can misread if they treat it as durable culture.

Patrick Hudson's maturity model helps frame this problem because a calculative or proactive label means little if the organization only looks mature right after it was told to look mature. A real shift is visible when the pattern stays stable after the campaign fades.

Distortion 4: averages that hide the weak pockets

One average can make two very different realities look like one story. A site with a strong main plant and a struggling maintenance group can still report a healthy overall score, which is why the average often flatters the leaders who are closest to the survey sponsor and hides the pressure on the crews most exposed to drift.

The only way around that problem is to break the score apart. Compare the top quartile with the bottom quartile. Compare direct employees with contractors. Compare day shift with night shift. Compare the unit that answers fast with the unit that answers late. The gaps are where the diagnostic work starts.

This is the point where Edgar Schein's view of culture becomes practical, because artifacts and declarations can look aligned while the underlying assumptions stay split. A site whose average looks good but whose subgroup data is jagged is not necessarily mature. It is often inconsistent.

Distortion 5: no field trace to challenge the score

The most expensive distortion appears when the survey is never compared with field traces. If leaders do not walk the job, read the permits, review the closeout quality, and listen to the crew whose work changed last week, the score becomes a standalone story. It may be a true story. It is still incomplete.

Field traces are the reality checks that keep culture work honest. They include recent stop-work examples, the quality of corrective action closure, the speed of escalation when a plan breaks, and the behavior of supervisors whose authority is visible when work has to slow down. A survey can point to a weak area, but the field tells you whether the weakness is procedural, relational, or structural.

That is also why Andreza Araujo's The Illusion of Compliance matters here. A clean report, which looks efficient on the screen, can still sit on top of poor reality. The only reliable correction is to compare the declared result with what the worksite would show if nobody was watching.

What leaders should compare before the next review

The cleanest way to read the survey is to put it beside three other evidence types. The table below keeps the comparison practical.

Evidence type What it can hide What leaders should check
Survey average Different realities inside the same number Break it by shift, contractor status, language, and function
Survey comments Polite answers that avoid risk Look for repeated concerns, not just positive tone
Field traces Gaps that no questionnaire can see Check permits, follow-up, supervision, and stop-work support
Management review Decisions that stay verbal Ask what changed in budget, cadence, authority, or design

This is where How to Run a Safety Culture Evidence Review in 90 Minutes becomes the practical companion. The survey gives direction, but the review should force a decision about the work that caused the score in the first place.

In more than 250 cultural transformation projects, Andreza Araujo has seen that leaders move faster when the evidence is separated this way. They stop arguing with the score and start asking why the score and the field disagree. That question is more uncomfortable, which is exactly why it works.

What to do when score and worksite disagree

When the survey and the field disagree, trust the field first. The worksite is where the exposure sits, and the score only tells you how people described that exposure. If the crew, the permit, the supervisor, and the corrective action log point in the same direction, the average should not be allowed to overrule them.

The next move is not to defend the questionnaire. It is to test the loop that produced the mismatch. Ask which subgroup pulled away from the average, which question blurred sentiment and control, which timing window distorted the answer, and which field trace should have caught the gap earlier. Those questions reveal whether the problem is sampling, wording, timing, or management follow-up.

If the answer is not obvious, run a short diagnostic cycle rather than another broad survey. Walk one area, read one permit chain, review one closeout trail, and speak to one contractor crew whose voice was weak in the data. A small but exact loop gives more truth than a wide but vague one.

What the best leaders do next

The best leaders do not throw away surveys. They keep them as one instrument inside a wider diagnostic system, and they use the mismatch between the survey and the worksite to find the hidden pressure point. That is how the score becomes useful without becoming the master of the conversation.

Andreza Araujo's view is simple enough to be operational. If the survey says one thing and the field says another, the field gets the first vote. If the field keeps saying the same thing, the organization has a culture problem that cannot be solved with a prettier average. It needs a better decision path.

The strongest culture programs do not ask the questionnaire to prove the culture. They ask it to point to the place where leadership should look next. Once the field is in view, the real work begins.

Topics safety-culture culture-survey culture-diagnosis field-evidence leadership

Frequently asked questions

Can a high safety culture survey score prove the culture is healthy?
No. A high score can only prove that people answered positively on that day. Culture still has to be tested against field evidence, because control quality, supervisor decisions, and contractor reality can look very different from the average.
What should leaders compare with the survey result?
Leaders should compare the result with field traces, permit quality, corrective action closure, shift differences, contractor feedback, and recent leadership decisions. That comparison shows whether the score reflects the work or only the questionnaire.
Which Andreza Araujo book fits this topic best?
Safety Culture Diagnosis: Learn how to do your own is the best fit because it treats the survey as one input inside a wider diagnostic method. Safety Culture: From Theory to Practice also fits because it links culture to repeated decisions.
What does low participation usually mean?
Low participation usually means friction, distrust, fatigue, or low relevance. It is rarely simple indifference, because people respond more when they believe the process will change something in the work.

About the author

Andreza Araújo

Safety Culture Expert | Senior EHS Executive

Andreza Araújo is a safety culture expert and senior EHS executive with more than 25 years of experience in environment, health and safety. She is a Civil Engineer and Occupational Safety Engineer from Unicamp, holds a Master's degree in Environmental Diplomacy from the University of Geneva, and completed sustainability studies at IMD Switzerland. Andreza has served in Global Head of EHS roles in Fortune 500 environments, leading cultural transformation programs across multinational operations. She has represented Brazil as a speaker at the United Nations in Paris and has spoken at the International Labour Organization in Turin. She is the author of more than 16 books on safety culture in Portuguese, Spanish, English and German. Her work has earned more than 10 EHS awards, including two recognitions from Indra Nooyi, former PepsiCo CEO.

  • Civil & Safety Engineer (Unicamp)
  • M.A. Environmental Diplomacy (University of Geneva)
  • Sustainability Cert (IMD Switzerland)
  • People Management & Coaching (Ohio University)
  • UN Paris speaker representative for Brazil
  • ILO Turin speaker
  • LinkedIn Top Voice
  • Indra Nooyi PepsiCo CEO recognition (2x)

Documentaries

Watch Andreza's documentaries

Three productions on safety culture, organizational failure and the human lessons behind major disasters.

Podcasts

Listen to Andreza's podcasts

She hosts three shows on safety leadership, EHS and organizational culture, in English and Portuguese.

Summarize with AI