Safe Behavior

How to Calibrate Safety Observations Across Shifts in 21 Days

A practical 21-day method for making safety observations comparable across crews, supervisors, and shifts, so observation data changes decisions instead of becoming a volume target.

By 6 min read
Supervisor and worker calibrating a safety observation across shifts

Key takeaways

  1. 01Observation volume does not show quality when supervisors use different definitions of safe and at-risk work.
  2. 02A useful calibration process defines observable criteria, practices on the same examples, and checks agreement in the field.
  3. 03The first 21 days should end with fewer scoring disputes, clearer coaching, and a decision rule for recurring exposure.
  4. 04Andreza Araujo's experience across 250+ cultural transformation projects supports a practical test: the observation must change a decision or remove a barrier.
  5. 05The process should protect workers from blame while still making weak controls, unclear expectations, and repeated drift visible.

Two supervisors can watch the same lifting task and record opposite conclusions. One sees a worker who skipped a step. The other sees a poor layout that makes the step awkward and a crew that has adapted to keep production moving. If both observations enter the same dashboard, the number looks precise while the decision signal is weak.

Safety observations become useful when different shifts describe comparable work with comparable language. That does not mean forcing every task into a rigid script. It means agreeing on what can be seen, what must be verified, and what should trigger a change in the work system. A 21-day calibration cycle gives a supervisor team enough time to establish that agreement without turning observation into another paperwork campaign.

What to prepare before day one

Choose one work family for the first cycle, such as mobile equipment interaction, line clearance, manual handling, or permit handover. Do not try to calibrate every task at once, because the team will spend its energy debating categories rather than improving the quality of one decision stream.

Bring together observers from different shifts, including at least one experienced operator or worker representative. The group needs a short observation standard, two or three recent examples with identifying details removed, and a record of what happened after each observation. Across 25+ years leading EHS programs, Andreza Araujo has repeatedly emphasized that culture is visible in the decisions that follow a conversation, not in the number of forms completed.

Define the output before the cycle starts. The team should finish with a shared rubric, a field agreement check, a list of recurring barriers, and one rule for escalating a pattern. As Andreza Araujo argues in Safety Culture: From Theory to Practice, compliance matters, although a compliant record that does not change the work remains a weak control.

Step 1: Define the observation question

Write one question that every observer will answer. For a lifting task, it might be, “Can the load be controlled throughout the movement without placing a person in the line of fire?” The question should direct attention to exposure and control, rather than inviting a general opinion about whether the worker looked careful.

Test the question against three recent tasks. If observers need long explanations before they can answer, narrow the wording. Verification is complete when two people can read the question and identify the same work element without asking what the program intended to measure. A common error is writing a question such as “Was the procedure followed?” before checking whether the procedure matches the physical conditions.

Step 2: Separate what was seen from what was assumed

Create three fields in the observation form: visible condition, worker action, and missing or ineffective control. This separation prevents an observer from turning an interpretation into a fact. “No gloves” is visible. “Worker does not care about safety” is an assumption that the observation cannot prove.

Ask the team to rewrite five old observations using the three fields. Verify that another observer can understand the event without knowing the person involved. If the record contains labels about attitude, motivation, or character, remove them. James Reason's work on active and latent failures supports this discipline because the visible action may be only one layer in a larger set of conditions.

Step 3: Build a small scoring rubric

Use no more than four outcomes for the first cycle: control effective, control needs reinforcement, control absent or ineffective, and not observable. Each outcome needs a plain-language definition and one example. A rubric that contains ten shades of performance may appear sophisticated while producing disagreements that nobody can resolve in the field.

Verification requires each observer to classify the same six examples independently, then explain the evidence behind the classification. Do not reward the person who gives the most severe score. The objective is a defensible classification that leads to the right next action. A common error is treating “not observable” as failure, which pressures observers to guess when the work was not actually visible.

Step 4: Practice on the same evidence

Run a 30-minute practice session with photographs, short videos from approved internal sources, or anonymized descriptions from the selected work family. Everyone must classify the evidence before the group discusses it. The facilitator should record where the disagreement occurs, because the disagreement often reveals an unclear criterion rather than a careless observer.

Repeat the exercise with a second set that contains a real tradeoff, such as a control that works in normal conditions but becomes difficult during a changeover. Verification is achieved when observers can state which fact changed their classification. The error to avoid is allowing the most senior person to speak first, since the rest of the group may align with authority instead of evidence.

Step 5: Test the rubric at the shift boundary

Pair observers from different shifts and ask them to watch the same task at a handover or another operational transition. The point is not to inspect a worker twice. The point is to see whether the observation language survives a change in crew, lighting, workload, supervisor presence, or production pressure.

Compare the records immediately after the task. Verify that differences are explained by changed conditions, not by different standards. If the day shift writes “control needs reinforcement” and the night shift writes “control absent,” the team must identify the factual distinction or revise the rubric. A common error is calibrating only in a meeting room, where the pressure that shapes behavior is absent.

Step 6: Turn observations into coaching decisions

Every observation should end with a defined response. The response may be a worker conversation, a supervisor clarification, a maintenance request, a task redesign, or escalation to the risk owner. The observer should record the decision and the owner, not just the category.

Use one question during the conversation: “What made the safer action easier or harder here?” That question keeps the discussion close to the work while still holding the team accountable for the control. In Sorte ou Capacidade, Andreza Araujo examines why visible outcomes can be mistaken for individual ability; the same caution applies when a single observation is used to judge a worker without testing the surrounding conditions.

Step 7: Review agreement and recurring barriers

At the end of the second week, compare a sample of observations from every participating shift. Do not focus only on agreement percentages. Review the disagreements, the control owners, and the time taken to close each action. A high agreement rate can still hide a poor rubric if everyone is recording the same shallow question.

Verification is complete when the team can name its three largest sources of disagreement and connect each one to a wording change, a training need, or a control problem. A common error is correcting the observer before checking whether the work standard is realistic. If the same workaround appears across shifts, the program has found a system issue, not merely a coaching opportunity.

Step 8: Set the 21-day decision rule

On day 21, agree on what happens when the same exposure appears repeatedly or when a control receives different classifications across shifts. For example, three comparable observations of an ineffective control can trigger a risk-owner review, while repeated disagreement can trigger a field verification by the supervisor and the worker representative.

Publish the rule with the rubric and show one completed example. Verify that a supervisor on each shift can explain who owns the next decision, when it is due, and what evidence closes it. A common error is ending the cycle with a report that describes the problem but gives nobody authority to change the work.

How to keep calibration alive after day 21

Repeat a short cross-shift comparison whenever the task changes, a serious near miss occurs, a new supervisor joins, or the observation categories begin to drift. The review can be brief when the standard is stable. It should become longer when observers disagree about a critical control or when workers stop offering context.

In more than 250 cultural transformation projects supported by Andreza Araujo, the durable pattern is not a permanent campaign. It is a management rhythm that makes weak signals visible and assigns a response. That is why the final measure should not be observation volume. It should be whether the program removed a barrier, clarified an expectation, improved a control, or escalated a risk before harm occurred.

Safety is about coming home, and observation quality matters because it determines what the organization notices before someone is hurt. If your team needs a practical way to connect field conversations with leadership decisions, explore the resources from Andreza Araujo and the English safety blog.

  • Choose one work family and one observation question.
  • Separate visible facts from assumptions about attitude.
  • Use a four-outcome rubric with examples.
  • Practice on shared evidence before field use.
  • Test the standard across a real shift boundary.
  • Record the decision owner after every observation.
  • Review disagreements as evidence about the system.
  • Set an escalation rule before the 21-day cycle ends.
Topics safe-behavior behavioral-observations safety-coaching supervision shift-handover leading-indicators

Frequently asked questions

What does calibration mean in a safety observation program?
Calibration means aligning observers on what counts as safe, at-risk, or not observable, then checking that they apply the same criteria to comparable work. It is a quality-control process for the observation system, not a contest between supervisors.
How long does it take to calibrate safety observations across shifts?
A focused 21-day cycle is enough to define the standard, practice with shared examples, test agreement in the field, and correct the largest scoring differences. The cycle should then become a recurring review rather than a one-time campaign.
Should safety observations be used to discipline workers?
Observations should first expose conditions, decisions, and barriers that make safe work difficult. A repeated deliberate violation may require a separate management process, but mixing discipline into every observation reduces trust and weakens the quality of the evidence.
What is the most useful result of a calibrated observation program?
The most useful result is a reliable decision signal. Leaders can see which exposures require engineering change, which expectations need clarification, which supervisors need support, and which recurring patterns deserve escalation.

About the author

Andreza Araújo

Safety Culture Expert | Senior EHS Executive

Andreza Araújo is a safety culture expert and senior EHS executive with more than 25 years of experience in environment, health and safety. She is a Civil Engineer and Occupational Safety Engineer from Unicamp, holds a Master's degree in Environmental Diplomacy from the University of Geneva, and completed sustainability studies at IMD Switzerland. Andreza has served in Global Head of EHS roles in Fortune 500 environments, leading cultural transformation programs across multinational operations. She has represented Brazil as a speaker at the United Nations in Paris and has spoken at the International Labour Organization in Turin. She is the author of more than 16 books on safety culture in Portuguese, Spanish, English and German. Her work has earned more than 10 EHS awards, including two recognitions from Indra Nooyi, former PepsiCo CEO.

  • Civil & Safety Engineer (Unicamp)
  • M.A. Environmental Diplomacy (University of Geneva)
  • Sustainability Cert (IMD Switzerland)
  • People Management & Coaching (Ohio University)
  • UN Paris speaker representative for Brazil
  • ILO Turin speaker
  • LinkedIn Top Voice
  • Indra Nooyi PepsiCo CEO recognition (2x)

Documentaries

Watch Andreza's documentaries

Three productions on safety culture, organizational failure and the human lessons behind major disasters.

Podcasts

Listen to Andreza's podcasts

She hosts three shows on safety leadership, EHS and organizational culture, in English and Portuguese.

Summarize with AI