New Reliability Manager in 45 Days: How to Turn Deferred Maintenance Into a Safety Escalation Plan
A new reliability manager should not begin with a larger backlog report. This 45-day plan connects deferred maintenance, critical controls, escalation rights, and field evidence.
Key takeaways
- 01Rank deferred maintenance by potential harm and control dependency, not backlog age alone.
- 02Separate routine backlog work from safety-critical decisions that require named acceptance and expiry dates.
- 03Set escalation triggers that operators, technicians, supervisors, and production leaders can use during real work.
- 04Verify completed repairs in the field so closure records do not substitute for control performance.
- 05Build your next 45-day leadership routine with Andreza Araujo's practical safety resources.
A new reliability manager can inherit thousands of open work orders and still miss the maintenance decisions that create the greatest safety exposure. The first 45 days should therefore produce a safety escalation plan, not only a cleaner backlog.
The central question is not how quickly the manager can reduce overdue tasks. It is which deferred conditions can weaken a critical control, who has authority to accept the exposure, and what evidence will show that the risk has actually moved.
A reliability manager should rank deferred maintenance by potential harm, connect each high-consequence item to an accountable decision owner, set escalation time limits, and verify the resulting control in the field. This approach prevents backlog age from becoming a misleading substitute for risk control.
What the reliability manager must own before starting
The reliability manager owns the link between equipment condition and operational risk, even when another leader owns production, engineering, or EHS. That ownership does not mean approving every repair. It means making the exposure visible, clarifying the decision rights, and refusing to let a technical backlog hide a safety-critical dependency.
Begin by asking three questions of the plant manager, maintenance planner, operations leader, and EHS manager. Which failures could remove a protection layer? Which temporary measures are being treated as permanent? Which decisions are being delayed because no one wants to stop production?
Use the first week to establish a shared definition of safety-significant maintenance. It may include an interlock that is bypassed, a firewater pump with an unresolved defect, a guarding repair that has been deferred, or an isolation point whose condition makes the permit process unreliable. The definition should fit the site rather than copy a generic list.
Andreza Araujo has spent more than 25 years in multinational EHS and safety leadership roles, where the practical test of leadership is whether a concern reaches the person who can change the condition. Her book Make The Difference: Be a Leader in Health & Safety reinforces that operational leaders must translate concern into a decision, an owner, and a follow-up.
Days 1 to 10: Map the maintenance exposures that can hurt people
The first 10 days should create an exposure map that shows where deferred maintenance can defeat prevention, detection, containment, or recovery. A list sorted only by age or cost will miss failures that are recent, poorly documented, or connected to a critical operating condition.
Pull the open backlog, temporary repair register, impairment log, bypass register, overdue inspection list, and recent incident investigations. Compare those records with what operators and technicians see during the shift. The gaps between the records and the field are often more valuable than either source by itself.
Classify each item by consequence, control dependency, exposure frequency, and time to safe recovery. A six-month-old defect on a rarely used noncritical asset may deserve attention, yet a two-day-old interlock bypass on a frequently accessed machine may require immediate escalation. The date is a clue, not the decision.
Then connect each high-priority item to the work process that depends on it. A pump defect may affect emergency response. A failed sensor may change the conditions under which a permit is valid. A damaged barrier may leave production technically possible while making the intended safe method impractical.
Days 11 to 20: Separate backlog pressure from safety-critical decisions
Between days 11 and 20, the manager should separate ordinary backlog management from safety-critical decisions that require named acceptance, a time limit, and visible compensating measures. This prevents urgent production requests from quietly becoming risk acceptance without an accountable signatory.
Create two review lanes. The first lane handles routine reliability work through normal planning and prioritization. The second lane handles conditions that can change a critical control, emergency function, isolation step, or exposure to serious harm.
The second lane needs a decision record with five fields: the condition, the credible consequence, the temporary control, the decision owner, and the expiry date. If the condition cannot be made safe within the agreed period, the escalation should move to the next level rather than renew itself automatically.
This is where the critical control register can help. It gives the reliability manager a way to connect an equipment condition to the control that should prevent harm, instead of discussing maintenance in isolation from the work it protects.
Quantify the review rhythm. A daily review may cover new impairments, a twice-weekly meeting may address items approaching expiry, and a weekly leadership review may examine repeat deferrals. The cadence should be small enough to sustain and serious enough to prevent silent drift.
| Backlog view | Safety escalation view |
|---|---|
| Age of the work order | Potential consequence and control dependency |
| Repair cost | Exposure created if the repair remains deferred |
| Planner ownership | Named decision owner with authority to accept or remove the risk |
| Completion date | Field evidence that the control works after completion |
Days 21 to 30: Set escalation rules that production can understand
By day 30, the site should know exactly when a maintenance condition leaves the normal planning process and enters leadership escalation. Clear triggers reduce negotiation at the moment when production pressure, incomplete information, and concern about delay are most likely to collide.
Write the triggers in operational language. Escalation may be required when a critical control is unavailable, when a temporary measure has reached its expiry, when the same defect returns after repair, when a safety function cannot be tested, or when the work instruction no longer matches the equipment condition.
Each trigger should answer four practical questions. Who must be informed? Who can authorize continued operation? What temporary protection is required? At what point does the work stop or the asset leave service?
Do not make the reliability manager the only route. A technician, operator, or supervisor who sees the condition should be able to raise it without first proving the entire causal story. The manager's role is to make the response disciplined, not to create a permission barrier around reporting.
Andreza Araujo's experience across 250+ cultural transformation projects is relevant here because escalation rules become part of culture only when they shape ordinary decisions. A rule that exists in a procedure but cannot be used during a night shift is not a reliable control.
Days 31 to 45: Verify that the repaired control works in the field
The final 15 days should test whether completed maintenance changed the exposure at the point of work. Closing a work order proves that a task was recorded. It does not prove that the barrier is available, understood, usable, and effective under operating conditions.
Choose a sample of completed safety-significant repairs and observe the task or operating condition they protect. Check the physical result, the alarm or interlock response, the isolation point, the updated drawing or instruction, and the operator's understanding of the change.
Include at least one item that was closed quickly and one that was deferred more than once. Repeat deferral is evidence that the original decision did not remove the constraint, which means the next review should examine design, resources, access, competence, or planning rather than simply assign another due date.
Use a simple evidence standard. The control should be present, available when needed, understood by the people who rely on it, and supported by a record that matches the field. If any one of those conditions fails, the maintenance action is incomplete from a safety perspective.
The control assurance guide provides a useful distinction between an audit, a check, and field evidence. For a new reliability manager, that distinction keeps verification focused on whether the control performs rather than whether the paperwork looks finished.
Common mistakes that weaken a new manager's first quarter
The most damaging mistakes are not technical errors alone. They are management choices that make a visible backlog look safer while leaving the underlying exposure unchanged.
- Ranking by age alone. Old work is not automatically the most dangerous work, and newer defects can affect a control that protects people every shift.
- Accepting temporary measures without expiry. A bypass, barricade, additional inspection, or manual check becomes a hidden permanent condition when no one owns its removal date.
- Measuring closure instead of control performance. A completed task can still leave a failed function, an unusable procedure, or an operator who was never briefed.
- Escalating only after failure. The manager should escalate when the control is weakening, not wait for an injury or serious near miss to make the condition credible.
- Keeping EHS outside the technical review. Maintenance and safety describe different parts of the same exposure, which is why the decision should bring both perspectives together.
Another trap is treating communication as a substitute for capacity. If the team has no access window, spare part, competent resource, or isolation opportunity, a stronger message will not repair the constraint. The decision must identify what the operation needs to make the safe method possible.
Resources to deepen the reliability manager's leadership practice
A new reliability manager should deepen technical knowledge and leadership judgment together, because deferred maintenance becomes a safety problem at the point where someone decides what can wait. The most useful resources connect equipment evidence with behavior, authority, and organizational culture.
Start with Make The Difference: Be a Leader in Health & Safety, which is written for leaders who need to move from intention to visible action. Safety Culture: From Theory to Practice is also useful when the manager needs to understand why repeated deferrals, weak escalation, and silent workarounds reveal the culture that the operation is actually rewarding.
Andreza Araujo's work across 30+ countries and 250+ companies supports a practical conclusion: leadership is visible in the conditions people are allowed to leave unresolved. The reliability manager who makes those conditions discussable, assigns decision rights, and verifies the result can improve safety without turning maintenance into a separate administrative campaign.
What should change after day 45?
After 45 days, the reliability manager should be able to show more than a backlog percentage. The site should have a ranked list of safety-significant deferrals, clear escalation triggers, named decision owners, expiry dates for temporary measures, and field evidence for completed repairs.
The practical decision is simple even when the work is difficult. If a deferred condition can weaken a control that protects people, it belongs in a leadership conversation before it becomes an incident. That is how reliability management becomes part of safety leadership rather than a separate maintenance report.
For readers who want to connect this plan to handover discipline, the shift handover guide shows how unresolved conditions can transfer between teams without a clear owner.
Andreza Araujo's safety leadership message is consistent with that operating principle. Make the risk visible, give someone the authority to act, and return to the field to confirm that the decision changed the work. Continue building practical safety leadership with Andreza Araujo.
Frequently asked questions
What should a new reliability manager do first for safety?
How can a reliability manager prioritize deferred maintenance?
Who should accept the risk of a deferred safety repair?
How long should a temporary maintenance control remain in place?
Does closing a maintenance work order prove that the safety risk is controlled?
About the author
Andreza Araújo
Safety Culture Expert | Senior EHS Executive
Andreza Araújo is a safety culture expert and senior EHS executive with more than 25 years of experience in environment, health and safety. She is a Civil Engineer and Occupational Safety Engineer from Unicamp, holds a Master's degree in Environmental Diplomacy from the University of Geneva, and completed sustainability studies at IMD Switzerland. Andreza has served in Global Head of EHS roles in Fortune 500 environments, leading cultural transformation programs across multinational operations. She has represented Brazil as a speaker at the United Nations in Paris and has spoken at the International Labour Organization in Turin. She is the author of more than 16 books on safety culture in Portuguese, Spanish, English and German. Her work has earned more than 10 EHS awards, including two recognitions from Indra Nooyi, former PepsiCo CEO.
- Civil & Safety Engineer (Unicamp)
- M.A. Environmental Diplomacy (University of Geneva)
- Sustainability Cert (IMD Switzerland)
- People Management & Coaching (Ohio University)
- UN Paris speaker representative for Brazil
- ILO Turin speaker
- LinkedIn Top Voice
- Indra Nooyi PepsiCo CEO recognition (2x)
Documentaries
Watch Andreza's documentaries
Three productions on safety culture, organizational failure and the human lessons behind major disasters.
Podcasts
Listen to Andreza's podcasts
She hosts three shows on safety leadership, EHS and organizational culture, in English and Portuguese.