Premises and facilities management

Reactive maintenance: how to run it well, and keep it in proportion

Reactive maintenance is a maintenance strategy in which equipment, building systems and other assets are repaired only after they break down or fail, rather than on a preventive schedule.

By SiteClaraPublished 13 minute read

A maintenance technician kneeling under a sink in an office restroom, tightening a pipe fitting with a wrench beside an open toolbox.

It is the restroom leak, the rooftop unit that trips, the door that will not latch. Every facility needs some, and too much of it is a sign that something else is wrong. This guide covers what the federal guidance and OSHA say, how to prioritize work orders and set response times, what reactive work really costs, how to manage vendors on site, and how to close the loop with the person who called it in.

01

What reactive maintenance is, and the duty behind it

Reactive maintenance is maintenance carried out in response to a failure: a clogged restroom drain, a rooftop unit that will not cool, a card reader that will not release. It is also called corrective, breakdown or unplanned maintenance, and its most urgent form is emergency maintenance, where the failure creates an immediate risk to people, the building or its security.

The Federal Energy Management Program's O&M Best Practices Guide, Release 3.0, Chapter 5: Types of Maintenance Programs describes reactive maintenance as "basically the 'run it till it breaks' maintenance mode," in which "no actions or efforts are taken to maintain the equipment as the designer originally intended to ensure design life is reached." The chapter sets it beside three other approaches:

  • Preventive maintenance: actions performed on a time- or run-based schedule to detect, preclude or mitigate degradation and extend useful life. See our preventive maintenance checklist.
  • Predictive maintenance: measurements that detect the onset of degradation, so work is based on the actual condition of the equipment rather than a calendar.
  • Reliability centered maintenance: a systematic approach that recognizes not all equipment is equally important to the facility or to safety, relies heavily on predictive maintenance, and accepts that equipment which is inexpensive and unimportant to reliability may best be left to a reactive approach.

Inside reactive work sits planned run-to-failure: a deliberate decision to let a non-critical item fail and then replace it. A light fixture in a storage closet is a fair example. A fire alarm control panel is not.

No federal law sets how much reactive maintenance a building may have. The duty behind it is the employer's. Section 5(a)(1) of the Occupational Safety and Health Act of 1970, Section 5, Duties requires each employer to furnish a place of employment "free from recognized hazards that are causing or are likely to cause death or serious physical harm." Two OSHA general industry standards bear directly on what happens after a failure:

  • 29 CFR 1910.22, General requirements (Subpart D, Walking-working surfaces), paragraph (d)(2): "Hazardous conditions on walking-working surfaces are corrected or repaired before an employee uses the walking-working surface again. If the correction or repair cannot be made immediately, the hazard must be guarded to prevent employees from using the walking-working surface until the hazard is corrected or repaired."
  • 29 CFR 1910.37, Maintenance, safeguards, and operational features for exit routes, paragraph (a)(4): "Safeguards designed to protect employees during an emergency (e.g., sprinkler systems, alarm systems, fire doors, exit lighting) must be in proper working order at all times." Paragraph (d)(2) adds that during repairs or alterations, employees must not occupy a workplace unless the exit routes are available and existing fire protections are maintained.

Check the state before relying on federal text. According to OSHA's State Plans page, 22 State Plans cover both private sector and state and local government workers, and seven cover only state and local government workers. The page says State Plans "must be at least as effective as OSHA in protecting workers and in preventing work-related injuries, illnesses and deaths." Fire and property maintenance codes are adopted locally, and the authority having jurisdiction (AHJ) decides what applies.

02

What reactive maintenance covers in practice

In an office building, a school, a retail center or a residential tower, reactive work orders fall into familiar groups:

  • Plumbing: leaks, clogged toilets, running flush valves, drains backing up.
  • HVAC: no heat, no cooling, a tripped rooftop unit, hot and cold calls from tenants.
  • Electrical: lights out, tripped breakers, damaged receptacles.
  • Doors and access control: doors that will not latch, broken closers, a failed card reader.
  • Envelope and interiors: roof leaks, stained ceiling tiles, torn carpet, broken fixtures.
  • Elevators: a car out of service, handled by the elevator contractor under its own agreement.
  • Life safety systems: a fire door that does not close, an exit sign dark, a trouble signal on the fire alarm panel.
  • Damage: vandalism, forced entry, storm damage, a burst pipe.

Whether the work is done by in-house building engineers or contracted out to vendors, every job follows the same path: the problem is reported, by a tenant, a custodian, a day porter, a security officer or an engineer on rounds; a work order is opened with the location and ideally a photo; it is prioritized and assigned; someone attends, makes it safe and fixes it or orders the part; the fix is verified and the work order closed; and the person who reported it is told.

Most facilities run the middle of that path through a facilities help desk or a computerized maintenance management system (CMMS). The weak points are the first step and the last two.

03

Priorities and response times

Reactive work needs a priority scale that the help desk, the engineers, the vendors and the property manager all use the same way. The times below show the shape only; yours belong in the contract or service level agreement.

  • Priority 1, emergency: an immediate risk to life safety, security or the building, such as flooding, a gas odor or a failed fire protection system. Respond and make safe within an hour or two, around the clock.
  • Priority 2, urgent: a failure that takes part of the building out of use or will get worse quickly, such as the only accessible restroom out of order. Respond the same day or within 24 hours.
  • Priority 3, routine: a fault that is inconvenient but not a risk, such as a dripping faucet or a stained ceiling tile. Respond within a few business days.
  • Priority 4, scheduled: work that can be grouped with other jobs or done after hours, within weeks.

Three habits make a priority scale work in practice:

  1. Measure response and completion separately. Response is when someone arrives or makes the problem safe; completion is when it is resolved. Otherwise a vendor can meet every response target while work orders wait weeks for parts.
  2. Rank the risk, not the caller. A light out is a nuisance in a conference room and a hazard in an exit stair.
  3. Plan for after hours. Say who answers an emergency call at night and on weekends, who may authorize an after-hours service call and up to what amount, and what the security officer or night custodian on site should do while they wait.

Making it safe is not a repair. A leak isolated at the valve or a wet area barricaded is the right first response, and for a floor or walkway it is what 29 CFR 1910.22(d)(2) asks for when the repair cannot be made at once. The work order stays open until the permanent fix; see also slip and fall prevention.

04

The balance between planned and reactive work, and the cost

Reactive maintenance is not a failure in itself. Some breakdowns cannot be foreseen, and for non-critical items running to failure can be the cheapest choice. The problem is a facility that is purely reactive.

FEMP's guide lists the advantages of reactive maintenance as low cost and less staff, and calls them "a double-edged sword." Its disadvantages are:

  • increased cost due to unplanned downtime of equipment;
  • increased labor cost, especially if overtime is needed;
  • cost involved with repair or replacement of equipment;
  • possible secondary equipment or process damage from equipment failure;
  • inefficient use of staff resources.

The guide adds that running to failure shortens equipment life and needs a larger stock of spare parts. It also says that a study from the winter of 2000 still found reactive work to be the predominant mode in the United States, at more than 55% of the average facility's maintenance resources and activities, against 31% preventive, 12% predictive and 2% other. It estimates that preventive maintenance saves 12% to 18% on average over a purely reactive program, and says the breakdown of continually top-performing facilities is under 10% reactive, 25% to 35% preventive and 45% to 55% predictive. Those figures are FEMP's, drawn from the studies it cites; treat them as a direction, not a benchmark for your own building.

There is no standard ratio of planned to reactive work that suits every facility. A more useful habit is to track reactive work orders by asset and by cause:

  • Repeat work orders on one asset point to equipment that needs replacing, or a preventive maintenance task that is missing or being skipped.
  • A cluster of calls in one area points to a root cause: a failing roof section, an undersized drain line, an air handler serving too much space.
  • Work that was put off and is now failing belongs on the deferred maintenance list, with a cost, so the people who fund the building can see it.

A CMMS makes that analysis easier, because every work order is tied to an asset and a location. The aim is not to eliminate reactive work but to choose it, asset by asset, and put critical assets such as fire protection, life safety and the main mechanical plant on a schedule. That is how a maintenance team reduces costly failures and unplanned downtime without planning work nobody needs.

An HVAC technician testing a rooftop unit with a multimeter while a building engineer looks on from the roof of a retail center.

05

Vendors on site, and closing the loop

Most facilities use outside vendors for some reactive work. Settle the arrangements before the first emergency:

  • Approved vendors: a short list per trade, each with a current certificate of insurance on file and any state or local license the trade requires.
  • How they are paid: time and materials with a not-to-exceed amount, a flat rate, or a service agreement, with trip charges and after-hours rates agreed in advance.
  • Authorization limits: who may issue a work order to a vendor, and up to what amount.
  • Safety on site: a vendor working on equipment must control hazardous energy. Under 29 CFR 1910.147, The control of hazardous energy (lockout/tagout), whenever outside servicing personnel do covered work, "the on-site employer and the outside employer shall inform each other of their respective lockout or tagout procedures."
  • Asbestos before intrusive work: OSHA's construction standard, 29 CFR 1926.1101, Asbestos, treats repair and maintenance likely to disturb asbestos-containing material as Class III work, and says that before work begins, "building and facility owners shall determine the presence, location, and quantity of ACM and/or PACM at the work site" and tell the employers and employees concerned.
  • Fire protection impairments: when a repair takes a sprinkler, standpipe or fire alarm system out of service, follow the impairment procedure in your locally adopted fire code and the AHJ's requirements, which may include notifying the fire department and posting a fire watch.

On a shared site, responsibility does not stop with the vendor who created a hazard. OSHA's Multi-Employer Citation Policy, CPL 2-0.124 defines a controlling employer as one "who has general supervisory authority over the worksite, including the power to correct safety and health violations itself or require others to correct them," and expects it to exercise reasonable care to prevent and detect violations.

Closing a work order has two halves, and the second is the one that gets forgotten:

  1. Verify the fix. Someone on site, not only the technician who did the work, confirms the problem is resolved before the work order is closed. A work order closed by the vendor who did the job is a claim, not a check.
  2. Tell the person who reported it. A short note that the problem is fixed, or when it will be, turns a request into trust. People who never hear back stop reporting.

Response and completion times, repeat calls and reopened work orders are the basis of vendor reviews. Keep work order history tied to the asset: when an insurer or an inspector asks what was done and when, the answer should be a record, not a recollection.

06

Where the record fails, and what SiteClara does about it

Reactive maintenance records are usually good from the moment a work order is opened in the CMMS. They are weak before and after. Before, because the leak a night custodian noticed or the exit sign a security officer saw dark on patrol is passed on verbally, or on a note that never reaches the help desk. After, because nobody verifies the fix or tells the person who reported it.

SiteClara works on the teams on the ground and the shared spaces they look after. A printed QR poster goes at each location, such as a restroom, a mechanical room, a stairwell, a loading dock or a parking level, with an optional NFC tag behind it. Staff scan the code or tap the tag on their own phone, with no app to install, and report a problem with a photo. The report already knows the building and the location, and it goes onto the team's list of jobs until someone closes it, so the leak has a date, a place, a photo and a name from the first day.

The supervisor's queue groups open jobs by building and floor, and a supervisor can escalate a job to the building manager, who can answer it. Routine checks can be scheduled at a location too, and marked done, or what stopped them said, where they happen. Each day the supervisor approves a report that goes to designated management or client contacts at 8 a.m. the next morning, showing what was reported, completed and still open.

07

Questions people ask

What are the four types of maintenance?

The Federal Energy Management Program's O&M Best Practices Guide, Release 3.0, Chapter 5: Types of Maintenance Programs compares four: reactive maintenance (breakdown or run-to-failure), preventive maintenance (time-based), predictive maintenance (condition-based) and reliability centered maintenance, which combines predictive and preventive techniques with root cause failure analysis. Other sources group them differently, but these are the four programs the federal guide sets side by side.

What is the difference between reactive and proactive maintenance?

Reactive maintenance waits for the failure: FEMP's O&M Best Practices Guide, Chapter 5 calls it the "run it till it breaks" mode and describes it as repairing or replacing damaged equipment "when obvious problems occur." Proactive maintenance acts before the failure. The same chapter describes preventive and predictive maintenance as repairing or replacing damaged equipment "before obvious problems occur," and labels reliability centered maintenance "Pro-Active or Prevention Maintenance." It notes that the secondary damage a failure can cause is a cost "we would not have experienced if our maintenance program was more proactive."

What are the disadvantages of reactive maintenance?

FEMP's O&M Best Practices Guide, Chapter 5 lists increased cost due to unplanned downtime, increased labor cost, especially if overtime is needed, the cost of repairing or replacing the equipment, possible secondary equipment or process damage, and inefficient use of staff resources. It adds that equipment run to failure has a shorter life, is likely to need more extensive repairs, will probably fail during off hours or close to the end of the normal workday, and needs a large inventory of repair parts.

08

Where to read more, and a list to take away

Read FEMP's O&M Best Practices Guide, Chapter 5 for the four approaches side by side, OSHA's 29 CFR 1910.22, 29 CFR 1910.37 and 29 CFR 1910.147 for the safety duties, and 29 CFR 1926.1101 before intrusive repairs in older buildings. Check OSHA's State Plans page to see whether your state runs its own plan, and ask the AHJ about the fire and property maintenance codes adopted locally.

To keep reactive maintenance under control, check that:

  • every problem becomes a work order with a location and a priority, whoever reports it and however;
  • priorities follow risk and how critical the asset is, with example problems for each level;
  • response and completion times are measured separately;
  • after-hours emergencies have a named route and a person who can authorize an after-hours service call;
  • a hazard that cannot be fixed at once is guarded, and its work order stays open until the permanent repair;
  • vendors are approved, insured and licensed, exchange lockout/tagout procedures, and are told about asbestos before intrusive work;
  • any impairment of a fire protection system follows the local fire code, with a fire watch where required;
  • repeat work orders are reviewed and fed into the preventive maintenance schedule or the deferred maintenance list;
  • the person who reported the problem is told when it is fixed.

Sources

Every document this guide quotes or links to, in the order it first cites them.

  1. O&M Best Practices Guide, Release 3.0, Chapter 5: Types of Maintenance Programs www1.eere.energy.gov
  2. Occupational Safety and Health Act of 1970, Section 5, Duties osha.gov
  3. 29 CFR 1910.22, General requirements (Subpart D, Walking-working surfaces) osha.gov
  4. 29 CFR 1910.37, Maintenance, safeguards, and operational features for exit routes osha.gov
  5. OSHA's State Plans page osha.gov
  6. 29 CFR 1910.147, The control of hazardous energy (lockout/tagout) osha.gov
  7. 29 CFR 1926.1101, Asbestos osha.gov
  8. Multi-Employer Citation Policy, CPL 2-0.124 osha.gov