The Short Answer
The same equipment keeps breaking because the repair fixed the symptom, not the root cause, and nothing in the process caught that before the next failure.
In multi-location facilities, the cause is usually one of a few things: a vague work order, a vendor who swapped a part without finding out why it failed, a preventive maintenance task that misses the actual failure mode, or an asset past its useful life. You stop repeat failures by tracking them as a metric, investigating the worst offenders, fixing the cause, and confirming the fix held.
This guide covers each cause with examples from rooftop HVAC units, walk-in coolers, ice machines, roofs, and plumbing. It then gives a 7-step framework, a repair vs. replace guide, and the metrics that show whether it is working.
What Is Repeat Equipment Failure?
A repeat equipment failure is when the same asset fails in the same way, or with the same symptom, within a set window after it was "fixed." Many maintenance teams use a 90-day window (Fabrico).
In work order terms, it is a repeat work order: a new corrective ticket for an asset that recently had a corrective ticket closed for the same problem.
Repeat failure vs. normal wear
Not every second repair means something is wrong. A belt that wears out after 18 months and gets replaced is normal wear. A belt replaced three times in one summer is a repeat failure, and it usually points to a misaligned pulley, the wrong belt size, or a motor problem.
The test: did the failure come back faster than the component's normal service life? If yes, the root cause is still there.
What is a "bad actor" asset?
A bad actor is an asset that fails more often, or costs more to maintain, than comparable assets. One oil and gas reliability program enrolled equipment once its maintenance cost passed 125% of the expected average, and removed it only after 12 months with no failures (Gas Processing News, 2025).
The idea transfers directly to facilities. If 40 stores run the same rooftop unit model and three stores generate half the HVAC tickets, those three units are your bad actors.
Why The Same Equipment Keeps Breaking: 10 Root Causes
Most repeat failures trace to one of these ten causes, and many involve two or three at once.
1. The repair fixed the symptom, not the cause
This is the most common cause, and reactive maintenance builds it in. A store calls it "warm walk-in cooler warm." A technician finds a tripped compressor, resets it, and the unit cools. Ticket closed, but nobody asked why the compressor tripped.
If the cause was a condenser coil packed with grease and dust, it will trip again. Federal Energy Management Program research found that a dirty coil raising condensing temperature from 95°F to 105°F cuts cooling capacity by 7% and raises power use by 10% (JADE Learning). Food Service Technology Center data estimates dirty coils waste $220 to $625 per refrigeration unit per year in electricity (DOE docket).
Low not-to-exceed (NTE) limits can make this worse. When a vendor is authorized only for a quick fix, a quick fix is what you get. See why low NTEs backfire.
2. Work orders don't capture what actually failed
"HVAC not working. Tech came out. Fixed." That work order cannot prevent the next failure, because it records no symptom, failed component, cause, or replaced part.
Without structured closeout fields, each ticket looks like a one-off and repeat failures stay invisible. A clear work order triage model and consistent closeout fields are the foundation for everything else in this guide.
3. There is no usable asset history
In many multi-location organizations, asset history lives in vendor invoices, email threads, a regional spreadsheet, and one long-tenured manager's memory. The technician on site cannot see that this unit had the same capacitor replaced in March and June.
History matters most for older equipment. A Plant Engineering maintenance study found aging equipment is the single biggest factor in unscheduled downtime, with equipment failure a close second (Plant Engineering, 2019). Without history, you cannot tell "wearing out" from "keeps getting misdiagnosed."
4. Vendor quality and callbacks go unmeasured
Most multi-location teams rely on outside vendors for HVAC/R, plumbing, electrical, roofing, and kitchen equipment. That makes vendor quality a root cause in its own right.
Aberdeen Group research puts the average first-time fix rate in field service at 75%, with best-in-class organizations at 89%. A job not fixed on the first visit takes about 1.6 extra visits (OptimoRoute). Failed first fixes trace mainly to missing or incorrect parts (51%), skills gaps (25%), and not enough time on site (13%) (AEX).
National Apartment Association data shows HVAC is both the most common work order category and the most common callback (NAA). If you don't track callbacks by vendor, a vendor who returns three times can look cheaper per invoice than one who fixes it once.
A repeatable vendor onboarding process that sets callback expectations up front helps.
5. Preventive maintenance is missed, or it's the wrong PM
The US Department of Energy estimates preventive maintenance saves 12% to 18% over a purely reactive approach, and predictive maintenance saves another 8% to 12% over preventive alone (DOE O&M Best Practices Guide; PNNL). Those savings only appear when PM is completed and targets how the asset actually fails. For a facilities view of the tradeoff, see reactive vs. preventive maintenance in facilities.
-
Missed PM: filter changes, coil cleanings, and belt checks get skipped in busy seasons, which is exactly when HVAC and refrigeration are under the most stress.
-
Wrong PM: the task exists but misses the failure mode. Quarterly filter changes do nothing for a condensate drain line that clogs every summer.
Over-maintenance is a real risk too, because every time someone opens equipment, a new problem can be introduced. Nowlan and Heap's 1978 study of airline components found only about 11% of failures followed a predictable wear-out pattern (Reliamag).
Calendar-based PM alone cannot stop every failure, so match each task to how the asset actually fails.
6. How site staff use the equipment
In facilities, the "operators" are store staff, kitchen crews, teachers, volunteers, and custodians. Common examples:
-
Walk-in cooler doors propped open during deliveries
-
Thermostats overridden or set to extremes
-
Ice machine filters never changed because nobody owns the task
-
Fryers and dish machines run without basic daily cleaning
A literature review found human error to be a major factor in 70% to 80% of equipment failures, incidents, and accidents, with maintenance error alone at 15% to 20% (MATEC Web of Conferences, 2020). Those studies come mostly from aviation and industry, so treat the figures as directional. If one site keeps breaking the same equipment, look at how it is used, not only how it is repaired.
7. Wrong or low-quality parts
A non-OEM capacitor, a belt one size off, a gasket that "almost fits": each can turn one failure into a series. With 51% of failed first fixes tied to missing or incorrect parts, technicians sometimes install whatever is on the truck to get the unit running.
Ask for part numbers on every closeout. Check warranty status before approving replacements, since non-approved parts can void coverage.
8. The operating environment changed
Equipment rarely runs under its design conditions for its whole life (6D Testing & Analysis). In facilities, that looks like:
-
A restaurant adds menu items and runs the kitchen exhaust harder
-
A school extends hours for after-school programs
-
A remodeled retail space keeps its original, now undersized, rooftop unit
-
A coastal site corrodes coils and fasteners faster than an inland one
When the operating environment changes, the failure rate changes with it, and replacing the same part will not keep up.
9. The asset is at the end of its life
Some equipment keeps breaking because it is worn out. ASHRAE data puts the median service life of a single-zone rooftop air conditioner at about 15 years (ASHRAE chart). In ASHRAE's public database, 215 of 220 packaged rooftop units were still in service at a median age of 16 years (ASHRAE).
A median is not a deadline. But once an asset is past its expected life and repair frequency is climbing, the question changes from "how do we fix it?" to "should we keep fixing it?"
10. Nobody verified the fix
A unit restarts, the ticket closes, and nobody checks whether it still runs correctly a week later. This is the gap between a closed ticket and a solved problem.
Verification can be simple: a walk-in temperature reading 48 hours after repair, a site manager sign-off after a leak repair, or a photo of a cleaned coil.
A 7-Step Framework To Stop Repeat Failures
The same seven steps work for one building or 500 sites; only the volume of data changes.
-
Capture every failure properly. Require structured closeout on every corrective work order: asset ID, symptom, failed component, cause found, action taken, parts used (with part numbers), and vendor. "Fixed" is not a closeout.
-
Define and measure repeat failures. Pick a window (90 days is a common start) and flag any corrective work order on the same asset with the same symptom inside it. A useful first rule: investigate any asset with more than three corrective work orders in 90 to 180 days (Limble).
-
Compare identical assets across sites. If 30 locations run the same equipment model, you have a natural experiment. When one site has three times the tickets of the others, the cause is almost always local: installation, usage, environment, or the vendor serving that site. Rank assets by repeat work orders and repair cost to build your bad actor list. This depends on managing maintenance across multiple locations without losing visibility.
-
Investigate the top offenders, not everything. Run root cause analysis on your top 5 to 10 bad actors. The 5 Whys method is enough for most facility assets (worked example below).
-
Fix the cause, not just the part. Match the action to the cause: change a PM task, shield or relocate equipment, retrain site staff, change the part specification, or change vendors. If the cause is end of life, start the replacement process.
-
Adjust preventive maintenance to what actually failed. If drain lines clog every summer, add a spring drain treatment. If coils foul quarterly at kitchen-adjacent units, clean those units more often than others. The same asset type should be able to run on different schedules at different sites.
-
Verify, then watch the metric. Confirm the fix held with a follow-up reading, photo, or site sign-off. Then track whether the asset drops off the repeat list; a common bar is no repeat failure for 12 months after corrective action.
Worked example: 5 Whys on a walk-in cooler
|
Question |
Answer |
|---|---|
|
Why is the walk-in cooler warm? |
The compressor tripped on high pressure. |
|
Why did it trip? |
The condenser could not reject heat. |
|
Why not? |
The coil was coated in grease and dust. |
|
Why was it coated? |
The coil is not on the PM schedule, and the unit sits next to the kitchen exhaust. |
|
Why is it not on the schedule? |
The PM template was copied from a site without a nearby exhaust and never adjusted. |
The root cause is a PM template, not a compressor. Replacing the compressor would have cost far more and fixed nothing.
Repair or Replace? A Decision Guide For Facility Assets
No single formula decides repair vs. replace, but weighing these six factors together makes the decision defensible to finance.
|
Factor |
Lean toward repair |
Lean toward replace |
|---|---|---|
|
Age vs. expected life |
Under about 75% of expected life (under 11 years for a 15-year rooftop unit) |
At or past expected service life |
|
Repair cost vs. replacement |
A single repair well under 50% of replacement cost |
A repair near 50% of replacement, or 12-month repairs above about 30% of replacement cost |
|
Repeat failures |
First or second occurrence with a fixable root cause |
Repeats continue after the root cause was corrected |
|
Parts and refrigerant |
OEM parts readily available |
Parts obsolete, or refrigerant phase-down raising service costs |
|
Energy use |
Near rated efficiency |
Well below current equipment |
|
Business impact |
Low-criticality asset with an easy workaround |
Failure closes a location, spoils inventory, or creates safety or compliance risk |
The 50% and 30% thresholds are contractor rules of thumb, not formal standards. Adjust them for criticality: a failing walk-in that spoils $4,000 of product a month deserves a lower threshold than a break room fan.
For a portfolio, build these decisions into asset lifecycle management and periodic facility condition assessments, so replacements show up in the capital plan before they become emergencies.
The Metrics That Tell You It's Working
Repeat work order rate is the one metric that directly shows whether repairs are solving problems; the other five explain why it moves.
|
Metric |
Formula |
What it tells you |
Starting target |
|---|---|---|---|
|
Repeat work order rate |
Corrective WOs on the same asset with the same symptom within 90 days of a prior closure ÷ total corrective WOs × 100 |
Whether repairs fix problems |
Steady decline |
|
Mean time between failures (MTBF) |
Total operating time ÷ number of failures |
Run time between breakdowns |
Rising |
|
Mean time to repair (MTTR) |
Total repair time ÷ number of repairs |
Speed of resolution |
Falling, without more repeats |
|
PM compliance |
PMs completed on time ÷ PMs scheduled × 100 |
Whether the plan is executed |
Above 90% (DOE) |
|
First-time fix rate, by vendor |
Jobs resolved on the first visit ÷ total jobs × 100 |
Vendor quality |
75% average, 89% best in class (Aberdeen) |
|
Repair cost vs. replacement value |
12-month repair spend ÷ replacement cost |
When repair stops making sense |
Review above about 30% |
One caution on MTTR: pushing repair speed alone rewards quick resets over real fixes. If MTTR improves while it repeatedly rises, the team is getting faster at fixing the same thing twice.
Track maintenance cost by asset and site so leadership can see what each bad actor actually costs. For the wider metric set, see facility maintenance KPIs.
Repeat Failure Checklist
Run these ten checks the next time an asset fails twice.
-
☐ Same asset, same symptom, within 90 days?
-
☐ Does the last work order record the failed component and the cause?
-
☐ Did the vendor provide part numbers and photos?
-
☐ Is the asset under warranty, and was the warranty provider used?
-
☐ Do identical assets at other sites show the same problem?
-
☐ Is there a PM task that targets this failure mode, and was it completed?
-
☐ Has anything changed at the site (hours, usage, remodel, weather)?
-
☐ Is the asset past about 75% of its expected life?
-
☐ Have 12-month repair costs passed about 30% of replacement cost?
-
☐ Did anyone verify that the last fix held?
Frequently Asked Questions
Why does my equipment keep breaking after it's repaired?
Equipment usually keeps breaking after a repair because the repair fixed the symptom rather than the root cause. Common hidden causes include dirty coils, the wrong replacement part, a preventive maintenance task that misses the actual failure mode, changed operating conditions, or an asset at the end of its useful life.
What is a repeat work order?
A repeat work order is a new corrective ticket for the same asset and the same problem within a set window, often 90 days, after a previous repair was closed. Tracking the repeat work order rate shows whether repairs are solving problems or only resetting them.
What is a bad actor asset in facilities management?
A bad actor asset is equipment that fails more often, or costs more to maintain, than comparable assets. In multi-location facilities, comparing the same equipment model across sites is the fastest way to find bad actors, because outliers stand out clearly against the rest of the fleet.
What are the most common causes of equipment failure in commercial facilities?
The most common causes are missed or poorly targeted preventive maintenance, symptom-only repairs, misuse by site staff, wrong or low-quality parts, environmental stress such as heat, grease, and corrosion, and aging equipment. Weak work order data and unmeasured vendor quality let these causes repeat unnoticed.
How do you do root cause analysis on facility equipment?
Start with a clearly defined failure, then ask "why" repeatedly (the 5 Whys method) until you reach a cause you can act on, such as a missing PM task or a mismatched part. Use work order history and cross-site comparisons as evidence, then verify that the corrective action held.
When should you replace equipment instead of repairing it?
Consider replacement when an asset is near or past its expected service life, when a single repair approaches 50% of replacement cost, when 12-month repair costs pass about 30% of replacement cost, or when repeat failures continue after the root cause was addressed. These thresholds are rules of thumb, not formal standards.
How much does preventive maintenance reduce costs compared to reactive maintenance?
The US Department of Energy estimates preventive maintenance saves 12% to 18% compared with a purely reactive program, and predictive maintenance saves another 8% to 12% over preventive alone. Actual savings depend on how reactive the starting point is.
What is a good first-time fix rate for maintenance vendors?
Aberdeen Group research puts the average first-time fix rate in field service at about 75%, with best-in-class organizations at about 89%. A job not fixed on the first visit takes about 1.6 additional visits on average, so tracking first-time fix rate by vendor exposes hidden repeat costs.
Where Software Fits (And Where It Doesn't)
Software does not fix equipment; technicians do. What the right system does is make repeat failures visible, so the framework above stops depending on someone stitching together invoices and emails by hand.
The capabilities that matter are structured closeout, asset history every vendor can see, repeat work order and callback tracking by site and vendor, PM schedules that vary by location, and repair vs. replace data in one place.
LeanSite was built for this problem in multi-location and mid-sized organizations. It combines work orders, asset management, vendor management, preventive maintenance, budget tracking, and analytics, and its AI assistant, Vera, helps teams spot patterns inside the workflow.
Whatever tool you use, judge it by one question: can it show you, in under a minute, which assets broke twice this quarter, and why?
Written by Pelumi Akinwande, Operations Content Lead at LeanSite, who works directly with multi-site facilities and property operations teams evaluating work order software. Connect on LinkedIn.
Tags: