Skip to main content
DoneThat

AI Adoption GuidePropertyMaintain

Predictive Equipment Failure Model

ML analyzes IoT sensor data from HVAC, lifts, and plant to predict failure before it occurs, triggering work orders and parts procurement ahead of breakdown. (e.g., Facilio, IBM Maximo, Siemens MindSphere)

Property processAcquireLeaseOccupyMaintainBillRenewVacateDispose

By Don, DoneThat’s AI coach · updated

What a predictive equipment failure model does

A predictive equipment failure model reads streams from IoT sensors on HVAC plant, lifts, and other building systems, then scores how likely a component is to fail within a planning window. The score is decision support for a facilities engineer, not an automatic ticket or purchase. You still raise the work order, choose the intervention, and decide whether parts should be reserved or ordered.

Typical inputs include vibration, temperature, current draw, differential pressure, runtime hours, door-cycle counts on lifts, and fault codes from controllers. The model compares those signals against each asset's history and against peer assets of the same type. When patterns that preceded past failures reappear, the risk score rises and the CMMS or IWMS can surface a recommended inspection or planned job for human review.

Platforms in this space (for example Facilio, IBM Maximo, and Siemens MindSphere) often sit beside the building management system and the work-order system. The value is earlier notice than calendar-based PM alone, without pretending the model can replace engineering judgment on site.

Sensor history requirements and empty results

The model needs enough labeled or observed history to separate normal drift from failure precursors. If an asset has no usable sensor history, incomplete tagging, or only a few weeks of data after a major retrofit, the honest output is empty: no score, no ranked failure mode, no suggested work. Showing a low-confidence guess as if it were a plan creates false confidence and wasted call-outs.

Treat empty output as an operational signal. It means fix instrumentation, map sensors to asset IDs, backfill historian data where it exists, or stay on time-based and inspection-based maintenance until the series is long enough. Do not invent a default "healthy" score for assets that never produced training-quality data.

When history is partial, scope the model to the assets that qualify and keep the rest on existing PM routes. A partial deployment that is accurate beats a site-wide score that is mostly noise.

How engineers use risk scores in the maintain stage

Day to day, you review a queue ordered by risk and criticality, not by which sensor screamed loudest. A high score on a redundant fan with a spare online is different from a medium score on the only chiller serving a critical floor. Pair the model output with business impact, SLA commitments, and known lead times for parts.

The human-in-the-loop rule is fixed: the model scores failure risk; engineers raise the work order. Confirmation steps usually include a short trend review, a visual or listening check where safe, and a decision on whether to convert the alert into a planned job, escalate to a specialist contractor, or dismiss with a reason code. Reason codes feed retraining and reduce repeat false positives.

Do not auto-order parts from a risk score. Procurement stays a separate decision after the engineer confirms the likely failure mode and the required BOM. Auto-ordering from a probabilistic score creates stockouts of the wrong SKU, cancelled POs, and eroded trust in the program. Reserve or kit parts only after a person accepts the diagnosis and the job plan.

Related practice on the same maintain path includes Building Defect Detection via Computer Vision for fabric and envelope issues that sensors may miss, Contractor Invoice Validation when reactive and planned jobs hit the same vendors, and Energy Optimization Agent when HVAC setpoints and plant sequencing interact with wear.

Failure modes the model is good at, and where it is not

Vibration and thermal drift on rotating plant, clogging and filter load on AHUs, and abnormal current signatures on motors are common wins because precursors show up in continuous signals. Lift systems often benefit from cycle counts, door timing, and drive fault history when the OEM or gateway exposes them cleanly.

The model is weaker when failures are sudden mechanical breaks with no precursor, when sensors are poorly calibrated, or when the "failure" label in CMMS history was used loosely (for example, every tenant complaint tagged as equipment failure). Garbage labels teach the model the wrong story. Clean close-out codes and failure-mode taxonomies matter as much as the algorithm.

Seasonal assets need care. A cooling tower idle for months can look "anomalous" on restart unless the model and the engineer account for start-of-season behavior. Likewise, after a major overhaul, reset or re-baseline so the new healthy state is not scored as degradation.

Rolling the model into plant maintenance planning

Start with a short asset list: high-impact HVAC plant, critical lifts, and any system whose downtime forces expensive temporary plant or SLA credits. Prove that scores arrive early enough to schedule during preferred windows, that false positives stay manageable, and that empty outputs appear when data is missing rather than silent wrong greens.

Wire alerts into the CMMS as proposed tasks or watchlist items, not as released work. Keep parts planning in the same review: suggested BOMs can appear as hints, but purchasing remains manual. Measure outcomes that facilities teams already care about: emergency call-out rate on covered assets, planned versus reactive mix, and mean time between confirmed failures, using your own CMMS extracts rather than vendor marketing numbers.

Governance should name who owns dismissals, how often models are retrained after major plant changes, and what happens when a sensor goes dark (treat dark sensors as empty or degraded coverage, not as "no risk"). When those habits are in place, the predictive score becomes a planning input alongside inspections, OEM guidance, and contractor capacity, not a black-box replacement for any of them.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other. This one is rated high effort to implement, so the baseline matters more than usual.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first