AI Adoption GuideManufacturingMake
Predictive Maintenance on Production Assets
Vibration and thermal sensor ML predicts bearing, motor, and hydraulic failures 48 to 72 hours ahead and automatically generates maintenance work orders.
Manufacturing processPlanSourceMakeInspectPackShipServiceReturn
By Don, DoneThat’s AI coach · updated
What this use case covers
Predictive maintenance on production assets uses vibration and thermal sensor streams, plus machine learning, to surface early signs of bearing, motor, and hydraulic degradation. The goal is not a fully autonomous maintenance program. It is earlier, more specific risk signals that a planner or maintenance supervisor can turn into scheduled work before a line stops.
On a make-stage line, that usually means continuous or near-continuous monitoring on critical rotating and fluid-power equipment: motors, gearboxes, pumps, fans, spindles, and hydraulic power units. Models score asset health, highlight likely failure modes, and can draft or trigger maintenance work-order candidates in the CMMS or EAM. A human still owns release to the floor.
This page is for maintenance managers and reliability leads who already have, or are adding, condition sensors and need a clear operating model: what the system should predict, what it must never invent, and how work gets approved.
Signals, models, and failure modes
Useful predictions start with signals that map to known wear mechanisms. Vibration spectra and envelope features help detect imbalance, misalignment, looseness, and bearing race or rolling-element defects. Thermal trends catch overheating from friction, electrical issues, or restricted cooling. Hydraulic assets often need pressure, flow, and temperature in addition to vibration on pumps and motors.
Models typically learn a healthy operating baseline for each asset or asset class, then score deviation severity and progression. Outputs that planners can act on include:
- Asset and component at risk (for example, drive-end bearing on Motor M-214)
- Likely failure mode family (bearing, motor winding or rotor, hydraulic pump/valve)
- Relative urgency or health trend (worsening vs stable)
- Recommended inspection or corrective action class (inspect, lubricate, align, replace, oil analysis)
Do not treat a vendor’s marketing window as a proven plant result. How far ahead a useful alert arrives depends on sensor placement, sample rate, load profile, and how aggressively the plant acts on warnings. Run your own lead-time measurement on closed work orders and confirmed failures before you set planning assumptions.
Vendor platforms in this space include Augury, Siemens Senseye, and PTC ThingWorx. Evaluate them on sensor and historian integration, failure-mode explainability, CMMS write-back, and how empty or low-confidence states are shown when data quality drops.
How predictions become work orders
A durable workflow keeps prediction and authorization separate.
- Sensors stream to the edge or historian; the model updates health scores and alerts.
- High-risk or rapidly worsening assets create a candidate work order or notification in the CMMS.
- Planner or maintenance supervisor reviews evidence: trends, spectra snapshots, recent loads, and related OEE or downtime history.
- Approver confirms priority, parts, craft, and window (planned downtime, changeover, or next opportunity).
- Released work is executed and closed with failure code and findings so the loop can be audited.
The planner or maintenance lead still approves the work order. Auto-creating a draft is fine; auto-releasing wrench time without review is not. False positives burn trust and spare parts. False negatives that were never reviewed leave the same liability as ignored paper checklists.
When sensors are offline, stale, or fail quality checks, the system should return an empty prediction (or an explicit “insufficient data” state), not the last good score dressed up as current health. Offline gaps must be visible on the same board used for live alerts so coverage holes are treated as reliability risk, not silence.
Operating model on the make floor
Start with a criticality-ranked asset list: bottleneck machines, single-point failures, long lead-time spares, and safety-related rotating equipment. Instrument those first. Pair each asset with a clear owner, spare strategy, and “what good looks like” for vibration and temperature under normal loads.
Define alert classes that match how your crew works:
- Watch: trend drifting; include in weekly reliability review
- Inspect: schedule a short diagnostic within a defined window
- Plan: prepare parts and craft for the next planned opportunity
- Escalate: risk of imminent functional failure; protect the schedule and notify operations
Tie predictive alerts to production reality. A rising bearing risk on a constrained cell should reach both maintenance and the scheduler so the outage can land in a known window instead of a scramble. Related make-stage practices include OEE Root Cause Classification for linking downtime codes to asset health, and Real-Time Schedule Reoptimization when a predicted outage forces sequence changes.
For fleets that sit outside the main line or across sites, compare patterns with Remote Asset Health and Failure Prediction. The sensing and model ideas overlap; the difference is whether the primary consumer is the plant maintenance planner on this line or a distributed service organization.
Governance, data quality, and cost outcome
Cost impact comes from fewer unplanned stoppages, less secondary damage, and better use of planned windows, not from eliminating all reactive work. Track a small set of measures that maintenance and finance both accept:
- Unplanned downtime hours on instrumented assets
- Emergency vs planned maintenance mix
- Work orders opened from predictive alerts vs other sources, and their hit rate (confirmed issue vs no-find)
- Mean time between failures or functional failures for the pilot asset set
- Sensor uptime and percent of assets in “empty prediction / no current score” state
Require named ownership for model thresholds, CMMS mapping, and sensor health. Document when a model version changed and whether alert volume shifted. Keep human override and deferral reasons in the work-order record so you can see whether the process or the model is drifting.
Procurement and IT should lock down who can write to the CMMS, what fields auto-populate, and how vendor cloud processing handles plant data. Reliability ownership stays with the plant even when a vendor hosts the analytics.
Practical checklist before you scale
- Critical assets identified; sensors mounted and validated against known good and known bad conditions
- Empty-prediction behavior confirmed for offline and stale feeds
- Draft work-order path tested; approval remains with planner or maintenance supervisor
- Parts and craft capacity reviewed so alerts do not create a queue you cannot staff
- Baseline downtime and emergency work measured before go-live on the pilot set
- Weekly review cadence for watches, no-finds, and missed detections
- Clear handoff to operations and scheduling when a planned intervention will cut into run time
Pilot on one line or one asset family, prove alert quality and work-order discipline, then expand. Predictive maintenance pays when it changes when and how you intervene, under human control, with honest silence when the sensors cannot speak.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other. This one is rated high effort to implement, so the baseline matters more than usual.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first