Content

From Maintenance Schedules to AI: How We Learned to Predict Machine Failure

Replace a component too early and you waste useful life. Replace it too late and you may be dealing with an unplanned shutdown.

Maintenance teams have always had to navigate somewhere between the two.

For a long time, the safest answer was often a schedule. Inspect the equipment every six months. Replace the bearing after a defined number of operating hours. Overhaul the pump during the next planned shutdown.

There is nothing inherently wrong with that approach. For many assets, preventive maintenance remains the right strategy. A component that is inexpensive, non-critical and easy to replace might even be allowed to run until failure.

But a maintenance interval has an inherent limitation: it tells you when you planned to intervene. It does not necessarily tell you what condition that particular component is actually in.

A bearing scheduled for replacement next month may still have considerable useful life. Another bearing of exactly the same type may already be deteriorating.

Reliability engineers have spent decades trying to understand that difference. The objective has not fundamentally changed. The evidence available to make the decision has.

Explore our "From Prediction to Prevention" series:

Part 1: From Maintenance Schedules to AI: How We Learned to Predict Machine Failure
How reliability engineering evolved from historical failure models and condition monitoring towards AI-assisted predictive maintenance.

Part 2: Predictive Maintenance Gives You 30 Days. Will That Be Enough?
What happens after a potential failure is detected, from identifying the correct spare part to checking availability, sourcing it and preparing the intervention.

Part 3: How Does a Machine Know It's About to Fail?
How vibration, temperature, acoustic signals, oil analysis and other condition-monitoring technologies reveal signs of equipment degradation.

Part 4: Robot Dogs, Drones and Smart Sensors: Meet the New Maintenance Team
How autonomous inspection, computer vision, robotics and other emerging technologies are changing how manufacturers monitor equipment on the shop floor.

Part 5: From Predictive Maintenance to Predictive Spare Parts Planning
How failure probabilities, equipment condition, lead times and operational data could enable more dynamic spare parts inventory decisions.

Part 6: The Autonomous Factory: Can a Machine Order Its Own Spare Part?
What happens when failure prediction, maintenance, inventory and procurement become increasingly connected, and how close manufacturers are to a truly autonomous maintenance workflow.

Engineers Were Predicting Failure Long Before AI

Predictive maintenance did not begin with artificial intelligence.

Long before machines were connected to IoT platforms, reliability engineers were using historical failure data and statistical models to estimate how equipment was likely to behave.

Mean Time Between Failures, or MTBF, is one familiar example. If a group of motors has an MTBF of 10,000 hours, that does not mean each motor will fail after 10,000 hours. MTBF describes average reliability across a population; it does not predict when a specific machine will fail.

Reliability engineers therefore use other statistical approaches where appropriate. One of the best known is the Weibull distribution. The NIST Engineering Statistics Handbook describes Weibull as a particularly flexible life distribution model because it can represent different failure behaviours, including situations where failure risk decreases, remains relatively constant or increases over time.

The point is not the particular calculation. It is that engineers were already using data to move maintenance decisions beyond intuition.

And a reliability engineer can build a sophisticated model in Excel. The limitation is not necessarily the mathematics in the spreadsheet. It is how much changing information humans can realistically feed into that model, how frequently they can do it and across how many assets.

This is something reliability engineer Sanjib Das also highlighted in our interview series on spare parts planning. Mathematical models can replace gut feeling with explainable decisions, but their inputs do not stand still. Failure behaviour changes. Operations change. Lead times change.

A calculation may be correct when it is made and progressively less representative of reality as the world around it changes.

But historical failure data can only tell us so much about the condition of an individual machine. A model might indicate when a component is statistically more likely to fail. It cannot, on its own, tell us whether that particular component is already showing signs of degradation.

This is where condition monitoring changes the picture.

From Failure History to Machine Condition

Take our bearing again.

Its age and historical failure data may suggest that it still has useful life remaining. But vibration measurements begin to change. Its temperature increases. Its acoustic signature no longer resembles its normal operating pattern.

These changes may be early indications that the bearing is beginning to deteriorate, even though it can still perform its intended function. Reliability engineers describe the period between the point at which such a potential failure can first be detected and the point of functional failure using the P-F Curve.

"P" represents the point at which a potential failure becomes detectable; "F" represents the point at which the equipment can no longer perform its required function. The time between them is known as the P-F interval. In practical terms, this is the time available to identify the developing problem, assess the risk and intervene before functional failure occurs.

Better condition monitoring can help detect signs of deterioration earlier within this process, potentially giving maintenance teams more time to decide how and when to intervene. Depending on the equipment and failure mode, those early signs might appear as changes in vibration, temperature, acoustic signals, lubricant condition or electrical behaviour.

We now have two different types of evidence: what tends to happen to bearings like this one, and what appears to be happening to this particular bearing.

This also helps distinguish concepts that are often used interchangeably. Condition monitoring observes the health of equipment. Condition-based maintenance uses that information to determine whether intervention is required. Predictive maintenance goes further by using available information to estimate how the condition is likely to develop and when intervention may become necessary.

The boundaries are not always this neat in industrial practice, but one distinction is particularly important.

A sensor doesn't predict anything.

From Sensor Data to Failure Prediction

A vibration sensor measures vibration. That information only becomes predictive when we interpret what it means.

At its simplest, that might mean defining a threshold and triggering an alert when vibration exceeds an acceptable level. Add historical measurements and we can identify a trend before that threshold is reached. More advanced analytical models can detect behaviour that differs from normal operating patterns and use the development of that degradation to estimate what might happen next.

This is where Remaining Useful Life, or RUL, comes into play. NASA describes prognostics as predicting when a component will no longer perform its intended function, with the time remaining until that point representing its Remaining Useful Life.

But predictive maintenance does not necessarily mean an algorithm announcing:

"This bearing will fail on Tuesday at 14:37."

Predictions contain uncertainty. Operating conditions and loads change, measurements contain noise, and new evidence continues to arrive. As that evidence changes, the prediction itself can change.

Predictive maintenance is therefore better understood as a continuously updated assessment of equipment condition and risk than as a crystal ball predicting an exact failure date.

And once information starts arriving continuously across hundreds or thousands of assets, another challenge emerges: analysing it.

So Where Does AI Come In?

If reliability engineers could already model failure and sensors could already monitor machine condition, what does AI actually change?

One important answer is scale.

A modern production environment can generate enormous volumes of condition data. A single asset might produce information about vibration, temperature, pressure, electrical behaviour and operating load. Machine-learning models can help identify patterns and relationships across those datasets that would be impractical to evaluate manually and can reassess them as new information becomes available.

A 2025 systematic review by Abdulrahman Alamin and his co-authors, covering 60 peer-reviewed studies published between 2020 and 2024, shows how widely machine learning is already being explored for predictive maintenance, from fault detection and diagnosis to Remaining Useful Life estimation. At the same time, the authors highlight persistent challenges around data quality and availability, model interpretability, scalability and deployment in real industrial environments.

AI is not inherently right because it processes more data. Industrial plants may have huge quantities of data showing normal operation but relatively few well-documented examples of actual failures. Models still need the right data, operational context, validation and engineering expertise.

AI therefore does not make reliability engineering obsolete. It gives reliability teams new ways to apply and scale their expertise.

This is perhaps the bigger shift behind predictive maintenance. A traditional reliability model might use historical failure information to calculate risk and inform a maintenance decision at a particular point in time. A connected predictive-maintenance system can incorporate new evidence as the machine continues to operate. Another 500 operating hours may pass, the load may change, vibration may increase or a new anomaly may appear. Each new piece of information can change the assessment.

The resulting recommendation does not always have to be to replace the component. Depending on how its condition develops, the appropriate response might be to continue monitoring it, increase inspection frequency, prepare a replacement or schedule an intervention during the next planned shutdown.

The objective is not to automate every engineering decision. It is to give the people making those decisions more timely and continuously updated evidence.

And the same shift is beginning to influence another critical maintenance decision: whether the right spare part will be available when it is needed.

What Happens to the Spare Part?

Suppose the analysis now indicates that our bearing is deteriorating and is likely to require intervention within the next 30 days.

The maintenance team has gained something extremely valuable: time.

But a prediction alone does not prevent downtime.

Someone still needs to know which bearing is required, whether the material master contains the correct information, whether the part is already in the warehouse or available at another site, and whether Procurement can source it before the intervention.

This is where predictive maintenance starts to intersect with spare parts management.

Spare parts planning faces a surprisingly similar decision problem. Reliability asks how likely we are to need a component. Inventory planning then needs to determine whether that probability, combined with the consequences of not having the part available, justifies stocking it.

At SPARROW, we see the same transition from static towards dynamic decision-making in spare parts planning. SPARROW.Plan uses reliability-based methods together with operational spare parts data to calculate stocking recommendations. As the underlying information changes, recommendations can be recalculated rather than leaving teams dependent on decisions made years earlier and never revisited.

As SPARROW CEO Meir Veisberg puts it:

“We are providing our customers with a competitive edge, enabling them to make informed decisions and optimize their operations.”

The technology matters because the decision matters.

Prediction Is Only the Beginning

Predictive maintenance gives manufacturers the opportunity to know more about the condition of their equipment, and potentially to know it earlier.

That advance warning is where much of its value lies.

If you know that a component may require intervention in 30 days rather than discovering the problem when production stops, you have potentially gained 30 days to prepare.

But can you identify the right spare? Is it already in stock? Could another plant have it? Can Procurement source it before the planned intervention?

Predicting the failure is only the first decision. Preventing the downtime requires everything that happens next to work too.

Next in the series: Your AI Predicts a Failure in 30 Days. Now What?

Ready to streamline your spare parts strategy?

Green stylized letter S with interlocking gear-like curved lines
See ROI in 3 months
Ready to use – no integration needed
From one module to full Hub — your way