Anomaly detection and remaining useful life answer different questions. The first says something has changed. The second says how much life is left and how sure it is. Most demonstrations show the first while the sales deck talks about the second, and the difference only becomes visible when maintenance planning tries to use the output. Six criteria, each with the question to ask.
No products are named here, ours included. The disclosure is at the end.
1. Component level or cabinet level
In an inverter the power modules, the DC link capacitors, the cooling fans and the contactors age through different mechanisms and at different speeds. One health index for the cabinet hides the component that will fail first, which is the only one you can act on. Ask for the list of components covered, and check it against your own failure history rather than against the brochure.
Ask: which components are modelled individually, and which are not modelled at all?
2. Anomaly detection or a degradation model
A degradation model turns operating history into consumed life through a known mechanism: thermal cycling into solder fatigue through Coffin-Manson, temperature into electrolyte ageing through Arrhenius, an irregular temperature series into cycles through rainflow counting. Anomaly detection compares the present against the past and flags a deviation. Both are useful. Only the first produces a remaining useful life you can put in a plan.
Ask: is the output a deviation score or a remaining life, and if it is a remaining life, which degradation mechanism produces it?
3. What data it needs, at what resolution, and how much history
Sampling rate decides what is visible. Ten minute averages hide the thermal cycles that cause the fatigue, so a model that claims to count cycles from averaged data is reconstructing them, and you should ask how. Then there is the cold start: many systems say nothing useful until they have seen a year of operation, which matters if the fleet you most want to cover is the one you just commissioned.
Ask: what sampling rate and how many months of history before the first usable output, and what does the system do until then?
4. Whether the horizon matches your planning cycle
A warning that arrives two weeks before a failure is worth very little if the part has a ten week lead time and the site is visited quarterly. Line the forecast horizon up against parts lead time, crew availability and scheduled outages before looking at any accuracy figure. And insist on an interval: a single date with no probability attached cannot be scheduled against.
Ask: at what horizon is the estimate stated, with what uncertainty, and how does that compare with my own lead times?
5. The false alarm rate at the operating point
Accuracy is the wrong metric when failures are rare: a model that predicts no failure ever is right almost all the time. What matters is precision and recall at the threshold you would actually run, and who sets that threshold. A crew dispatched three times for nothing stops believing the fourth alert, and that is a real cost that no supplier puts in a proposal.
Ask: at the recommended threshold, how many alerts per hundred units per year and what fraction of them were real, on data the model had not seen?
6. Fleet learning, and what happens with your unusual units
Models that learn across a fleet get better as it grows, which is an advantage until you get to the two units of an old model that nobody else has. Ask what happens there, and ask what happens after a firmware update changes the meaning of a signal. Then ask where the output lands: a work order in the maintenance system, a warranty claim with evidence attached, or a dashboard that someone has to remember to open.
Ask: how does the system behave on a model with very few units in the fleet, and how does its output reach the maintenance system?
How to run the comparison
Pick a period from your own history that contains at least one real failure, remove it, and ask each candidate to run blind on the preceding data. Keep the failure dates to yourself until the answers are in. Score against the current practice, whether that is a fixed calendar or reaction to alarms, and count the false alarms as carefully as the hits. A supplier who is comfortable with a blind test on your data is telling you something before the results arrive.
Disclosure
We build one of these systems, InverterAI. The criteria above are the ones we would want a buyer to apply to us, including the ones about false alarms and cold start.