Every supplier will show you an accuracy figure, usually one number with a decimal place. On its own it tells you almost nothing: not what was predicted, not on whose data, not whether any of that data came from an asset like yours. The six criteria below are the ones that decide whether a system survives contact with an integrity department. Each ends with the question to put in writing.
No products are named here, ours included. The disclosure is at the end.
1. What it predicts, and at what granularity
A corrosion rate in millimetres per year, a remaining wall thickness, a probability of exceeding a threshold before a given date and a time to a limit state are four different outputs, and they feed four different decisions. Granularity matters as much as the variable: a number for a whole line is not a number for a weld, an elbow, or the low point where water drops out. If the output does not land on the same unit your inspection plan is written against, someone will be translating it by hand forever.
Ask: what is the predicted variable, in what units, over what spatial unit, and which decision is it meant to support?
2. Whether the physics is inside the model or beside it
There are two families. One fits a statistical model to inspection history and operating data. The other constrains the model with the equations that govern the mechanism, which is what a PINN does: electron transfer kinetics, temperature dependence, the thermodynamic driving force. Inside the range covered by the training data both can look equally good. The difference shows up outside it, and outside it is where the decisions that cost money are taken.
Ask: what happens to the prediction when an input moves outside the range seen in training, and can you show me that curve rather than describe it?
3. Which inputs it needs, and whether you actually have them
Read the input list against your own tag list before anything else. Temperature, pressure, flow regime and water cut are usually there. Partial pressures of carbon dioxide and hydrogen sulphide, inhibitor dosing records, sand production and the inspection history in a machine-readable form frequently are not. A model that needs data you do not collect is a data project first and a prediction project second, and the schedule should say so.
Ask: which inputs are mandatory and which optional, and what does the performance become with only the mandatory set?
4. Where the accuracy figure comes from
The same number means very different things depending on how it was produced. Measured on data held out in time, on assets never seen in training, is a strong claim. Measured by cross-validation on segments of the same line is a much weaker one, because neighbouring segments share almost everything. Measured on the training set is not a claim at all. And a point estimate with no interval cannot be used for risk-based inspection, which needs a probability, not a value.
Ask: on what split was that figure measured, how far ahead did it predict, and does the output carry an uncertainty interval?
5. Whether a single prediction can be explained
Sooner or later an inspector, an insurer or a regulator will ask why a segment was deprioritised. The system has to be able to answer with the inputs used, the model version, the mechanism that drove the result and the date it was produced. If a prediction cannot be reconstructed six months later, it cannot support an integrity decision, whatever its accuracy.
Ask: show me the audit trail of one prediction, from raw inputs to the number, as the system stores it.
6. How it fits your operating model
Where does it read data from: historian, SCADA, the laboratory system, inspection records? Where does the output go: a risk-based inspection plan, work orders, a report someone retypes? Who retrains or recalibrates when the process changes, a new inhibitor is dosed or a line is tied in? A system nobody owns after the pilot degrades quietly, and the first sign is that people stop opening it.
Ask: what does handover look like, who recalibrates, on what cadence, and at what cost?
How to run the comparison
Give every candidate the same extract of your own history, with the last period removed, and ask for predictions on that period. Agree in advance what a good result is, in your units, before you see the outputs. Compare against your current method rather than against nothing: if the current method is a constant rate per line, that is the baseline the system has to beat. And make the pilot big enough to include at least one segment that behaved unusually, because average behaviour is not what you are buying protection against.
Disclosure
We build one of these systems, CorrosionAI. The six criteria above are the ones we would want a buyer to apply to us, and we would rather be measured against them than against an accuracy figure with no context.