Ask a reliability engineer what they do with the monitoring system and you will often get a version of the same answer: they check it when something else has already gone wrong. The dashboard is not off. It is just no longer the first thing anyone looks at.
That is what a 70 percent false positive rate buys. It is roughly what legacy reliability systems run at, and it is not a tuning problem. It is the reason programs get funded, deployed, and then quietly abandoned with the models still technically running.
The math of ignoring things
Two out of three alerts are noise. An engineer investigates the first, the second, the third. By the tenth, the rational move is to wait for a second signal before spending an hour on any of them. That is not negligence, it is triage under a bad prior.
The cost is not the wasted hours. It is the real detection that arrives in month seven and gets treated exactly like the six false ones before it. A program with a 30 percent hit rate does not deliver 30 percent of the value. It delivers close to zero, because nobody is acting on the output.
Why detectors cry wolf
Most false positives are not model errors. They are context errors.
The equipment is behaving differently because the plant asked it to. Feed changed, ambient changed, the unit is at turndown, a parallel train came offline, an operator moved a setpoint. A statistical detector sees deviation from a learned baseline and fires. Nothing is wrong.
Second cause: a fixed baseline in a process that never repeats. Catalyst decays, exchangers foul, seasonal ambient swings 25 degrees. Normal moves. A threshold set at commissioning is wrong within months, in both directions.
Third: no failure mode behind the signal. A detector that can say a bearing temperature is elevated but cannot say which degradation mechanism produces that pattern has no basis to distinguish a real precursor from a transient. Every excursion looks the same.
What separates a suppressible alert from an actionable one
Four questions. An alert that cannot answer them should not reach a human.
Is the deviation explained by current operating conditions? If the expected value was computed for the load, feed and ambient the unit is actually running at, most operational noise disappears before anyone sees it.
Is there a mechanism? Not just that a value moved, but which failure mode produces this combination of signals on this asset class, and whether the supporting evidence is present or absent.
Is it progressing? A single excursion is an event. A trend with a slope, on multiple correlated tags, is degradation. The second one is worth waking someone up for.
Does it have a consequence? Criticality, production at risk, and how long the asset can run. An alert with no consequence attached is information. An alert with one is a decision.
An alert that clears all four is worth an engineer's hour. One that clears none is a suppression rule waiting to be written.
Measure the number that decides everything
Most reliability programs report alert counts and detection lead time. Neither tells you whether the system is trusted.
Track precision instead: of alerts raised, what share were confirmed at execution. Track it per asset class, because a program can be excellent on rotating equipment and useless on heat exchangers, and a blended number hides that. Track suppression volume too, since alerts silenced before reaching a human are the difference between a filter and a firehose.
Then track the one that matters most and is almost never captured: how many alerts got investigated at all. Investigation rate is the honest measure of trust, and it falls before anything else does.
Precision is what makes the rest possible
Every downstream benefit of a reliability program depends on someone believing the alert. Lead time is worth nothing if the work order is never raised. Planned-window execution requires a planner who will hold capacity for the finding. Model improvement requires a technician's confirmation coming back.
None of that happens at 30 percent precision. All of it becomes routine somewhere north of 80. That single number is not one metric among many. It is the one that determines whether the program is a decision system or a dashboard nobody opens.
