Comment by cpgxiii
2 hours ago
The "naive notion of system performance" is that the system is primarily operating in the ideal regime, and thus all deviations deserve alerts (and receive attention). The reality is that often the system is operating often enough in some failure mode that any such alerts would be regularly triggered by accepted behavior, and thus the alerts are systematically ignored/disabled/normalized.
E.g. nominally you should never be mixing traffic types (aircraft and helicopters, civilian and military) in close proximity to a major airport and in a regime where TCAS is unlikely to offer sufficient protection. So in theory, any mixing should immediately trigger an alert and investigation to develop new procedures. But in practice, if you routinely allow such mixing under what you believe are "safe" practices (and get away with such mixing for a long time) then when a real accident happens you will have plenty of "proto-accidents" to look back on, but the warning signs from those near-accidents will have become accepted practice.
My response would be to use more sophisticated anomaly detection.
Problems managing a complex system? Why not consider adding another complex system!