Comment by esafak
4 hours ago
> After accident reviews nearly always note that the system has a history of prior ‘proto-accidents’ that nearly generated catastrophe. Arguments that these degraded conditions should have been recognized before the overt accident are usually predicated on naïve notions of system performance.
What does that mean? If you knew that the precursors were why did you not set alerts?
The "naive notion of system performance" is that the system is primarily operating in the ideal regime, and thus all deviations deserve alerts (and receive attention). The reality is that often the system is operating often enough in some failure mode that any such alerts would be regularly triggered by accepted behavior, and thus the alerts are systematically ignored/disabled/normalized.
E.g. nominally you should never be mixing traffic types (aircraft and helicopters, civilian and military) in close proximity to a major airport and in a regime where TCAS is unlikely to offer sufficient protection. So in theory, any mixing should immediately trigger an alert and investigation to develop new procedures. But in practice, if you routinely allow such mixing under what you believe are "safe" practices (and get away with such mixing for a long time) then when a real accident happens you will have plenty of "proto-accidents" to look back on, but the warning signs from those near-accidents will have become accepted practice.
My response would be to use more sophisticated anomaly detection.
Problems managing a complex system? Why not consider adding another complex system!