← Back to context

Comment by bri3d

4 hours ago

* Yes, TRACON have their own dedicated links. You can look into ASTERIX and STARS to learn some of the cursed ways the data processing and dataflow work.

* There is supposed to be a primary and a secondary link, in this case the primary failed and the fail-over also failed. It's unclear from the reporting if they were damaged in the same incident or if the failover was not tested or monitored adequately.

The Philadelphia TRACON site has been notoriously unreliable and was supposedly improved in 2025, it's also unclear if these issue actually could stem from that implementation.

Reporting is saying that the backup was cut a while ago and they only discovered it when the primary failed. Apparently nobody was ping testing the backup link.

I’ve had enough double-redundant links fail - and I mean pretty thoroughly double-redundant, different provider, different media, different physical path - that I’m quite surprised something like the air traffic control system only has two links.

  • Contractually dual-path, or actually dual-path? My understanding is that there's enough infrastructure horse-trading going on behind the scenes that it's very difficulty to be certain that two circuits between points A and B don't share the same infrastructure somewhere in between.

    First big oops of this form that I remember:

    "In December 1986, the ARPANET had 7 dedicated trunk lines between NY and Boston, except that they all went through the same conduit -- which was accidentally cut by a backhoe. "

    https://www.csl.sri.com/~neumann/insiderisks06.html

    • I'm talking about data not able to leave the building, not even data getting lost on the way.

      One example was a site that had fiber and coax, from different companies. They might have shared a pipe at some point along their length, hard to say. But both connections went down at the same time for digital reasons, not physical. The providers had simultaneous unrelated backend router problems and the site lost contact to both gateways.

      This was with a nice SD-WAN system that routed everything dynamically across both pipes to address latency and errors, but if the packets get dropped at the first hop in both networks, there's not much you can do!

      The saving grace in that case was the third redundant connection, a 4G cell modem, with a lot less bandwidth but able to keep the critical transactions going. These days I'd certainly want Starlink on the roof as well.

    • Yup. Seven different paths comes with seven different sets of equipment, and seven different rights-of-way that have to be negotiated, purchased, leased, whatever.

  • One thing that needs to be reviewed is if they are dual links or redundant links.

    In any system where both sides of redundancy always carry traffic then loss of one link can cause congestion failure if any link fails. This is a very common means of failure in electrical networks that requires load shedding. Well, you can't load shed air traffic.

    A more complex system that I like, but comes with it's own set of constraints and implementation issues is a system where both lines carry all the traffic at all times. This way the default state of the system is always working and your first failure isn't invisibly critical.

    But this is very hard as we see in TCP when things get out of order and high latency creeps in. You have to manage a lot more state at the data level.