Over ten years ago, my employer was spinning up a new DC across the state and had three links between it and the primary DC. Two were pretty direct, but we needed a third because at one point in the 300 mile path, the two main links went within 400 meters of each other.
So how national-security-adjacent critical systems like an airport system doesn't have a larger set of backup links, and validates that they are geographically separate up until the connections, and have realtime alerting on the connection status, is surprising to me.
Similar events happened in Europe disguised as thieves stealing fiber optic. This does not have any sense economically, as the value in the market is zero so... either is an honest accident and is cleared in a few days, or is sabotage
That's very uncharacteristic, especially given all those contingency requirements (backup policies, failovers, testing schedules etc.) imposed on corporations/companies in the wake of 9/11.
You actually have to ensure that someone isn't just checking boxes the tests passed and the tests actually passed.
I worked for a regional ISP that had a major outage when the redundant fiber provided by the telephone company was cut in one place triggering a full loss of connectivity. It also caused a massive 911 outage for 200,000 people as it isolated the 911 center from the city core.
Turns out the phone company didn't connect one side of the ring topology even though they certified they did. Needless to say lawsuits abounded.
I've worked on backup/failure systems since the mid-90s, and I've found there's one universal truth: If you don't fully test your backup/failure system, you don't have a backup/failure system.
There's generally two wrong responses: (1) We spent a lot of money on 'blah blah blah', a lot of other companies use it, so yeah, we've got a backup/failure system. And, (2) inadequate testing - either, we tested 1 of 50 services, and it worked, so the whole system can be restored; or, we gracefully tested, and it worked, so it will obviously work during not-graceful incidents.
And the root cause of this is generally that no one gets promoted for implementing an adequate backup/failure system, or it's extremely rare.
Besides what the others have said, speaking from a neteng perspective, sometimes backup lines (and the core infrastructure in general) are engineered in such a way that it can't easily be tested properly without taking other things down, or manually rolling a truck specifically to test it in isolation with extra equipment.
Not saying that's what is going on here, just that it's possible.
Building multiple diverse paths and monitoring for fiber cuts is not hard. Having 2 fiber paths is insufficient for even only moderately important workloads at a tech company. Overlapping fiber cuts happen. For something with significant economic and safety impact this is just crazy.
Sometimes the level of incompetence / lack of care in organizations like this astounds me. I understand issues like this can be complicated and systemic but it honestly makes me think very poorly of the technologists building these systems in government.
They had two diverse paths, but apparently did not have a (reliable?) process for ensuring the backup path was functional during fail-over. I imagine there will be some soul-searching over that.
Nothing I've ever seen or experienced with FAA indicates a lack of care; the parsimonious explanation is almost always that people who care a great deal are working within complex systems that don't always have a consistent or externally legible set of priorities. Or more intuitively, the failures we see are the "acceptable" ones versus the unacceptable ones (like planes falling out of the sky).
> Nothing I've ever seen or experienced with FAA indicates a lack of care
Hard to fathom someone still holding that view after the 737MAX disasters.
The NTSB has repeatedly and publicly criticized the FAA for failing to implement their recommendations, including as recently as last week (Amazon Prime crash resulting in 5 deaths). They've explicitly called out the FAA's failure to act as a contributing factor in crashes, multiple times.
The problem is with considering the second path a backup path. It is a common problem with things that are considered to be a backup. In contrast, both paths should ideally have been in regular use. As a general policy, for resiliency, one should always exercise all of one's routes/suppliers/vendors, even if some are suboptimal, not just the primary one.
I've told this story so many times I've worn the corners off it. BT was giving a presentation to MSN WAN OPS about the new datacenter buildout in London for us. They kept going on and on about physical security, man traps, and that our fiber left the building on each side and didn't get close to each other for so many km away.
BT guy ends that part of the presentation with "the IRA will really have to get shit together to take you off the net"
This was during the troubles. Same planet, different world.
That sounds pretty plausible. But you still have to have enough internal competency to oversee a contractor, and ask the right questions/build proper requirements and verify they are meeting them.
Finding out your backup fiber has been cut, aparently some time
ago, only when you try to use it as the backup isn’t an issue of “aging infrastructure.” It’s an issue of letting completely incompetent people run your tech.
Over the weekend, I got linked to https://admiralcloudberg.medium.com/reaping-the-whirlwind-in..., which is an exhaustive look at the causes of the first airspace collision in the US in several decades. After reading that, my conclusion is that, if the software being rolled out is what I think it is, it would be a borderline impeachable offense to not first trying it out in the DC airspace.
The fundamental problem with DCA, and one of the two main causes of the crash [1], is that they are required by political pressure to operate at a higher operational tempo than they can safely operate at. DCA has essentially 1½ usable runaways--large planes can only use the larger runway, and given the capacity restrictions, airlines have been pushing to use fewer small planes at the airport. This shift in plane size means the effective safe slot capacity has gone down, but people still keep citing the same number as justification for safe numbers, and the politicization of the issue has shut down everyone who complained that the actual traffic just couldn't be safely handled.
If DCA's theoretical slot capacity (36 landings and takeoffs each per hour) is to be reached one just one runway, you have about 100 seconds to go from plane 1 touchdown; exit the runway, letting plane 2 on to take off; plane 2's wheel leaving the runway, clearing plane 3 touchdown. That's doable, but with essentially 0 margin for error. Alternating the runways used for each landing would give you closer to 30s of margin, but if only 10-20% of the planes can use the alternate runway, you can't divert enough planes to use the alternate runway.
Now DCA and the FAA aren't stupid enough to actually schedule 36 landing slots every hour, but the problem is that because of a few various factors, what was scheduled as, say, 30 slots for an hour (giving 120 seconds between landings, probably sufficient margin) ends up being 12 slots used in the first half-hour and 18 slots used in the second half-hour, which means a lot of the actual operation ends up having no safety margin even though on paper you have sufficient margin.
One of the ways you can rectify that is to assign slots further in advance so that you don't get the bunching. Effectively saying "oh, if you leave right now, you'll arrive at 3:30 with three other planes, but if I hold you for 10 minutes, I can push you into a less busy arrival time." Doing this requires good, accurate prediction of the actual flight travel times, and my understanding is that this is what the new software is meant to provide.
[1] The other main cause is essentially that the DoD's aviation practices in the area is a giant clusterfuck that endangers lives, and unfortunately that sentence is not relegated to the past tense.
I don't quite understand this sort of thing happening. Wasn't the whole point of the internet to be a self-healing network where we route around severed cables, etc?
Or is it that these ATC networks are their own air-gapped network with less redundancy? That just doesn't add up. Or maybe there was only one line going to the ATC, with no multiple "ISPs" like a datacenter would have?
* Yes, TRACON have their own dedicated links. You can look into ASTERIX and STARS to learn some of the cursed ways the data processing and dataflow work.
* There is supposed to be a primary and a secondary link, in this case the primary failed and the fail-over also failed. It's unclear from the reporting if they were damaged in the same incident or if the failover was not tested or monitored adequately.
The Philadelphia TRACON site has been notoriously unreliable and was supposedly improved in 2025, it's also unclear if these issue actually could stem from that implementation.
I’ve had enough double-redundant links fail - and I mean pretty thoroughly double-redundant, different provider, different media, different physical path - that I’m quite surprised something like the air traffic control system only has two links.
Not to be conspiratorial, but about an hour ago I saw this in The Economist: "Russia’s grey-zone attacks on Europe are growing more brazen"
[QUOTE:]
Russia has been staging covert operations and deniable provocations against European countries for many years. Since its full-scale invasion of Ukraine began in 2022 they have increased. These have included incursions by reconnaissance drones, hacking, arson, parcel bombs targeting cargo flights and plots to kill executives of leading European arms firms. In recent months the pace of such shadowy attacks appears to have risen again. The Russian operations have three main goals, according to analysts and security officials: trying to coerce European countries into halting their support for Ukraine; raising the costs of providing that support; and attempting to make NATO look powerless.
[/QUOTE]
And of course there've been "suspicious" cable cuts in the Baltic Sea and other European waters ....
As a backup for when ground communications are cut? As opposed to having no backup in that situation? Yeah, maybe we should, no matter who controls the company.
Fiber was all installed long after we knew locating it was important. I suspect it all had a tracer wire installed with it. Of course sometimes the tracer wire can break.
Before fiber a lot of things were put into the ground with no thought how you would find it again.
Indeed, tracer wire is usually bonded to a (sacrificial) magnesium anode, but it’s a crapshoot whether the tracer wire is still intact.
I sometimes hire directional boring and excavation contractors and if there’s any doubt as to where an electrical conduit or natural gas pipe is, I opt for the hydrovac truck to minimize risk.
Always bring a length of fiber optic cable when you're in the wilderness. If you get lost or stranded, bury the cable. Within a few hours, a backhoe will stop by to dig it up, and the crew will rescue you.
Local people know where they are because once upon a time a D8 drove through pulling a big orange spool of fiber conduit and passers by chatted about how another fiber was going in. Also around here the contractors seeded the disturbed ground with a wildflower mix (presumably to discourage weeds) so you can also look for straps of wildflowers.
And of course there are marker posts that say "buried fiber optic cable do not dig".
Once you get out farther from civilization these signs become less common and there are more situations where they can get removed.
Most cheap hired labor doesn't realize you shouldn't dig up the market posts and throw them in the dumpster on construction sites. Add a little bit of rain and suddenly the contractor showing up with an excavator that was told everything was properly marked, and can see markings in other places but not where they are digging, and disaster occurs.
Also soils aren't static, I've been on sites where the midpoint of a buried telecommunications cable had drifted over 8 feet from where the markers were a few hundred yards apart. Just looking at the markers and assuming a straight line isn't a safe bet. Looking at the fence rows to the left and right of the cable and you could see a matching bow in the fence rows.
The UK ATC system was down again today for the second time this month. I'm no conspiracy theorist but hard to believe those two + this are 'accidents'. I believe the Netherlands has also had several issues recently.
So of course if the FAA is going to put AI into the watchtowers, they're obviously going to use local AI and not rely on some random Amazon cloud infra in some middle eastern country.
right. of course, no ones going to cut corners on the safe use of AI and secure infrastructure in this administration.
> When it went to flip into the backup, we discovered that the backup fiber had a break
Pretty grim that a life critical system wasn't designed to report that the backup fibre was unserviceable until they attempted to switch over to it.
I wonder how long it was down? Days, weeks, months?
All these cable/internet service companies are incompetent as it gets.
I am paying, $1000, $1800 & $1900 for the same service at 3 different location (20 mins from each other).
Two locations, I also have old coax lines that are still active, but not paying for it.
When I bought two businesses, I learned that they were paying for a dedicated fiber but using coax service.
At one of the location, we had fiber, and paying for backup coax and wireless. But if you turn off fiber box, it wont fail over to either one.
I wouldn't surprise it was down for weeks and no one bothered about it.
[dead]
Sometimes your multiple fiber paths end up in the same bundle, severed by the same backhoe. It's always a fun day when that becomes apparent.
Put all your backups in one basket, and then pray that nobody crushes the basket.
1 reply →
last time this happened in my area a local farmer was burying a cow and took out the whole towns connection
1 reply →
That is negligence, practical if not contractual.
Over ten years ago, my employer was spinning up a new DC across the state and had three links between it and the primary DC. Two were pretty direct, but we needed a third because at one point in the 300 mile path, the two main links went within 400 meters of each other.
So how national-security-adjacent critical systems like an airport system doesn't have a larger set of backup links, and validates that they are geographically separate up until the connections, and have realtime alerting on the connection status, is surprising to me.
1 reply →
Could this be sabotage?
That seems more than merely plausible given current tensions and the location/timing connection with NYC and UN.
I would put my money on it.
Similar events happened in Europe disguised as thieves stealing fiber optic. This does not have any sense economically, as the value in the market is zero so... either is an honest accident and is cleared in a few days, or is sabotage
2 replies →
trump's escalator revenge?
Yes please. Russia did it. Please move the patriotic war to the Bering Strait (Strait of America?) and leave Europe alone.
That's very uncharacteristic, especially given all those contingency requirements (backup policies, failovers, testing schedules etc.) imposed on corporations/companies in the wake of 9/11.
You actually have to ensure that someone isn't just checking boxes the tests passed and the tests actually passed.
I worked for a regional ISP that had a major outage when the redundant fiber provided by the telephone company was cut in one place triggering a full loss of connectivity. It also caused a massive 911 outage for 200,000 people as it isolated the 911 center from the city core.
Turns out the phone company didn't connect one side of the ring topology even though they certified they did. Needless to say lawsuits abounded.
2 replies →
A lot of people don't check their backups until they need to restore.
A lot of people are incompetent.
Nobody wants backups as a feature. The feature is restore.
1 reply →
A backup without a restore test isn't a backup at all
3 replies →
I've worked on backup/failure systems since the mid-90s, and I've found there's one universal truth: If you don't fully test your backup/failure system, you don't have a backup/failure system.
There's generally two wrong responses: (1) We spent a lot of money on 'blah blah blah', a lot of other companies use it, so yeah, we've got a backup/failure system. And, (2) inadequate testing - either, we tested 1 of 50 services, and it worked, so the whole system can be restored; or, we gracefully tested, and it worked, so it will obviously work during not-graceful incidents.
And the root cause of this is generally that no one gets promoted for implementing an adequate backup/failure system, or it's extremely rare.
Besides what the others have said, speaking from a neteng perspective, sometimes backup lines (and the core infrastructure in general) are engineered in such a way that it can't easily be tested properly without taking other things down, or manually rolling a truck specifically to test it in isolation with extra equipment.
Not saying that's what is going on here, just that it's possible.
Building multiple diverse paths and monitoring for fiber cuts is not hard. Having 2 fiber paths is insufficient for even only moderately important workloads at a tech company. Overlapping fiber cuts happen. For something with significant economic and safety impact this is just crazy.
Sometimes the level of incompetence / lack of care in organizations like this astounds me. I understand issues like this can be complicated and systemic but it honestly makes me think very poorly of the technologists building these systems in government.
They had two diverse paths, but apparently did not have a (reliable?) process for ensuring the backup path was functional during fail-over. I imagine there will be some soul-searching over that.
Nothing I've ever seen or experienced with FAA indicates a lack of care; the parsimonious explanation is almost always that people who care a great deal are working within complex systems that don't always have a consistent or externally legible set of priorities. Or more intuitively, the failures we see are the "acceptable" ones versus the unacceptable ones (like planes falling out of the sky).
> Nothing I've ever seen or experienced with FAA indicates a lack of care
Hard to fathom someone still holding that view after the 737MAX disasters.
The NTSB has repeatedly and publicly criticized the FAA for failing to implement their recommendations, including as recently as last week (Amazon Prime crash resulting in 5 deaths). They've explicitly called out the FAA's failure to act as a contributing factor in crashes, multiple times.
Regulatory capture ruined the FAA.
The problem is with considering the second path a backup path. It is a common problem with things that are considered to be a backup. In contrast, both paths should ideally have been in regular use. As a general policy, for resiliency, one should always exercise all of one's routes/suppliers/vendors, even if some are suboptimal, not just the primary one.
1 reply →
[flagged]
I've told this story so many times I've worn the corners off it. BT was giving a presentation to MSN WAN OPS about the new datacenter buildout in London for us. They kept going on and on about physical security, man traps, and that our fiber left the building on each side and didn't get close to each other for so many km away.
BT guy ends that part of the presentation with "the IRA will really have to get shit together to take you off the net"
This was during the troubles. Same planet, different world.
I’m pretty sure that type of project is contracted to the private sector and isn’t caused by government technologists in any way…
That sounds pretty plausible. But you still have to have enough internal competency to oversee a contractor, and ask the right questions/build proper requirements and verify they are meeting them.
Finding out your backup fiber has been cut, aparently some time ago, only when you try to use it as the backup isn’t an issue of “aging infrastructure.” It’s an issue of letting completely incompetent people run your tech.
There is also a new ATC system being rolled out, as early as today
https://www.airwaysmag.com/new-post/faa-smart-first-deployme...
There’s been “a new ATC system” rolling out for the last 25 years.
Over the weekend, I got linked to https://admiralcloudberg.medium.com/reaping-the-whirlwind-in..., which is an exhaustive look at the causes of the first airspace collision in the US in several decades. After reading that, my conclusion is that, if the software being rolled out is what I think it is, it would be a borderline impeachable offense to not first trying it out in the DC airspace.
The fundamental problem with DCA, and one of the two main causes of the crash [1], is that they are required by political pressure to operate at a higher operational tempo than they can safely operate at. DCA has essentially 1½ usable runaways--large planes can only use the larger runway, and given the capacity restrictions, airlines have been pushing to use fewer small planes at the airport. This shift in plane size means the effective safe slot capacity has gone down, but people still keep citing the same number as justification for safe numbers, and the politicization of the issue has shut down everyone who complained that the actual traffic just couldn't be safely handled.
If DCA's theoretical slot capacity (36 landings and takeoffs each per hour) is to be reached one just one runway, you have about 100 seconds to go from plane 1 touchdown; exit the runway, letting plane 2 on to take off; plane 2's wheel leaving the runway, clearing plane 3 touchdown. That's doable, but with essentially 0 margin for error. Alternating the runways used for each landing would give you closer to 30s of margin, but if only 10-20% of the planes can use the alternate runway, you can't divert enough planes to use the alternate runway.
Now DCA and the FAA aren't stupid enough to actually schedule 36 landing slots every hour, but the problem is that because of a few various factors, what was scheduled as, say, 30 slots for an hour (giving 120 seconds between landings, probably sufficient margin) ends up being 12 slots used in the first half-hour and 18 slots used in the second half-hour, which means a lot of the actual operation ends up having no safety margin even though on paper you have sufficient margin.
One of the ways you can rectify that is to assign slots further in advance so that you don't get the bunching. Effectively saying "oh, if you leave right now, you'll arrive at 3:30 with three other planes, but if I hold you for 10 minutes, I can push you into a less busy arrival time." Doing this requires good, accurate prediction of the actual flight travel times, and my understanding is that this is what the new software is meant to provide.
[1] The other main cause is essentially that the DoD's aviation practices in the area is a giant clusterfuck that endangers lives, and unfortunately that sentence is not relegated to the past tense.
I suspect it is vibe coded. Contract won in June. Deploying in September.
Good luck to all of us.
I can guarantee you that it was not. Check out DO-278A, and its standards and guidelines that must be followed.
1 reply →
Deploying first to DC? Yeah, that’ll end well.
1 reply →
I don't quite understand this sort of thing happening. Wasn't the whole point of the internet to be a self-healing network where we route around severed cables, etc?
Or is it that these ATC networks are their own air-gapped network with less redundancy? That just doesn't add up. Or maybe there was only one line going to the ATC, with no multiple "ISPs" like a datacenter would have?
* Yes, TRACON have their own dedicated links. You can look into ASTERIX and STARS to learn some of the cursed ways the data processing and dataflow work.
* There is supposed to be a primary and a secondary link, in this case the primary failed and the fail-over also failed. It's unclear from the reporting if they were damaged in the same incident or if the failover was not tested or monitored adequately.
The Philadelphia TRACON site has been notoriously unreliable and was supposedly improved in 2025, it's also unclear if these issue actually could stem from that implementation.
I’ve had enough double-redundant links fail - and I mean pretty thoroughly double-redundant, different provider, different media, different physical path - that I’m quite surprised something like the air traffic control system only has two links.
2 replies →
It is, but you need to set it up correctly.
Do you want air traffic control to be on the general internet?
That seems like an extremely foolhardy thing to do.
As a tertiary backup: Yeah, maybe I do want that.
It seems like it would present a less-chaotic solution than that provided by having no data communications at all.
Why not? Assuming it's competently set up - serious encryption, details not blabbered about (to attracted DDoS or whatever), etc.
You certainly don’t need to run this over public Internet to get resiliency over multiple paths.
I thought Musk got an FAA telecom contract for Starlink back when he was effectively running the government (or at least running it into the ground)?
https://www.cnn.com/2025/02/25/business/musk-faa-starlink-co...
Not to be conspiratorial, but about an hour ago I saw this in The Economist: "Russia’s grey-zone attacks on Europe are growing more brazen"
[QUOTE:]
Russia has been staging covert operations and deniable provocations against European countries for many years. Since its full-scale invasion of Ukraine began in 2022 they have increased. These have included incursions by reconnaissance drones, hacking, arson, parcel bombs targeting cargo flights and plots to kill executives of leading European arms firms. In recent months the pace of such shadowy attacks appears to have risen again. The Russian operations have three main goals, according to analysts and security officials: trying to coerce European countries into halting their support for Ukraine; raising the costs of providing that support; and attempting to make NATO look powerless.
[/QUOTE]
And of course there've been "suspicious" cable cuts in the Baltic Sea and other European waters ....
https://www.economist.com/europe/2026/09/21/russias-grey-zon...
Well.. ATC in the north of the Uk was down today as well. Fun.
ATC seems like a great place to have redundancy with a starlink dish on the roof.
Put two for redundancy
Ah yes let’s make our entire infrastructure depend on a company controlled by a single person.
As a backup for when ground communications are cut? As opposed to having no backup in that situation? Yeah, maybe we should, no matter who controls the company.
It has been conjectured that the easiest way to locate a buried fiber optic cable is to start digging.
(The actual easiest method is to install the conduit or direct burial cable with tracer wire)
Fiber was all installed long after we knew locating it was important. I suspect it all had a tracer wire installed with it. Of course sometimes the tracer wire can break.
Before fiber a lot of things were put into the ground with no thought how you would find it again.
Indeed, tracer wire is usually bonded to a (sacrificial) magnesium anode, but it’s a crapshoot whether the tracer wire is still intact.
I sometimes hire directional boring and excavation contractors and if there’s any doubt as to where an electrical conduit or natural gas pipe is, I opt for the hydrovac truck to minimize risk.
Always bring a length of fiber optic cable when you're in the wilderness. If you get lost or stranded, bury the cable. Within a few hours, a backhoe will stop by to dig it up, and the crew will rescue you.
Local people know where they are because once upon a time a D8 drove through pulling a big orange spool of fiber conduit and passers by chatted about how another fiber was going in. Also around here the contractors seeded the disturbed ground with a wildflower mix (presumably to discourage weeds) so you can also look for straps of wildflowers.
And of course there are marker posts that say "buried fiber optic cable do not dig".
Once you get out farther from civilization these signs become less common and there are more situations where they can get removed.
Most cheap hired labor doesn't realize you shouldn't dig up the market posts and throw them in the dumpster on construction sites. Add a little bit of rain and suddenly the contractor showing up with an excavator that was told everything was properly marked, and can see markings in other places but not where they are digging, and disaster occurs.
Also soils aren't static, I've been on sites where the midpoint of a buried telecommunications cable had drifted over 8 feet from where the markers were a few hundred yards apart. Just looking at the markers and assuming a straight line isn't a safe bet. Looking at the fence rows to the left and right of the cable and you could see a matching bow in the fence rows.
2 replies →
The UK ATC system was down again today for the second time this month. I'm no conspiracy theorist but hard to believe those two + this are 'accidents'. I believe the Netherlands has also had several issues recently.
[flagged]
[flagged]
[flagged]
So of course if the FAA is going to put AI into the watchtowers, they're obviously going to use local AI and not rely on some random Amazon cloud infra in some middle eastern country.
right. of course, no ones going to cut corners on the safe use of AI and secure infrastructure in this administration.
Are we sure this wasn't Claude? Has Dario made an announcement yet?
Maybe Sean Duffy was just binge-watching his old reality tv shows and gobbled up the bandwidth?