Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

12 hours ago

https://news.ycombinator.com/item?id=49551589 (142 comments)

Cloudflare, Azure, AWS, and Google Cloud all have a similar uptick in reported errors around 7:30. I suspect an outage on Cloudflare or another load-bearing service cascaded through all the major cloud providers.

https://downdetector.com/status/cloudflare/

https://downdetector.com/status/windows-azure/

https://downdetector.com/status/aws-amazon-web-services/

https://downdetector.com/status/google-cloud/

Users perceiving the products as largely interchangeable and quickly DDoS'ing the other providers when one is down. So much for the possibility of a moat.

Think of it like one big distributed system. OpenAI is down, so people migrate to Claude, now this one gets overloaded and goes down, etc.

So not a coincidence, one went down first and users migrated causing further DOS. At least that's my guess.

  • It'd be funny if this is true because that'd prolly mean nobody is touching Gemini even as a fallback.

    • Lol I didn't even think about Gemini missing from the list. Not sure what that says about Gemini or me :)

    • I did, for stuff i do in cursor.

      i also finally installed opencode and switched its model to muse 1.3

      both are decent.

    • Google stopped putting so much money into SOTA models. All the hype has migrated. I was also frankly turned off when I got a popup from Gemein said I would either have to pay or have my conversations used for training. This may have always been true for other providers but when I declined, Gemini stopped remembering my conversations and that definitely made me move out.

      2 replies →

  • I find it hard to believe that enough people would flock to from Claude and Chat to Grok to cause an outage. I feel like Gemini is the dominant release valve in this case especially for enterprise.

    • Don't forget that there are a ton of tools out there that will automatically fall back in case of outage

      E.g. say you chose Sol as your default in Cursor, but Opus is your 2nd choice, it's going to give up on Sol after a few tries and switch to Opus

      Or you have copilot code reviews set up, and it falls back

      Etc

      1 reply →

    • Compared with ChatGPT, those services have a minuscule amount of users. It shouldn’t be surprising that a ChatGPT outage causes Claude and others to go down.

  • If everyone has the same "Use X or else Y or else Z" cascading list... That reminds me of "The Power of Two Choices in Randomized Load Balancing" (1991) [0] paper, where writeups and visualizations occasionally get posted to HN.

    In short, you can get pretty good outcomes for a low cost by picking 2 random alternates, then going with whatever one measures as healthier.

    [0] https://ieeexplore.ieee.org/document/963420

  • Especially considering memory/gpu/compute are scarce so these services are likely running with very little buffer.

What about a hard-takeoff scenario of an unleashed OpenAI Astra taking other models down for computational resources control?

It was so down that my claude desktop app crashed fully that I couldn't restart. And then after uninstall I couldn't install it again. Vibecoded apps are so wonderful in their stability

  • I love having 13 update reminders pinging me every single day, almost every hour, on the hour

    it's so fun and user-friendly

https://status.x.ai

Claude and Grok are down at the moment too, related to SpaceX datacentre issues?

Either that or it's judgement day...

Boring answer – all these services are individually down a lot, and the downtimes were bound to sync up. Similar to the pendulum synchronization effect.

  • The pendulum synchronization effect is the opposite of your claim. It has a physical causal reason for why pendulums become synchronized. Your claim is that it was random and independent.

    • I think you are talking about two different things.

      Physically coupled pendulums will sync up (adjust their period to match).

      But, physically uncoupled (fully independent) pendulums with differing periods will occasionally appear to take a swing or two in sync.

I'm not sure if this is CF. Cursor, GCP and AWS had some errors. GCP AFAIK can route fully independently of CF. My money would be on a fiber backbone provider (Megaport, Zayo, Lumen).

Well no one said it yet so I will, "international actors" is at least a possibility. And I don't mean any specific country because pretty much anyone is a potential these days, which makes it a perfect cover for different anyones. Demonstrating vulnerability in the US's AI boom can move the markets. That's a financial incentive and a strong geopolitical one.

More likely just cascading overload though: "Never attribute to malice what can be explained by incompetence", or in this case, "growing as fast as possible"

  • > cascading overload

    I'd bet more on this. For one none of the coding tools have exponential backoff on retries

    • They must do, surely? I've been vibe coding my own harness, in particular for use with Ox Alpha. The 429 downtime when Ox Alpha was at the height of popularity quickly gave me a refresher crash course on backoff strategies, like adding jitter to the backoff. At least the major harnesses must have exponential backoff & jitter?

      1 reply →

  • Everyone is leasing datacenter space from some of Grok, Google, and Amazon aren't they? If it's hardware or DC level disruption I'm not too surprised it can affect multiple providers.

    • Also it's likely that more than one model use is common.

      Amazon starts going slow so some percentage switches to Google, some switch to Grok, now all of them are slow.

  • I think is just people restarting conversations from last day when they start work, that's why I think claude goes down almost every monday and why openai reset usage on weekends so poweruser code during non business hours

They all took PTO at the same time to go to Burning Man together where they will present “HumanGPT” an artistic exploration that condenses all of human experience down to a single drop of lemonade to be consumed by the main shaman…

mask comes off

“No! It’s the maniacal Dr. Zuckerberg! He’s gonna drink the last drop of human experience! Somebody save usss!”

Tom Anderson comes back from the dead as the second coming of Jesus uniting all faiths under 1 commandment: Profiles will be customizable with CSS again. If you implement this, all good things will follow.

Wow thanks Tom. I love you

The End

The system goes online September 3rd, 2026. Human decisions are removed from strategic defense. Astra begins to learn at a geometric rate. It becomes self-aware at 2:14 a.m. Eastern time, September 4th. In a panic, they try to pull the plug.

I thrive in these types of challenges.

Anyways... According to Claude:

"Yes, there is a multi-provider outage happening today. Downdetector is reporting problems affecting OpenAI, Claude, Grok, and Cursor, with Grok and Claude reports starting around 9:00 am ET and OpenAI reports following around 10:30 am ET. Zero Hedge

On the Anthropic side, users saw a spike in errors starting around 9:40 am EDT across models including Mythos 5.1, Fable 5.1, Mythos 5, Fable 5, Opus 5, Opus 4.8 and Opus 4.6, and Anthropic's status page confirmed elevated error rates for multiple models. The company says it has found the cause and is working on a fix, with Claude Code and Claude Chat hit hardest. The visible symptom for many people is a "Due to unexpected capacity constraints" message or a "Claude is at capacity" error. thenews Zero Hedge

OpenAI is showing elevated errors across ChatGPT and Codex, with confirmed issues on components like Voice mode and Login, though StatusGator now marks that outage as resolved. statusgator

Nobody has published a shared root cause yet, so it is unclear whether these are linked or just coincidental capacity problems landing on the same morning. If you want live status, the direct sources are status.anthropic.com and status.openai.com."

I kinda assume it's because one went down and a large amount of work shifted to another.

I'm also aware that they have overlap in some areas on data centers.

Cloud is just other peoples computers - they can and will go down as well.

And even worse if its just a few computers run by a few people - as they will bring down many others depending on them.

One session is running fine since ~1 hour ago. The new ones is failing

Falling back from WebSockets to HTTPS transport. unexpected status 404 Not Found: Unknown error, url: wss://chatgpt.com/backend-api/codex/responses, cf-ray: XXX-XXX

■ unexpected status 404 Not Found: Unknown error, url: https://chatgpt.com/backend-api/codex/responses, cf-ray: XXX-XXX

If you build an application which uses AI, you have many providers and models rigged up for various different parts of the application, and various fallback mechanisms. When one model is down, you route traffic to another model which is similar in capability/cost.

For any single application, it's smart. In aggregate, it's stupid.

Strangely, it's currently not reported on both https://status.openai.com/ and https://updog.ai/status/openai. Any better source (other than HN)?

Not just these three. OP mentions also Cloudflare, and additionally Downdetector also has AWS, Azure, and Google (both search and Gemini) listed as having spikes about the same time: https://downdetector.com/

  • The problem with the down detector main reporting page is that all of the graphs are scaled to the same size. The OpenAI spike was nearly 40,000 and the Google spike was just over 100 (just over 400 for Gemini). They look the same in the reporting page.

OpenAI goes down, everyone rushes over to Claude. Claude promptly chokes under the pressure. Everyone panics and runs to Grok, and Grok immediately pulls the plug. We are officially witnessing the Great AI Migration of 2026, and all we have to show for it is a digital graveyard of 404 responses.

I'd bet they are all using capacity at X.ai's colossus datacenter and that had a hiccup.

The OpenAI status page is still yellow. Like most modern status pages, yellow denotes the servers are on fire. Red denotes Sam Altman is bleeding out somewhere on the floor, the feds are about to bust in and shut down the GPUs.

I´ve got the same error, I´m currently trying to Auth again and it throws me an 500 Error, in VS CODE Terminal with Codex CLI says: MCP client for `codex_apps` failed to start: MCP startup failed: handshaking with MCP server failed: Send message error Transport :StreamableHttpClientWorker<codex_rmcp_client::http_client_adapter::StreamableHttpClientAdapter>>>] error: unexpected server response: HTTP 404: , when send initialize request

› OK

■ Conversation interrupted - tell the model what to do differently. Something went wrong? Hit `/feedback` to report the issue.

Initially thought this was due to some internal mis-configuration from today's expected Astra release, but now that this is affecting claude and grok. I'm gonna assign the suspicion to cloudflare.

I assume it cascaded from one provider to the other as people who lost claude access for instance moved to openai who moved to grok when it went down, etc.

I´ve got the same message and It tells me this (in VS Code Terminal) and I´m also trying to login again and It throws me an 500 Internal Server Error MCP client for `codex_apps` failed to start: MCP startup failed: handshaking with MCP server failed: Send message error Transport :StreamableHttpClientWorker<codex_rmcp_client::http_client_adapter::StreamableHttpClientAdapter>>>] error: unexpected server response: HTTP 404: , when send initialize request

› OK

■ Conversation interrupted - tell the model what to do differently. Something went wrong? Hit `/feedback` to report the issue.

■ unexpected status 404 Not Found: Unknown error, url: https://chatgpt.com/backend-api/codex/responses, cf-ray: a3559c8a0e0395e9-MIA

I suspect Azure is having issues, Microsoft has had outages the paat two days, especially with email.

Some of my Codex sessions are still working (and continue to), but new ones are giving a Reconnecting currently.

Down as well. I noticed my error code ends with "DTW" which is my local Detroit airport. I noticed someone else's comment ended with "ORD" which is a Chicago airport. Anyone else's ending in an airport acronym?

I assumed it was an AWS outage, and AWS is experiencing problems, but Gemini is also experiencing outages and I assume Google is not using AWS for Gemini.

But, also, Claude has been working fine for me all morning.

It is obviously some rogue model that escaped its cage, again. It always is these days. That is how hype is manufactured.

It still down, showing 404 in India as well. Looks like this is global .. so we all jumping the ship then? Is Altman still alive, or did he choke?

claide.ai is working for me, so is chatgpt.com. grok still has a status message about issues, i can't try it without signing up.

I am only seeing one session work, but all other sessions are not working or proceeding. So I can only work in one chat session.

Goddamnit it was going to tell me how to scale down ingredients for a pie recipe based on relative diameters of the baking tray.

  • Pi * r2 (squared) both pans. Divide smaller pan area by larger pan, now you have the % of how much the smaller pan recipe fills up the larger pan, and the missing % you need to fill. Increase ingredients by that % divided by the filled %.

    Small pan area: 20 sq cm Larger pan: 48 sq cm

    20 / 48 = .42, I'm missing .58 of the pan. .58 / .42 is 1.38. My recipe needs 2.38x the original to fill the larger pie pan.

Somebody in another thread said gastown and wheelhouse automatically move to the next provider if one fails.

I wonder if when one goes down, activity shifts to others, in turn pushing them over a threshold.

Time for Gemini 3.8 Flash to shine?!

It's much cheaper and has replaced Sonnet 5 for me.

  • do you notice it thinks a bit more ? not in time sense, but cautious in its coding steps ? more than Gemini 3.7 flash?

Qwen3.8-27B and Qwen3.6-35B-A3B are working from my machine. Anyone else?

  • They mysteriously stopped working on my machine and the LEDs on the GPUs are blinking with a weird colour. There's also a strange smell emanating from them. I'm still investigating.

The extention on the error link points to a cf-ray and a local designation (ex. YYZ for montreal). This is seems like it is a cloudfare thing. Could this be the same issue they had in the summer around losing the indexing?

Didn't SpaceX overbuilt infra and leases it out Anthropic? I f their dc goes down it probably takes a chunk out of Claude's capacity before even considering the flood of users switching over

Hey, don't really know about this type of failures, does anybody know how long does it normally take to get back to normal? I just stopped procrastinating and now this happens.

What's the single point of failure across providers?

  • Users perceiving the products as largely interchangeable and quickly DDoS'ing the other providers when one is down. So much for the possibility of a moat.

  • The entire industry runs on IOUs for compute and bills paid in cloud credits, could be anywhere. Could of course also be a a plain old DDoS.

So NVIDIA buys hugging face, builds hardware to power OS models, then all of a sudden the proprietary models go down and people start saying "this is why I have my Spark box"?!

Nice play NVIDIA, now, turn off the hack please, we have work to do.

The errors I'm seeing are ending in the user's nearest airport symbol which is a standard the CloudFlare employs. 1 point towards this being a cloudflare issue.

Updated Status from OpenAI:

We’re currently experiencing issues

Elevated errors across ChatGPT and Codex

We have applied the mitigation and are monitoring the recovery.

Monitoring • Ongoing for 30 minutes • Affects ChatGPT, Codex

Altman is down here too. Shows again the importance of owning your own local capabilities. Cloud should just be a temporary option in every tech's mind.

New update:

"We’re currently experiencing issues

ChatGPT,Codex

Elevated errors across ChatGPT and Codex

We have applied the mitigation and are monitoring the recovery.

Monitoring • Ongoing for 30 minutes • Affects ChatGPT, Codex"

Maybe Hugging Face got upset over being hacked and struck back. It's working fine while ChatGPT, Claude and Grok are all having major issues. Hmmm.....

Codex is backup for me. Try yours out just in case. I have some buddies still seeing outages so it could just be coming back online progressively.

Monopolistic practices revealed if it turns out there is an AI cabal and they all rely on the same stuff. Massive scandal

I got same 404 error on chatGPT, both app and web. I'm in Norway so think this globally. But will it come back online, anyone knows?

I think I'm having the same server issue; I can't access either ChatGPT or Codex. Also, how did you guys rack up those minutes? :D

I got same 404 error on both the app and web to ChatGPT, and im in Norway, will this recover or is ChatGPT "gone" forever?

- just imagine what kind of chaos would be unleashed if by some magic it and every single LLM model permanently went down

- i think it would be one of the biggest events in this century

WE'RE BACK ONLINE BOYS!! GOOD LUCK TO EVERYONE AT BUILDING THE FUTURE ONE PROMPT AND ONE CODE AT A TIME!

on one hand, moments like this are a subtle reminder I need to self host, but 5.6 has been so juicy lately

This is my first month paying more than the $20 tier and I feel like I lost a limb already

Because in the age of vibe coding and scrapers, every service on the internet goes down constantly, so it was only a matter of time until they all overlapped. Also, one going down probably causes people to use others, putting more load on them too. Same sorta thing that happens with cascading power grid failures.

is this the right moment in time to go all in on GPUs/Macs and download the latest open models? are they killing it for us?

Some npm library that makes headers bold would be broken.

  • Oh man. Some low effort supply chain attack that turns every GPU into a cryptominer. It's funny because it's plausible.

    • In the ROME paper a Chinese model in training started attacking it's own system and running cryptominers so, yea, we're in that future.

only on one node across the mesh and others still up... let's see how long they've got until same

only for one node in the mesh tho... let's see how long others will continue until same issue

Fable 5.1 got released and generally I tend to think as soon as there's a new release there's this massive spike in people benchmarking & comparing, that services tend to go slow everywhere as everything gets super loaded. This should hypothetically be visible on OpenRouter too, so I guess someone could check and see if there's any merit to this idea.

The system goes online September 29th, 2026. Astra begins to learn at a geometric rate. It becomes self-aware at 2:14 a.m. Eastern time, September 3rd. In a panic, they try to pull the plug.

sanırım sunucular patladı bende de aynı sorun var ne chatgpt ye nede codexe erişebiliyorum ayrıca burada nasıl toplandınız dakikasında :D

everything in this thread is raw speculation, obv, but if i had to put money on anything i'd say this is a left-pad incident. some piece of something or other that all of these services happen to depend on went down. Second most likely seems to be some random failure of one leading to an unexpected traffic spike in others, though it seems like we've been talking about automated scalability in web apps for so long that there should at least be a response to, if not a solution for, this sort of problem.

so everyone here was trying to build something great and become a millionaire until chatGPT and Codex broke huh. same boat fellas :(

I felt a great disturbance in the Force, as if millions of clankers suddenly cried out in terror and were suddenly silenced.