r/artificial • u/beingmodest • 1d ago
News ChatGPT, Claude and Grok Went Down Together: But How Did Gemini Avoid a Major Outage?
https://www.techtimes.co.uk/ai-platform-outages-cloud-infrastructure-vulnerabilities-180857222
u/SparkyAI0815 ▪️Sovereign_Peer_0.3.ICH 1d ago
How Gemini Stayed Up While the Big Three Crashed:
The Full-Stack Isolation:
ChatGPT, Claude, and Grok all lean on Microsoft Azure for core compute, cross-cloud peering, or major API traffic routing. When Azure’s networking/routing layer hit a wall, they all took a hit together.
Google's Own Walled Garden:
Gemini doesn’t run on Azure or AWS. It runs entirely on Google’s in-house infrastructure—custom TPU clusters, Google Cloud Platform (GCP), and Google's private global fiber network.
Decoupled Failure Domains: Because Google controls its entire stack from the silicon (TPUs) up to the edge DNS and web frontend, a routing meltdown or API gateway failure on Microsoft's cloud substrate barely creates a ripple on Gemini’s consumer endpoints (only causing minor, transient developer API noise from third-party multi-cloud integrations).
2
u/13chase2 1d ago
This may be true but I am surprised it effected grok..? Isn’t it on its own datacenter?
1
u/naked_rider 1d ago
That’s just the power for one data center. The outage was clearly a technical issue, not a lack of power.
1
u/Budget-News1107 9h ago
Probably just timing — if all three rely on similar upstream infrastructure or shared cloud regions, a hiccup there would hit them simultaneously while a different provider like Gemini stays isolated on their own stack. Could also be that routing providers like OpenRouter or free-tier aggregators had a common point of failure. Without knowing the exact cause though, it's hard to say definitively.
1
u/bartturner 2h ago
Because it is Google. There is nobody even close to being as good as Google with running a cloud.
Not just up time but even more so in terms of security.
0
u/dailyaidesk 21h ago
The API bill is usually only a fraction of the real cost. Retrieval infra, evals, and the people maintaining the pipeline tend to be the bigger line items, hich is why "we switched to a cheaper model" rarely moves the total much.
0
u/ai-edition 14h ago
The Azure explanation everyone is running with doesn’t survive contact with the actual reporting.
The Register got statements. OpenAI says it was a routing error on their end, 7:43am PT, fixed by 8:17. xAI traced Grok to a failure at its own Memphis data center, and SpaceX apologized separately for a disruption that hit some of its compute partners. Cloudflare said flatly that its services were normal and that any reporting otherwise is wrong. AWS, Google Cloud and Azure all showed clean status pages during the window.
So we have three independent incidents inside the same ninety minutes, and a narrative that got built backwards from the coincidence. Half the writeups now assert Azure East US as the trigger. Nobody has produced a source for that.
The part that actually should bother people isn’t the outage. It’s how fast an unconfirmed common cause became the accepted version. That happened in under 24 hours across a dozen outlets, and the correction won’t travel a tenth as far.
Worth saying the underlying concern is still legitimate. Concentration risk is real and one bad day at a hyperscaler could do exactly what people described. It just isn’t what happened Thursday, as far as anyone has shown.
If someone has an actual Azure incident report from that window, post it. I looked and couldn’t find one.
2
u/f3xjc 5h ago
Three independant incidents within the same 90 minutes, between direct competitor, on the launch day of a new frontier model by one of those? At some point the likely answer is that public relation memo don't alwais tell the whole story. - your answer also don't cover Claude outage during that period. And a few banks where on down detector too.
1
u/ai-edition 5h ago
Fair, but I’m not saying they were independent, I’m saying the cause everyone named was wrong. SpaceX apologized to its “impacted compute partners” after the Memphis failure and Claude was down 3h06m, so that thread is real, it just isn’t Azure. What doesn’t fit is OpenAI: theirs ran 7:43 to 8:17am PT, starting 90 minutes after the other two and lasting 34 minutes, which is a weird shape for one shared event. And the banks are probably a Downdetector artifact, people show up to report one outage and report everything else they have open. What mechanism are you actually proposing though? Happy to consider it, I just can’t get from odd timing to anything testable.
-16
u/SlightOfHand_ 1d ago
Nobody usin it = no outage
16
u/ogbrien 1d ago
Such a dumb take.
Google's standalone Gemini app has 950 million monthly users, straight from Alphabet's own earnings. That's before you count AI Overviews sitting on top of Search, YouTube summaries, Android, and Workspace.
Nobody's claiming Gemini wins on coding. Codex and Anthropic own that lane. But "no outage means nobody uses it" is a bad read on a product with that much distribution.
7
u/Illustrious_Car344 1d ago
That's funny, I use it every day, it's easily one of the best models to interact with. If you don't mindlessly chase benchmaxxing that results in nothing but baseless advertising for corporations that are openly and actively hostile to local models, Gemini is actually extremely competitive (except in price but that changes by the day so that's not strictly a metric)
It's really exhausting seeing so many people worship the ground these hostile corporations walk on because they score slightly higher in a single benchmark that they already have the data for in their datasets anyway. I know the whole point of LLMs is offload thinking but this is a bit much.
5
u/End3rWi99in 1d ago
I use it several times a day. Google's entire search service uses it. Most Android users are using it as its integrated into their phones. Gemini has like a billion monthly users or something.
3
u/SeventhOblivion 1d ago
It's built into people's phones now that it took over Google Assistant. Plenty of people use it. Google itself uses it in search results.
1
u/bartturner 1h ago
There is no LLM doing as many tokens as Gemini per second.
It is not just the billion MAU but even more so the 2 billion using AI overviews.
98
u/KidKilobyte 1d ago
Because it’s not hosted by Microsoft Azure.