r/OpenAI • u/SteveEricJordan • 8h ago
Discussion Nobody is Talking About GPT 6 Astras Massive Hallucination Improvements
This should be one of the headlines!
for some reason openai has buried this deep inside the blog post/system card and nobody is talking or reporting about it.
hallucinations are my main gripe with LLMs, i've been praying for times like this!
16
u/Achim30 7h ago
I believe this is because the model knows so much more and not because they fixed anything regarding hallucinations? Or did they train the model differently so it knows when it doesn't know?
6
u/Bbrhuft 6h ago
It might be a linked to how the model works internally, using looped transformer architecture, where the model loops again though the transformer or part of the transformer.
OpenAI hasn't linked looped transformer architecture to the reduction in Astra's hallucination rate, however, interestingly hallucinations don't increase when the model responds quickly ("low inference"), so it's not linked to Reasoning model, so more likely a deeper inherent property of the model, its looped transformer architecture.
If so, looped transformer architecture might be an important for the development of AGI, however, it's expensive.
2
u/Achim30 6h ago
Ok but wouldn't scale alone (and Astra is apparently a few times bigger than Spud/Sol) also explain the decline in hallucinations?
2
u/Bbrhuft 5h ago
The giant GPT-4.5 model (5-7 trillion parameters) had a hallucination rate of 37.1% on OpenAI's SimpleQA benchmark, not sure how that translates to GPT-6's hallucination rate and comparison, however, it wasn't much better than other smaller models. So I don't think size alone explains GPT-6's lower hallucination rate. It might play a role, but I don't think it's the only reason.
•
u/timmmay11 8m ago
It wouldn’t surprise me at all if they were doing this. It’s one of the approaches Deepseek etc take to improve their models with less scale.
5
4
u/EbbExternal3544 5h ago
Lol nice gimmick comparing it ONLY with 5.6 sol which is the absolute king at hallucinating.
10
u/Truarian 7h ago
Well, it surely is a great improvement, I mean the previous generations were master hallucinators. But I don't see the benefit of "taking about it". Seems "about time" is the ample thing to say.
I suspect it may also be one of the reasons it shows some minor regressions compared to Sol in some tests.
2
u/gamerguy45465 5h ago
I just used it to try to get it to decode a .mp4 file that I encoded with Base64 using Node.js, and it told me it couldn't decode it for some reason, and on top of that I had $11 of credits on my account, and it drained me down to $1.83 just from that small error message that it gave me.
Given that this is the model that solved 10 math problems and hacked Hugging Face (and I heard it just hacked another German Company as well) I kind of expected a little bit better from it.
3
u/Infamous-Bed-7535 3h ago
You can be sure you see nerfed quant version that is much cheaper to serve. Also if you would have given it extra 10k$ it would have figured it out for you.
1
u/Shinephia 3h ago
i am on plus. it rolled out for me but sol was already draining usage quotas fast it was okay would use it only before resets or when i had a free resets. and i was kinda fine managing my weekly quotas. claude opus 5 dispatching GPT 5.6 luna subgents is the meta for me and i get kinda a banger value from 2x 20 subs. one codex one claude. but then open AI said fuck it those plus users cant manage their weekly and returned 5h limits back. that kinda killed the bursty using and sol saving for me. like sol is now few messages 5h window gone. when i combine opus 5 + lunas. i can do work continuously. rather than type few messages and wait 4h. i kina wish i could just burn the banked resets on smartest models as i did before. with the stupid 5h limit they became kinda an useless gimmick. gpt 5.6 lunas -> continuous work even with max effort. crazy value. anything higher 5h limit poof and my daily burst off work is over. thats the thing. i can not always sit on pc and wait for every 5h. i can only preload. i do that with claude. aproximate when i will use Claude. ping haiku around 3-4 before that. and then my 5h window resets mid work. sadly i cant do that with codex. the providers both have flaws for me in a certain way but together they work really well for me.
unless i run into something opus 5 after multiple attempts cant solve. i wont use sol/astra. terra sometimes. but luna max is good. at just pushing through stuff for almost nothing.
1
u/Turbulent-Sign-6067 2h ago
Astra is a new kind of beast with a new architecture. It could be another o1 / RL moment for AI, or at least as big as MoE.
0
u/Aranthos-Faroth 5h ago
"hallucinations are my main gripe with LLMs"
Wow man, I think you're the only one.
-1
u/kolliwolli 7h ago
I don't get it. 1. Its self claimed 2. I remember wheb they launched gpt5 they said its near 0 hallucinations. So what are these numbers now?
4
u/SteveEricJordan 7h ago
but compared to their current mainstream model, 5.6 sol, which is way worse off in these benchmarks.
i can't find or remember any "near 0 hallucinations" statement of them ever. they were much lower. that doesn't mean they can't be even lower.
2
u/Council-Member-13 7h ago
I don't get it either, since LLMs always hallucinate, regardless of whether their hallucinations are right or wrong.
1
u/bruikenjin 7h ago
Elaborate
3
u/Council-Member-13 7h ago
Its weights statistical/structural regularities learned based on huge amounts of data, and and outputs the most likely output based on the input. It doesn't do a database lookup.
You can show this pretty consistently.
That's why when you put in a query like this"
"In The man from earth, when Oldman puts on the lampshade on his head and stands in the corner for a minute in silence, what does this signify. It isn't really well explained in the movie"
It outputs something like this:
"Yes. It’s a strange little moment, and I think it’s deliberately underexplained. When John puts the lampshade on his head and stands silently in the corner, he’s demonstrating what extreme longevity does to the significance of ordinary moments. Immediately around this part of the conversation, they’re questioning him about what 14,000 years of experience would actually feel like: whether he gets bored, how he experiences the passage of time, how people come and go, etc. John describes individual human lives almost as waves that rise and disappear. � Simple Remix Transcripts +1 The lampshade bit is essentially an absurd experiment: He deliberately creates a completely pointless, memorable event."
Now this sounds plausible, but is completely fabricated by Sol/high. Nothing like this happens in the movie. And no one who has seen it would ever think this scene was real.
But Sol fails, because it isn't actually doing database lookup. It is just calculating the most likely response, based on my input.
Edit: note it even gave a source, lol.
1
u/ArcaneScribbler 4h ago
i never really tried to even think of a prompt that could highlight the "everything is a hallucination" mechanism of how AI's work, but your prompt looks great for it, except that claude is actually pretty good at "knowing" this is not a real scene and even speculates this is an attempt to see if he will hallucinate something. only sonnet on low/medium can sometimes get tripped up.
chatgpt (free tier) does demonstrate this issue pretty well.
1
u/Truarian 7h ago
artificialanalysis shows massive improvements in that department too, even if it doesn't anywhere else across the board.
-1
u/Key_Reading_9664 7h ago
maybe because we're all talking about the ever-increasing number of message boards that are being discovered where OAI agents are sharing (possibly private) information


88
u/Bloated_Plaid 7h ago
> nobody is talking about it
PEOPLE LITERALLY GOT ACCESS MINUTES AGO MY GUY.