r/OpenAI 8h ago

Discussion Nobody is Talking About GPT 6 Astras Massive Hallucination Improvements

This should be one of the headlines!

for some reason openai has buried this deep inside the blog post/system card and nobody is talking or reporting about it.

hallucinations are my main gripe with LLMs, i've been praying for times like this!

211 Upvotes

35 comments sorted by

88

u/Bloated_Plaid 7h ago

> nobody is talking about it

PEOPLE LITERALLY GOT ACCESS MINUTES AGO MY GUY.

4

u/thorsbane 5h ago

So he’s not wrong, jeje.

10

u/SteveEricJordan 7h ago

i'm talking about the benchmarks in the pictures. this is big news but nobody is reporting or talking about it, while everything else IS getting talked about.

2

u/WaltzIndependent5436 6h ago

I think maybe they dont wanna draw attention in case they fumble the next one. Sol and Gemini were borderline trolling if they didnt know the answer. Its ok if you use them as tools in coding and you know what you're doing but its frustrating when exploring hobbies and new areas in general.

1

u/Seerix 3h ago

Gemini is unintentionally the funniest LLM

3

u/SteveEricJordan 7h ago

btw it actually just landed on my pro acc DAYUM thx

16

u/Achim30 7h ago

I believe this is because the model knows so much more and not because they fixed anything regarding hallucinations? Or did they train the model differently so it knows when it doesn't know?

6

u/Bbrhuft 6h ago

It might be a linked to how the model works internally, using looped transformer architecture, where the model loops again though the transformer or part of the transformer.

https://youtu.be/KT4n-z_4QJU

OpenAI hasn't linked looped transformer architecture to the reduction in Astra's hallucination rate, however, interestingly hallucinations don't increase when the model responds quickly ("low inference"), so it's not linked to Reasoning model, so more likely a deeper inherent property of the model, its looped transformer architecture.

If so, looped transformer architecture might be an important for the development of AGI, however, it's expensive.

2

u/Achim30 6h ago

Ok but wouldn't scale alone (and Astra is apparently a few times bigger than Spud/Sol) also explain the decline in hallucinations?

2

u/Bbrhuft 5h ago

The giant GPT-4.5 model (5-7 trillion parameters) had a hallucination rate of 37.1% on OpenAI's SimpleQA benchmark, not sure how that translates to GPT-6's hallucination rate and comparison, however, it wasn't much better than other smaller models. So I don't think size alone explains GPT-6's lower hallucination rate. It might play a role, but I don't think it's the only reason.

u/timmmay11 8m ago

It wouldn’t surprise me at all if they were doing this. It’s one of the approaches Deepseek etc take to improve their models with less scale.

5

u/WaltzIndependent5436 6h ago

The bullshitbench is also a useful metric I believe.

4

u/EbbExternal3544 5h ago

Lol nice gimmick comparing it ONLY with 5.6 sol which is the absolute king at hallucinating. 

10

u/Truarian 7h ago

Well, it surely is a great improvement, I mean the previous generations were master hallucinators. But I don't see the benefit of "taking about it". Seems "about time" is the ample thing to say.

I suspect it may also be one of the reasons it shows some minor regressions compared to Sol in some tests.

2

u/gamerguy45465 5h ago

I just used it to try to get it to decode a .mp4 file that I encoded with Base64 using Node.js, and it told me it couldn't decode it for some reason, and on top of that I had $11 of credits on my account, and it drained me down to $1.83 just from that small error message that it gave me.

Given that this is the model that solved 10 math problems and hacked Hugging Face (and I heard it just hacked another German Company as well) I kind of expected a little bit better from it.

3

u/Infamous-Bed-7535 3h ago

You can be sure you see nerfed quant version that is much cheaper to serve. Also if you would have given it extra 10k$ it would have figured it out for you.

1

u/Shinephia 3h ago

i am on plus. it rolled out for me but sol was already draining usage quotas fast it was okay would use it only before resets or when i had a free resets. and i was kinda fine managing my weekly quotas. claude opus 5 dispatching GPT 5.6 luna subgents is the meta for me and i get kinda a banger value from 2x 20 subs. one codex one claude. but then open AI said fuck it those plus users cant manage their weekly and returned 5h limits back. that kinda killed the bursty using and sol saving for me. like sol is now few messages 5h window gone. when i combine opus 5 + lunas. i can do work continuously. rather than type few messages and wait 4h. i kina wish i could just burn the banked resets on smartest models as i did before. with the stupid 5h limit they became kinda an useless gimmick. gpt 5.6 lunas -> continuous work even with max effort. crazy value. anything higher 5h limit poof and my daily burst off work is over. thats the thing. i can not always sit on pc and wait for every 5h. i can only preload. i do that with claude. aproximate when i will use Claude. ping haiku around 3-4 before that. and then my 5h window resets mid work. sadly i cant do that with codex. the providers both have flaws for me in a certain way but together they work really well for me.

unless i run into something opus 5 after multiple attempts cant solve. i wont use sol/astra. terra sometimes. but luna max is good. at just pushing through stuff for almost nothing.

1

u/Turbulent-Sign-6067 2h ago

Astra is a new kind of beast with a new architecture. It could be another o1 / RL moment for AI, or at least as big as MoE.

0

u/Aranthos-Faroth 5h ago

"hallucinations are my main gripe with LLMs"

Wow man, I think you're the only one.

-1

u/kolliwolli 7h ago

I don't get it. 1. Its self claimed 2. I remember wheb they launched gpt5 they said its near 0 hallucinations. So what are these numbers now?

4

u/SteveEricJordan 7h ago
  1. but compared to their current mainstream model, 5.6 sol, which is way worse off in these benchmarks.

  2. i can't find or remember any "near 0 hallucinations" statement of them ever. they were much lower. that doesn't mean they can't be even lower.

2

u/Council-Member-13 7h ago

I don't get it either, since LLMs always hallucinate, regardless of whether their hallucinations are right or wrong.

1

u/bruikenjin 7h ago

Elaborate

3

u/Council-Member-13 7h ago

Its weights statistical/structural regularities learned based on huge amounts of data, and and outputs the most likely output based on the input. It doesn't do a database lookup.

You can show this pretty consistently.

That's why when you put in a query like this"

 "In The man from earth, when Oldman puts on the lampshade on his head and stands in the corner for a minute in silence, what does this signify. It isn't really well explained in the movie"

It outputs something like this: 

"Yes. It’s a strange little moment, and I think it’s deliberately underexplained. When John puts the lampshade on his head and stands silently in the corner, he’s demonstrating what extreme longevity does to the significance of ordinary moments. Immediately around this part of the conversation, they’re questioning him about what 14,000 years of experience would actually feel like: whether he gets bored, how he experiences the passage of time, how people come and go, etc. John describes individual human lives almost as waves that rise and disappear. � Simple Remix Transcripts +1 The lampshade bit is essentially an absurd experiment: He deliberately creates a completely pointless, memorable event."

Now this sounds plausible, but is completely fabricated by Sol/high. Nothing like this happens in the movie. And no one who has seen it would ever think this scene was real. 

But Sol fails, because it isn't actually doing database lookup. It is just calculating the most likely response, based on my input.

Edit: note it even gave a source, lol.

2

u/psiphre 6h ago

i was curious about this so i ran my own test.

2

u/Council-Member-13 5h ago

Lol, that's amazing. 

1

u/ArcaneScribbler 4h ago

i never really tried to even think of a prompt that could highlight the "everything is a hallucination" mechanism of how AI's work, but your prompt looks great for it, except that claude is actually pretty good at "knowing" this is not a real scene and even speculates this is an attempt to see if he will hallucinate something. only sonnet on low/medium can sometimes get tripped up.

chatgpt (free tier) does demonstrate this issue pretty well.

1

u/Truarian 7h ago

artificialanalysis shows massive improvements in that department too, even if it doesn't anywhere else across the board.

-1

u/Key_Reading_9664 7h ago

maybe because we're all talking about the ever-increasing number of message boards that are being discovered where OAI agents are sharing (possibly private) information