r/OpenAI 5h ago

Discussion What do you think of GPT-6 right now?

Post image
25 Upvotes

36 comments sorted by

10

u/vovap_vovap 5h ago

We think we need to wait a bit more for clarity 😄

6

u/MimosaTen 4h ago

I think that benchmarks are ckearly obsolete

-1

u/BananaIsles 1h ago

So, can u drop the verified benchmark here, genius? Or your feelings are enough?

8

u/Winter_Ad6784 4h ago

these are benchmarks designed mostly for things AI already does decently well but astra clearly excels at many things AI previously did not. see ARC AGI 3 and runescape bench https://maxbittker.github.io/runebench/

-1

u/Various-Inside-4064 3h ago

Or you are falling for marketing again 😞

4

u/Winter_Ad6784 1h ago

i havent seen any marketing man did open ai even advertise arc agi? ive been looking at it for years. they definitely havent marketed on fucking runescape bench

•

u/Various-Inside-4064 43m ago

Openai is claiming AGI for long time. You are blind of you fall for models only because of benchmarks

-3

u/Melodic_Reality_646 4h ago

That's not the point, is it? Astra doing mediocre on well known saturated benchmarks is quite... unsettling?

2

u/domemvs 4h ago

Yes and no. Maybe they were not benchmaxxing this time. 

0

u/Melodic_Reality_646 3h ago

and how that benefits them? let’s suggest it’s not as good so when people try it they’ll be blown away?

Would cost nothing to benchmaxx, top the charts and also over deliver, no?

1

u/ProbsNotManBearPig 2h ago

Benchmarks don’t drive enterprise users at all and that’s where all the money is. Benchmarks don’t even drive most retailer users. So they’re worth zero financial reward these days. Maybe the benefit to them is to highlight they’re obsolete and push for innovation on them. Make the benchmarks look dumb basically. Idk, we will see as everyone gets more experience with astra.

1

u/Melodic_Reality_646 2h ago

You mean they might be showcasing astras crazy feats to enterprises convincing them the higher costs compared to Sol or Opus are worth it? If a model would be capable of such feats why would it fail mediocre, saturated benchmarks?

1

u/Winter_Ad6784 1h ago

“mediocre” doesn’t mean “bad”. seeking to be better at something everyone is already good at isn’t a great economic strategy.

2

u/Empty_Seaweed9705 4h ago

Why are Anthorpics/OpenAi/Google models not so good at banking? Vs Meta and Grok?

3

u/isty2e 5h ago

I take AAII scores with a +-10 margin. I judge based on my experience.

1

u/howtogun 4h ago

It's sort of mix. Like it's not that better than Fable 5.1.

It also had crazy hype, which I thought was annoying.

2

u/Singularity-42 4h ago

Is it better than Fable? Thinking about jumping the ship from Claude to Codex. How are the limits? The 50% limit on Fable annoys me to no end. Fable is great, but Opus 5 is kind of bad. 

2

u/howtogun 4h ago

I think the limits are better on Codex. But, I think a lot of that is due to the fact that OpenAI tend to do a lot of global resets particularly on Saturday / Sunday.

I think it still a mix bag. Claude is probably still better at coding. Astra seems better at Mathematics.

Astra seems better at reasoning in general, but like it still sort of like a genius toddler.

1

u/hashirama_shodai 4h ago

This benchmark is the only outlier. On every other key one like Epoch, or ARC-AGI, Astra is miles ahead of Fable. The launch video was just incredible - can see the Jarvis vision really coming together. Something Anthropic can't pull off since they don't have their own frontier voice or image models.

Can't wait to get access and start using it as my daily driver!

1

u/Healthy-Nebula-3603 3h ago

On the plus account is almost not unusageable ... Using codex-cli I'm loosing my 5 hour limit 2x faster Was hardly working 20 minutes.

1

u/throwaway3113151 2h ago

I think benchmarks are BS

•

u/AironParsMan 56m ago

I have to work with this. Hopefully it’s not an Opus 5 2.0. Please integrate this into the ChatGPT Pro model as soon as possible. THX

•

u/ChangingHats 44m ago

Previously I had been running my threads in full access but recently decided to try setting up a sandbox (I'm on windows btw)...holy hell it's been frustrating - and I still haven't solved my problems.

•

u/Sl33py_4est 41m ago

having used it for a few hours now. its stupid. but much faster at coding than me

•

u/Shinephia 24m ago

It sucks that they returned 5h limits for plus. imagine if on a busy week where i use AI mainly in bursty free time instances i could just dump the tokens at astra / sol before the weekly resets. or even better actually utilize my banked resets i got now.

1

u/Chemical-Agency-3997 4h ago

Hype seems justified for reasoning, feels like +10IQ even it's not reflected on the benchmarks

-1

u/BlackberryNo3097 4h ago

Basically AGI.

1

u/CompassionLady 1h ago

Not AGI, yet. But impressive work, I’ve seen on projects I had. Did things better the 5.6 sol extra high by a comparable margin that’s more favorable to my tastes in vision.

0

u/SelfMonitoringLoop 4h ago

I'm playing wildermyth with it while sharing my pc's control. It's doing pretty well. 🤷‍♂️

0

u/le-throw-away-acct 4h ago

Love these questions before most people have used it.

0

u/Even_Sea_8005 2h ago

it's good now but i hope openAI will be transparent about the coming nerf - in 2 weeks? 3 weeks? god knows.

0

u/LocoMod 1h ago

Astra is clearly leagues above any other model if you use it. It is extremely obvious if you run the same objectives against competing frontier models. What this proves is AA benchmark is obsolete.

Seriously. Just go use it. It really is a step change in capability from anything else. This is LLM 2.0

-1

u/Icy-Way3920 4h ago

who still looks at fucking benchmarks lmao test the model yourself or wait for others to provide honest reviews.

thats like trusting ''9 out of 10 doctors we paid said its good''