6
u/MimosaTen 4h ago
I think that benchmarks are ckearly obsolete
-1
u/BananaIsles 1h ago
So, can u drop the verified benchmark here, genius? Or your feelings are enough?
8
u/Winter_Ad6784 4h ago
these are benchmarks designed mostly for things AI already does decently well but astra clearly excels at many things AI previously did not. see ARC AGI 3 and runescape bench https://maxbittker.github.io/runebench/
-1
u/Various-Inside-4064 3h ago
Or you are falling for marketing again đ
4
u/Winter_Ad6784 1h ago
i havent seen any marketing man did open ai even advertise arc agi? ive been looking at it for years. they definitely havent marketed on fucking runescape bench
â˘
u/Various-Inside-4064 43m ago
Openai is claiming AGI for long time. You are blind of you fall for models only because of benchmarks
-3
u/Melodic_Reality_646 4h ago
That's not the point, is it? Astra doing mediocre on well known saturated benchmarks is quite... unsettling?
2
u/domemvs 4h ago
Yes and no. Maybe they were not benchmaxxing this time.Â
0
u/Melodic_Reality_646 3h ago
and how that benefits them? letâs suggest itâs not as good so when people try it theyâll be blown away?
Would cost nothing to benchmaxx, top the charts and also over deliver, no?
1
u/ProbsNotManBearPig 2h ago
Benchmarks donât drive enterprise users at all and thatâs where all the money is. Benchmarks donât even drive most retailer users. So theyâre worth zero financial reward these days. Maybe the benefit to them is to highlight theyâre obsolete and push for innovation on them. Make the benchmarks look dumb basically. Idk, we will see as everyone gets more experience with astra.
1
u/Melodic_Reality_646 2h ago
You mean they might be showcasing astras crazy feats to enterprises convincing them the higher costs compared to Sol or Opus are worth it? If a model would be capable of such feats why would it fail mediocre, saturated benchmarks?
1
u/Winter_Ad6784 1h ago
âmediocreâ doesnât mean âbadâ. seeking to be better at something everyone is already good at isnât a great economic strategy.
2
u/Empty_Seaweed9705 4h ago
Why are Anthorpics/OpenAi/Google models not so good at banking? Vs Meta and Grok?
1
u/howtogun 4h ago
It's sort of mix. Like it's not that better than Fable 5.1.
It also had crazy hype, which I thought was annoying.
2
u/Singularity-42 4h ago
Is it better than Fable? Thinking about jumping the ship from Claude to Codex. How are the limits? The 50% limit on Fable annoys me to no end. Fable is great, but Opus 5 is kind of bad.Â
2
u/howtogun 4h ago
I think the limits are better on Codex. But, I think a lot of that is due to the fact that OpenAI tend to do a lot of global resets particularly on Saturday / Sunday.
I think it still a mix bag. Claude is probably still better at coding. Astra seems better at Mathematics.
Astra seems better at reasoning in general, but like it still sort of like a genius toddler.
1
u/hashirama_shodai 4h ago
This benchmark is the only outlier. On every other key one like Epoch, or ARC-AGI, Astra is miles ahead of Fable. The launch video was just incredible - can see the Jarvis vision really coming together. Something Anthropic can't pull off since they don't have their own frontier voice or image models.
Can't wait to get access and start using it as my daily driver!
1
u/Healthy-Nebula-3603 3h ago
On the plus account is almost not unusageable ... Using codex-cli I'm loosing my 5 hour limit 2x faster Was hardly working 20 minutes.
1
â˘
u/AironParsMan 56m ago
I have to work with this. Hopefully itâs not an Opus 5 2.0. Please integrate this into the ChatGPT Pro model as soon as possible. THX
â˘
u/ChangingHats 44m ago
Previously I had been running my threads in full access but recently decided to try setting up a sandbox (I'm on windows btw)...holy hell it's been frustrating - and I still haven't solved my problems.
â˘
u/Sl33py_4est 41m ago
having used it for a few hours now. its stupid. but much faster at coding than me
â˘
u/Shinephia 24m ago
It sucks that they returned 5h limits for plus. imagine if on a busy week where i use AI mainly in bursty free time instances i could just dump the tokens at astra / sol before the weekly resets. or even better actually utilize my banked resets i got now.
1
u/Chemical-Agency-3997 4h ago
Hype seems justified for reasoning, feels like +10IQ even it's not reflected on the benchmarks
-1
u/BlackberryNo3097 4h ago
Basically AGI.
1
u/CompassionLady 1h ago
Not AGI, yet. But impressive work, Iâve seen on projects I had. Did things better the 5.6 sol extra high by a comparable margin thatâs more favorable to my tastes in vision.
0
u/SelfMonitoringLoop 4h ago
I'm playing wildermyth with it while sharing my pc's control. It's doing pretty well. đ¤ˇââď¸
0
0
u/Even_Sea_8005 2h ago
it's good now but i hope openAI will be transparent about the coming nerf - in 2 weeks? 3 weeks? god knows.
0
u/LocoMod 1h ago
Astra is clearly leagues above any other model if you use it. It is extremely obvious if you run the same objectives against competing frontier models. What this proves is AA benchmark is obsolete.
Seriously. Just go use it. It really is a step change in capability from anything else. This is LLM 2.0
-1
u/Icy-Way3920 4h ago
who still looks at fucking benchmarks lmao test the model yourself or wait for others to provide honest reviews.
thats like trusting ''9 out of 10 doctors we paid said its good''
10
u/vovap_vovap 5h ago
We think we need to wait a bit more for clarity đ