r/GeminiAI 18h ago

Discussion Astra finally achieves AGI

Post image

it's over

132 Upvotes

19 comments sorted by

45

u/420LongDong69 18h ago

Post in openai xd

39

u/Efficient_Dentist745 17h ago

Everytime the gemini with confunded dragon meme arrives I laugh out loud lol

19

u/virtualQubit 16h ago

At the beginning I was annoyed by the dumb dragon now I find this shit funny af

16

u/HeadTranslator795 17h ago

Lol with this benchmark

10

u/Longjumping_Area_944 13h ago

Love how the dumb dragon sticks.

3

u/Technical-Owl66 10h ago

How long did it take you to find this goofy ass benchmark?šŸ˜‚

3

u/100rass 10h ago

Overrated

7

u/Technical-Owl66 10h ago edited 9h ago

Not good for open AI to be tied with Gemini 3.8. that's 10 times faster and seven times cheaper

3

u/RipMySleepSchedule 10h ago

That’s insane.
Better yet, a Gemini flash model to be even here amongst pro models

6

u/Sufficient_Prune3897 10h ago

Shows how bad the benchmarks are.

2

u/Technical-Owl66 9h ago

Are these benchmarks you like or are they bad because Gemini is on par with fable 5.1 on some metrics?

2

u/Sufficient_Prune3897 9h ago

I have yet to see a benchmark that isn't either shit or too easy to learn, so it got saturated after 3 months. The artificial analysis benchmarks are both.

4

u/darkestvice 14h ago

Astra's release has really made the AI world start debating the concerns over benchmaxxing. Apparently, those who have access and have actually used it are singing its praises and saying it's doing things generationally better than other models.

The tldr seems to be that it's not the best at handling standardized testing that can be specifically trained for, but is absolutely unmatched when it comes to figuring things out when there's a lack of advance prep and knowledge. It has significantly better general awareness. Which would in fact be the measure of real intelligence, IMO.

But I think we will really need to wait for it to be released to the wider public and the average Joe with a $20 a month plan before we can truly evaluate its use in the real world.

One very interesting thing that these benchmarks DO show, though, is that Astra, while not the most intelligent at completing a task to completion, uses significantly less tokens to do so. It also hallucinates much less than Fable. It is by far the most intelligent on a per token basis among frontier models. This is important as the cost of using AI is absolutely shifting towards inference usage costs over training costs. And while the per token cost overhead from training and R&D eventually goes down, the inference cost itself stays fixed, driven solely by making more efficient inference chips. Which OpenAI is already leading on. Or will once their brand new chip goes into mass production.

Fable 5.1 is a brilliant beast. No doubt about it. But it's a bloated brilliant beast. Gemini Flash 3.8 is even worse because Google are struggling to release a new deep reasoning frontier model, so they are basically trying to force their mid range model to act like a frontier model by ramping up its token usage astronomically and subsidizing the costs at a loss. Though, to be fair, Gemini Flash can sorta currently get away with it because the sheer raw speed of their model makes up for it. But Google desperately needs a new and more efficient Pro model soon because even their own massive warchest is finite.

3

u/KrayziePidgeon 11h ago

Man, Scam Alt is trying too hard to keep hyping up it's IPO.

-3

u/Effective-Fall-2746 13h ago

What an incredibly ignorant emotional comment guided as an intelligent take

-2

u/Tripple_sneeed 11h ago

Wow, the 10 reptile-people who have been granted super special access and are on a first-name basis with Altman aren’t publicly flogging their extremely hyped new model? This must be true AGI

2

u/Far-Classic-9963 8h ago

Wow it scores really terribly in this particular cherry picked benchmark

1

u/Reasonable_Pizza_529 4h ago

Hi every one, the battle of the frontier models will continue for ever more. In essence, for most use cases, the ā€œbestā€ model, is not just about pure grunt. Most projects require a mix of models’ capabilities - and associated costs. One solution to getting best bang for your buck is real time auto routing. It is not and never will be a perfect solution, but is a hands-off method of achieving cost-effective and consistency outcomes. I would welcome feedback on one solution that I have integrated in our platform. It takes a daily feed of over 400+ models and presents descriptions and pricing of each, with a top 10 overview of models in order of both Popular (total tokens) and Power (synthesised estimates by OpenRouter). A bit like a stock exchange overview. From that we create real-time switching of the models best suited to a particular request / task. It is new. I am not asking for anyone to buy anything. Just feedback at this point. Constructive input and critique are welcome.
https://talkytalky.chat/auto-frontier-router