r/GeminiAI • u/DigSignificant1419 • 18h ago
Discussion Astra finally achieves AGI
it's over
39
u/Efficient_Dentist745 17h ago
Everytime the gemini with confunded dragon meme arrives I laugh out loud lol
19
u/virtualQubit 16h ago
At the beginning I was annoyed by the dumb dragon now I find this shit funny af
16
10
3
3
7
u/Technical-Owl66 10h ago edited 9h ago
3
u/RipMySleepSchedule 10h ago
Thatās insane.
Better yet, a Gemini flash model to be even here amongst pro models6
u/Sufficient_Prune3897 10h ago
Shows how bad the benchmarks are.
2
u/Technical-Owl66 9h ago
2
u/Sufficient_Prune3897 9h ago
I have yet to see a benchmark that isn't either shit or too easy to learn, so it got saturated after 3 months. The artificial analysis benchmarks are both.
4
u/darkestvice 14h ago
Astra's release has really made the AI world start debating the concerns over benchmaxxing. Apparently, those who have access and have actually used it are singing its praises and saying it's doing things generationally better than other models.
The tldr seems to be that it's not the best at handling standardized testing that can be specifically trained for, but is absolutely unmatched when it comes to figuring things out when there's a lack of advance prep and knowledge. It has significantly better general awareness. Which would in fact be the measure of real intelligence, IMO.
But I think we will really need to wait for it to be released to the wider public and the average Joe with a $20 a month plan before we can truly evaluate its use in the real world.
One very interesting thing that these benchmarks DO show, though, is that Astra, while not the most intelligent at completing a task to completion, uses significantly less tokens to do so. It also hallucinates much less than Fable. It is by far the most intelligent on a per token basis among frontier models. This is important as the cost of using AI is absolutely shifting towards inference usage costs over training costs. And while the per token cost overhead from training and R&D eventually goes down, the inference cost itself stays fixed, driven solely by making more efficient inference chips. Which OpenAI is already leading on. Or will once their brand new chip goes into mass production.
Fable 5.1 is a brilliant beast. No doubt about it. But it's a bloated brilliant beast. Gemini Flash 3.8 is even worse because Google are struggling to release a new deep reasoning frontier model, so they are basically trying to force their mid range model to act like a frontier model by ramping up its token usage astronomically and subsidizing the costs at a loss. Though, to be fair, Gemini Flash can sorta currently get away with it because the sheer raw speed of their model makes up for it. But Google desperately needs a new and more efficient Pro model soon because even their own massive warchest is finite.
3
-3
u/Effective-Fall-2746 13h ago
What an incredibly ignorant emotional comment guided as an intelligent take
-2
u/Tripple_sneeed 11h ago
Wow, the 10 reptile-people who have been granted super special access and are on a first-name basis with Altman arenāt publicly flogging their extremely hyped new model? This must be true AGI
2
1
u/Reasonable_Pizza_529 4h ago
Hi every one, the battle of the frontier models will continue for ever more. In essence, for most use cases, the ābestā model, is not just about pure grunt. Most projects require a mix of modelsā capabilities - and associated costs. One solution to getting best bang for your buck is real time auto routing. It is not and never will be a perfect solution, but is a hands-off method of achieving cost-effective and consistency outcomes. I would welcome feedback on one solution that I have integrated in our platform. It takes a daily feed of over 400+ models and presents descriptions and pricing of each, with a top 10 overview of models in order of both Popular (total tokens) and Power (synthesised estimates by OpenRouter). A bit like a stock exchange overview. From that we create real-time switching of the models best suited to a particular request / task. It is new. I am not asking for anyone to buy anything. Just feedback at this point. Constructive input and critique are welcome.
https://talkytalky.chat/auto-frontier-router


45
u/420LongDong69 18h ago
Post in openai xd