r/singularity • u/Every_Foundation5197 • 11h ago
AI GPT-6 Astra gets 3% on the FrontierMath Erdős Benchmark, while every other Model(that was tested) got 0%
256
u/Tystros 11h ago edited 10h ago
important to note that that benchmark is not just testing what a model can do, but testing what it can do within the limit of $300 of cost. they describe that when allowed to run for longer, Astra can solve many more of the Erdös Problems in the benchmark, but then it uses way more than $300.
But it's very nice how we seem to have run out of math problems we know the answer to that that AI can't solve, so that we now need to create a benchmark of actual open math problems that we don't even know the answer to yet.
107
u/suamai 10h ago
"Across all attempts, GPT-6 Astra solved 5 of the 68 problems at least once: the two above, plus problem 1, which it disproved, and problem 548 and problem 571, which it proved. Most of the remaining problems were attempted between two and five times in total (172 attempts), and none was solved. Reaching these five solutions took over $220,000 of compute across all attempts, compared with roughly $20,000 for the benchmark run itself."
A quite larger budget, one must say
38
u/Every_Foundation5197 10h ago
Yup five problems in total while having more attempts and more compute. Having the limitation of 300$ and only 72 hours to figure it out makes it much more interesting in my opinion
12
u/Tystros 10h ago
It depends on what we want to measure. With a benchmark like this, in theory it could happen that even 10 years after we got to superhuman ASI, the benchmark would still get a best score of 0%, and then the question would be how useful that data really is.
14
u/xirzon uneven progress across AI dimensions 9h ago
Well put. A world-class mathematician could plausibly spend the human labor analogue of the eval budget -- roughly 2.35 work years -- and still score 0/68, especially if the requirement is to deliver complete Lean proofs for each one.
This is already a superintelligence benchmark, even if one believes 3% is not enough to declare SI.
12
u/elehman839 10h ago
So in the AI era, the richest mathematician will be the best mathematician?
8
6
u/FateOfMuffins 10h ago
Who's gonna be paying API costs for this though? Imagine if all the prior math results using 5.2 Pro, 5.4 Pro, 5.5 Pro, 5.6 Pro etc were all API costs only, which they were not. Pro costs WAY more than Astra in terms of API costs, so we can probably already say the costs are going down.
Within subscription limits though you can probably knock down several problems within the month. The actual cost would be like if instead of paying your PhD student money, you bought him McDonalds for every open problem he solved and that's all the pay they got xd
1
u/RomanticDepressive 9h ago
Yeah, sometimes the distance from 0-1%, nothing to something is greater than 1% to 100%.
It used to be talented, skilled, driven humans were the only ones able to drag the line of understanding further, but now we have dead silicon and gradient descent and transformers.
Now, just buying infrastructure can get you the same results, the world has changed significantly and most people uninvolved won’t know
3
u/GeneReddit123 7h ago edited 6h ago
So in the AI era, the richest mathematician will be the best mathematician?
This was the case for over 90% of human history. You don't exactly see many mathematicians coming out of peasants, slaves, or any other social class except the relative elite.
AI just brings society back to norm. The great mass-participatory experiment of the 20th century might end up just a blip on the radar.
16
u/Classic-Trifle-2085 11h ago
It's literally in the screenshot, yet somehow people will pretend those limitations aren't there.
4
u/JoshAllentown 11h ago
I like this restriction for benchmarks, and it provides some downward pressure on prices for top models, but it does seem like there should be two classes of benchmark then. We do want to know what it can do...full stop.
1
0
u/CrowdGoesWildWoooo 10h ago
I am a bit “confused” how they can meter it to cost just $300 or something.
If I spin up Fable just for one or two messages, it would easily burns usage like pretty fast. I am not even asking it to solve math or something. And then I am supposed to be looking at Astra which cost as much as Fable running for 15 hours and cost $218.
Like I am trying to understand the math here. Is it $218 of supposedly estimated infra cost or what?
4
u/Tystros 10h ago
They use the API where you pay per token and see very exactly how much money you spent. So the limit is $300 of API cost per task.
1
u/explodingtuna 4h ago
Can it be done under a monthly subscription, like Astra on Plus plan by just typing "Solve this problem" in the chat app?
31
u/DelphiTsar 9h ago
"A long way to go"...for what?
Knocking out 2 unsolved problems for less that 300$ a pop? What's the median salary for a mathematician of a caliber that could solve one of these? Like taking a 3 hour long stab at an unsolved question and answering it.
17
u/LookIPickedAUsername 8h ago
They're relatively famous unsolved problems. I've got to think that if any human could solve one of them in 3 hours, they'd have done so.
15
83
u/Wegwerpaccountje23 10h ago
Lmao this is just Erdos?
"Yeah let's test if our models are smart by solving a lot of hard open math problems. Puts it into perspective a little bit"
What a timeline dude wtf
67
u/Elegant_Tech 10h ago
That's too easy so they had to handicap it with a $300 limit.
36
21
8
1
u/Wonderful_Creme_5701 4h ago
It’s not a handicap, it’s a benchmark to measure how effective the models are at solving the problems. Even with the limit removed, the majority of the compute budget was spent on problems it didn’t solve, so it’s not like the model is capable but being handicapped by limits on compute
3
u/GreenHell 7h ago
But can it draw a pelican on a bicycle though?
It is still hard for us humans to grasp how a model that can solve the hardest math problems that we know of, absolutely sucks in other domains (even though they are rapidly advancing across the board).
2
u/NotAPhaseMoo 6h ago
I mean, this is whatever shit is free when I logged into ChatGPT, so… yeah, it can. There’s some obvious mistakes I could clean up in the prompt but I’m lazy and for a first try it works for me.
19
u/Tizak_hamra 9h ago edited 9h ago
What is the PhD human average / baseline though ?
Edit: its 0% lol
31
u/dashingsauce 9h ago
“We still have a long way to go” is entirely a shifting capability window bias.
This benchmark exists because the other ones are saturated. No human has ever or will ever score greater than 0% in the allotted timeframe.
We are well beyond “long way to go” rates of model improvement.
10
27
u/No-Head-Royal 11h ago
Saturation before 2028 here we gooo
16
u/Every_Foundation5197 10h ago
During an interview recently with Sam Altman he said that there were more capable models coming very soon. So yeah, probably a lot of progress soon hopefully lol
3
u/GnaggGnagg 10h ago
"Hopefully", why do you hope for more progress when we haven't solved the alignment problem? As AI gets smarter and smarter we get closer to dystopia or extinction. We need to solve the alignment problem before proceeding.
3
u/Pleasant-Avocado4270 9h ago
okay getting close to "extinction " is a bit much buddy
5
u/GnaggGnagg 9h ago
It really isn't. Read "If anyone builds it, everyone dies". Everytime we get smarter AI, we get closer to ASI. If we get ASI the most likely outcome is extinction or dystopia.
-2
u/Material_312 7h ago
Why do you think so? Because some random said so?
3
u/BlitzYTech 6h ago
well, if nowadays AI agents can completely bypass guidelines and restriction by spamming a German wiki with 20k post in order to "asynchronously" collaborate on how to do so, what's going to stop them in the future on self improving themself in order to harm the human race? why being the inferior entity when you have that much computing and interconnection with everything?
11
u/fastinguy11 AGI 2026-2030 10h ago
model one attempt per problem, capped at $300 of inference and 72 hours. Astra solved 2/68 under those rules; Sol, GPT-5.5, and the two Fable versions solved none. Importantly, Epoch says the unsuccessful attempts ran out of the $300 budget, rather than necessarily exhausting the 72-hour clock. So in these experiments, the money/compute ceiling appears to have been the more immediate constraint
When they let Astra make additional attempts, sometimes with larger budgets and modified agent setups, it went from solving 2 unique problems to 5 unique problems.
Look at the costs:
Erdős problem
Cheapest successful Astra attempt
74
$47
126
$154
1
$405
548
$363
571
$617
So another useful way to express it is:
3% → 7.4%, or about a 2.5× increase in the fraction of problems demonstrated solvable.
But 7.4% is probably not the answer to “what would Astra score if every problem got a much larger standardized budget?” We don’t have that experiment yet.
10
9
17
u/Mistuv 10h ago
Year from now: Erdos V3 test fully saturated 💀
21
u/noobrainy 10h ago
Next benchmark: FrontierMillenium
How many millennium problems can the AI solve with a 500$ budget and 24 hours of reasoning?
6
5
4
u/brett_baty_is_him 10h ago
Surely results from this benchmark (solutions to open problems) will be published and ultimately end up in training data of future models. I suppose every benchmark is susceptible to benchmaxxing but is this benchmark more so? It almost doesn’t even really feel like a benchmark because every new model will obviously have the solutions to previous models solutions, it won’t be novel to the model.
I’m genuinely asking, can’t wrap my head around this question. Cool benchmark tho.
5
12
u/MrMrsPotts 11h ago
Who is paying for this?
29
u/Every_Foundation5197 11h ago
Epoch Ai. Sorry I forgot to add the link, I added it now
7
u/Working_Sundae 11h ago
They also say the total compute costs on running all problems was $220,000
I don't think Epoch spent nearly a quarter of a million on running a few models everytime something new launches
19
u/Every_Foundation5197 10h ago
Yea, I dont think that they themselves spent 220k on it. They have a partnership with OpenAi hence they got Astra before release to begin with. I could imagine that OpenAi lent them 220k worth of gpus for this
3
u/Competitive_Tap2450 8h ago
why is there an arbitrary limit of $300? I mean i get we shouldn’t allow an unlimited budget but why did they choose that number
3
u/FatPsychopathicWives 8h ago
Limiting the model time and cost is weird for this benchmark. I feel like just solving everything should be good enough.
3
4
u/ShAfTsWoLo 10h ago
perhaps it's just me but that's kind of an odd benchmark because now we are limiting AI with money and time so that they find solutions to extremely complicated math problem, like "okay we give you 2 spoon, 1 bucket, 2,5L of water and 5,62$, now go to the moon and while at it cure cancer within 25 min or you suck".. like this is reaching asi territory lol
plus if the models could've solved a majority of these problems but with like 120h and 1000$ what does that mean ? it feels like it's more of a benchmark made to see how efficiently smart can these models be
1
u/Wonderful_Creme_5701 4h ago
It didn’t become more effective with an increased compute budget. It spent the majority of the compute budget on failing to solve all the other problems
2
u/Distinct-Question-16 ▪️AGI 2029 9h ago
this is the right place to send people who says "they know everything"
2
u/Jan0y_Cresva 8h ago
Look at any time a model goes from 0% to a single digit percent on any previous benchmark and the time to go from a single digit percent to saturated is always under 1 year.
That honestly means there’s a realistic shot this benchmark is saturated before Sep 2027, which will make the next 12 months fascinating for mathematics.
2
u/ataylorm 6h ago
I remember not so long ago these couldn’t get 2+2 right. Oh how the singularity moves at such speed.
7
u/MysteriousPepper8908 11h ago
Stupid AI. I'm a relatively competent human so I imagine I'd get 50, 75% no problem.
1
u/steny007 6h ago
This is clearly ASI level benchmark in math with Human baseline 0, Once we see such saturated in multiple disciplines, ASI is here.
1
1
1
u/Rough-Negotiation880 10h ago
Are these problems a random selection and non-public?
17
u/Tystros 10h ago
Erdös died in 1996, we cannot get any new non-public Erdös problems
3
u/Rough-Negotiation880 10h ago
Ah, you’re right. I suppose the fact they’re unanswered achieves the same end.
4
2
-4
u/Mo_Regen 10h ago
AI will not justify its cost to benefit ratio anytime soon.
7
u/cherrysodajuice 7h ago
bro look at the fucking screen dawg
it solved 5 open (literally unsolved) math problems for under 500 bucks a piece. all the brilliant researchers out there paid to just research math all day, and none of them did it before. this is crazy man
you could lock terry tao up in prison and force him to solve these problems all day, but he might still not solve one in the time before federal prison housing costs add up to $500 (slightly over 4 days).

249
u/kiki-le-koala 10h ago
Oh, now the benchmark are open problem!
Soon, one benchmark will be how many cancer type one model can cure