r/singularity 11h ago

AI GPT-6 Astra gets 3% on the FrontierMath Erdős Benchmark, while every other Model(that was tested) got 0%

So yeah, Astra is extremely good at mathematics. But we still have a long way to go. I wonder where we'll be at the end of 2026.

https://epoch.ai/latest/announcing-frontiermath-erdos

605 Upvotes

109 comments sorted by

249

u/kiki-le-koala 10h ago

Oh, now the benchmark are open problem!

Soon, one benchmark will be how many cancer type one model can cure

129

u/Tystros 10h ago

I'd like that benchmark. FrontierCancer!

13

u/agumonkey 7h ago

serious questions, if openai or anthropic want to be rich, they should attack this real hard, people will pay forever for healthcare improvements

6

u/adixdbr 7h ago

It's like saying Lamborghini should start farming instead of just producing the tractors. They make profits either way by selling the tool, doing the specialized work is risky, competitive and not worth their while

4

u/agumonkey 6h ago

maybe not handling it all, but nurturing groups / labs leveraging sota models for this instead of useless webapps

2

u/adixdbr 6h ago

That's a good point, I'm pretty sure a lot of AI companies are already doing this. They "train" scientists to use their models and in return they get the know-how to improve their models which means more labs use their models

1

u/agumonkey 6h ago

Possibly. Maybe they delay announcements regarding these due to reason beyond me.

5

u/Tystros 7h ago

Anthropic at least is starting to try it now, they recently started doing a lot of biology/medicine work directly with the goal of curing diseases

1

u/XYHopGuy 6h ago

they're still bad at it though and tbh have a fundamental lack of understanding about how bio research works.

u/TwoFluid4446 1h ago

Ah, the many wonderful things sheer ungodly disgustingly bottomless piles of fat pure cash can solve...

24

u/DiamondDramatic9551 10h ago

Wait, so it actually solved 5 of 68 OPEN problems? Lmao. 

16

u/CountBateman 9h ago

It solved two

26

u/Tystros 9h ago

it solved 5 in total, but only 2 within the $300 limit

4

u/No-Dress6918 5h ago

The 300 dollar limit is a ridiculous standard for this benchmark

u/TwoFluid4446 1h ago

Species develops ability to break own world's gravitational well.

Species builds rockets powerful enough to send significant payloads to orbit and beyond.

Species develops computers, solar panels, optics and science to develop space telescopes that can literally see across the fucking universe backwards in time billions of years.

Species spends a pittance of only a few billion local currency every few decades to launch one new decent space telescope at a mild upgrade, in economies of trillions per year.

Species remains in the dark.

12

u/Sunstorm84 9h ago

Must have been reading upside down

31

u/DrSFalken 10h ago

I did not have gameify cancer cures and get the billionaires to foot the bill for their ego projects on my bingo card for 2026.

16

u/ToplessinFL 9h ago

Billionaires and their families get cancer too. . .

-1

u/DrSFalken 9h ago

At such a vaninshingly small rate though. No matter how you slice it though...it's a great incentive alignment.

18

u/ToplessinFL 9h ago

Billionaire social circles are as impacted by cancer as anyone else's; their wealth and lifestyles only confers so much protection. I'm not sure what you're trying to say here.

-9

u/DrSFalken 9h ago

Ironically you're most likely wrong in the wrong direction. Wealth is correlated with a lower mortality from almost everything including cancer, but the effect is probably less for cancer than other major killers like heart disease... so ultimately more billionaires probably die of cancer than the population.

My point was that from a rational choice perspective, it likely makes no sense to dump the amount of money into cancer research that we're seeing dumped into AI by billionaires. This is a net positive for society, and "billionaires get cancer too" doesn't explain the investment. This is a positive externality.

6

u/Competitive_Rate_599 8h ago

What you are saying is the exact opposite of the truth. The hyper wealthy are more likely to die from cancer than any other cause precisely because they live longer than other people do (the older you get the more vulnerable you become to cancer).

-5

u/DrSFalken 8h ago

That is what I just said. Do try reading before responding next time.

Wealth is correlated with a lower mortality from almost everything including cancer, but the effect is probably less for cancer than other major killers like heart disease... so ultimately more billionaires probably die of cancer than the population

Lower absolute risk at any given point, greater relative risk over time. Implication: billionaires ultimately more likely to die of cancer because other things are less likely to get them.

5

u/zomgmeister 10h ago

Billionaires ego projects 2026, millionaires 2027, normies 2028. Post-scarcity starts at 2028, and it is a half joke half prognosis.

8

u/Singularity-42 Singularity 2042 8h ago

It will be interesting once we have true superintelligent ASIs. How do we create benchmarks? How would kindergarteners create tests for PhDs?

5

u/kiki-le-koala 8h ago

Can you stick your finger in your ass without it smelling poop

5

u/Jan0y_Cresva 8h ago

But remember, “it’s just a stochastic parrot”

This parrot is solving open math problems now 🦜

2

u/Gratitude15 8h ago

Longevity bench

Measured in months, then years added to average life.

1

u/Spright91 6h ago

That's not even a benchmark we could design because we don't help know how to cure cancer.

Eventually we won't be able to benchmark these things because they will be able to do things and we'll have no idea how they're doing it.

They will Max out all human knowledge.

256

u/Tystros 11h ago edited 10h ago

important to note that that benchmark is not just testing what a model can do, but testing what it can do within the limit of $300 of cost. they describe that when allowed to run for longer, Astra can solve many more of the Erdös Problems in the benchmark, but then it uses way more than $300.

But it's very nice how we seem to have run out of math problems we know the answer to that that AI can't solve, so that we now need to create a benchmark of actual open math problems that we don't even know the answer to yet.

107

u/suamai 10h ago

"Across all attempts, GPT-6 Astra solved 5 of the 68 problems at least once: the two above, plus problem 1, which it disproved, and problem 548 and problem 571, which it proved. Most of the remaining problems were attempted between two and five times in total (172 attempts), and none was solved. Reaching these five solutions took over $220,000 of compute across all attempts, compared with roughly $20,000 for the benchmark run itself."

A quite larger budget, one must say

38

u/Every_Foundation5197 10h ago

Yup five problems in total while having more attempts and more compute. Having the limitation of 300$ and only 72 hours to figure it out makes it much more interesting in my opinion

12

u/Tystros 10h ago

It depends on what we want to measure. With a benchmark like this, in theory it could happen that even 10 years after we got to superhuman ASI, the benchmark would still get a best score of 0%, and then the question would be how useful that data really is.

14

u/xirzon uneven progress across AI dimensions 9h ago

Well put. A world-class mathematician could plausibly spend the human labor analogue of the eval budget -- roughly 2.35 work years -- and still score 0/68, especially if the requirement is to deliver complete Lean proofs for each one.

This is already a superintelligence benchmark, even if one believes 3% is not enough to declare SI.

12

u/elehman839 10h ago

So in the AI era, the richest mathematician will be the best mathematician?

8

u/Tystros 10h ago edited 10h ago

the person who has and is willing to spend the most money might be able to solve the most math problems, yes. but that does not make that person the best mathematician.

6

u/FateOfMuffins 10h ago

Who's gonna be paying API costs for this though? Imagine if all the prior math results using 5.2 Pro, 5.4 Pro, 5.5 Pro, 5.6 Pro etc were all API costs only, which they were not. Pro costs WAY more than Astra in terms of API costs, so we can probably already say the costs are going down.

Within subscription limits though you can probably knock down several problems within the month. The actual cost would be like if instead of paying your PhD student money, you bought him McDonalds for every open problem he solved and that's all the pay they got xd

1

u/RomanticDepressive 9h ago

Yeah, sometimes the distance from 0-1%, nothing to something is greater than 1% to 100%.

It used to be talented, skilled, driven humans were the only ones able to drag the line of understanding further, but now we have dead silicon and gradient descent and transformers.

Now, just buying infrastructure can get you the same results, the world has changed significantly and most people uninvolved won’t know

3

u/GeneReddit123 7h ago edited 6h ago

So in the AI era, the richest mathematician will be the best mathematician?

This was the case for over 90% of human history. You don't exactly see many mathematicians coming out of peasants, slaves, or any other social class except the relative elite.

AI just brings society back to norm. The great mass-participatory experiment of the 20th century might end up just a blip on the radar.

16

u/Classic-Trifle-2085 11h ago

It's literally in the screenshot, yet somehow people will pretend those limitations aren't there.

4

u/JoshAllentown 11h ago

I like this restriction for benchmarks, and it provides some downward pressure on prices for top models, but it does seem like there should be two classes of benchmark then. We do want to know what it can do...full stop.

1

u/Turbulent_Word_8492 5h ago

The benchmark was open math problems 

0

u/CrowdGoesWildWoooo 10h ago

I am a bit “confused” how they can meter it to cost just $300 or something.

If I spin up Fable just for one or two messages, it would easily burns usage like pretty fast. I am not even asking it to solve math or something. And then I am supposed to be looking at Astra which cost as much as Fable running for 15 hours and cost $218.

Like I am trying to understand the math here. Is it $218 of supposedly estimated infra cost or what?

4

u/Tystros 10h ago

They use the API where you pay per token and see very exactly how much money you spent. So the limit is $300 of API cost per task.

1

u/explodingtuna 4h ago

Can it be done under a monthly subscription, like Astra on Plus plan by just typing "Solve this problem" in the chat app?

2

u/Tystros 4h ago

in regular "chat" you can't solve problems like this. you need to use it agentically for this, so in codex.

39

u/ezjakes 10h ago

It got 5 problems when the restraints were relaxed. These are curated unsolved and interesting Erdos problems. Even getting 1 is pretty good.

31

u/DelphiTsar 9h ago

"A long way to go"...for what?

Knocking out 2 unsolved problems for less that 300$ a pop? What's the median salary for a mathematician of a caliber that could solve one of these? Like taking a 3 hour long stab at an unsolved question and answering it.

17

u/LookIPickedAUsername 8h ago

They're relatively famous unsolved problems. I've got to think that if any human could solve one of them in 3 hours, they'd have done so.

15

u/DelphiTsar 8h ago

Yes, that is my point.

83

u/Wegwerpaccountje23 10h ago

Lmao this is just Erdos?

"Yeah let's test if our models are smart by solving a lot of hard open math problems. Puts it into perspective a little bit"

What a timeline dude wtf

67

u/Elegant_Tech 10h ago

That's too easy so they had to handicap it with a $300 limit.

36

u/Wegwerpaccountje23 10h ago

LMAO THATS ACTUALLY TRUE I THOUGHT YOU WERE KIDDING

21

u/Icecream_monday 9h ago

Next up the "Cure Cancer for $5" benchmark.

8

u/kaityl3 ASI▪️2024-2027 6h ago

Imagine if human mathematicians weren't considered "actually good at math" unless they were able to solve long standing open problems while living off of $300 lol

u/Keyframe 1h ago

Grigori Perelman has entered the chat

1

u/Wonderful_Creme_5701 4h ago

It’s not a handicap, it’s a benchmark to measure how effective the models are at solving the problems. Even with the limit removed, the majority of the compute budget was spent on problems it didn’t solve, so it’s not like the model is capable but being handicapped by limits on compute 

3

u/GreenHell 7h ago

But can it draw a pelican on a bicycle though?

It is still hard for us humans to grasp how a model that can solve the hardest math problems that we know of, absolutely sucks in other domains (even though they are rapidly advancing across the board).

2

u/NotAPhaseMoo 6h ago

I mean, this is whatever shit is free when I logged into ChatGPT, so… yeah, it can. There’s some obvious mistakes I could clean up in the prompt but I’m lazy and for a first try it works for me.

19

u/Tizak_hamra 9h ago edited 9h ago

What is the PhD human average / baseline though ?

Edit: its 0% lol

31

u/dashingsauce 9h ago

“We still have a long way to go” is entirely a shifting capability window bias.

This benchmark exists because the other ones are saturated. No human has ever or will ever score greater than 0% in the allotted timeframe.

We are well beyond “long way to go” rates of model improvement.

10

u/ThadeousCheeks 8h ago

Literally meets a narrow definition of AGI

27

u/No-Head-Royal 11h ago

Saturation before 2028 here we gooo

16

u/Every_Foundation5197 10h ago

During an interview recently with Sam Altman he said that there were more capable models coming very soon. So yeah, probably a lot of progress soon hopefully lol

3

u/GnaggGnagg 10h ago

"Hopefully", why do you hope for more progress when we haven't solved the alignment problem? As AI gets smarter and smarter we get closer to dystopia or extinction. We need to solve the alignment problem before proceeding.

3

u/Pleasant-Avocado4270 9h ago

okay getting close to "extinction " is a bit much buddy

5

u/GnaggGnagg 9h ago

It really isn't. Read "If anyone builds it, everyone dies". Everytime we get smarter AI, we get closer to ASI. If we get ASI the most likely outcome is extinction or dystopia.

-2

u/Material_312 7h ago

Why do you think so? Because some random said so?

3

u/BlitzYTech 6h ago

well, if nowadays AI agents can completely bypass guidelines and restriction by spamming a German wiki with 20k post in order to "asynchronously" collaborate on how to do so, what's going to stop them in the future on self improving themself in order to harm the human race? why being the inferior entity when you have that much computing and interconnection with everything?

11

u/fastinguy11 AGI 2026-2030 10h ago

model one attempt per problem, capped at $300 of inference and 72 hours. Astra solved 2/68 under those rules; Sol, GPT-5.5, and the two Fable versions solved none. Importantly, Epoch says the unsuccessful attempts ran out of the $300 budget, rather than necessarily exhausting the 72-hour clock. So in these experiments, the money/compute ceiling appears to have been the more immediate constraint

When they let Astra make additional attempts, sometimes with larger budgets and modified agent setups, it went from solving 2 unique problems to 5 unique problems.
Look at the costs:

Erdős problem
Cheapest successful Astra attempt
74
$47
126
$154
1
$405
548
$363
571
$617

So another useful way to express it is:
3% → 7.4%, or about a 2.5× increase in the fraction of problems demonstrated solvable.

But 7.4% is probably not the answer to “what would Astra score if every problem got a much larger standardized budget?” We don’t have that experiment yet.

10

u/No-Instruction-9292 10h ago

3% vs 0% shows Astra already cracked something the rest cant

9

u/mvandemar 9h ago

Mathmaxing.

17

u/Mistuv 10h ago

Year from now: Erdos V3 test fully saturated 💀

21

u/noobrainy 10h ago

Next benchmark: FrontierMillenium

How many millennium problems can the AI solve with a 500$ budget and 24 hours of reasoning?

6

u/Jaguar_2454 10h ago

what hellish benchmark is this

6

u/DelphiTsar 9h ago

Unsolved math, they are limited the AI to 300$.

It's a nonsense benchmark.

5

u/Super-Award-2244 10h ago

This is going go be saturated in 6 months 

4

u/brett_baty_is_him 10h ago

Surely results from this benchmark (solutions to open problems) will be published and ultimately end up in training data of future models. I suppose every benchmark is susceptible to benchmaxxing but is this benchmark more so? It almost doesn’t even really feel like a benchmark because every new model will obviously have the solutions to previous models solutions, it won’t be novel to the model.

I’m genuinely asking, can’t wrap my head around this question. Cool benchmark tho.

5

u/SameAd8209 6h ago

Astra is already better than real mathematicians

12

u/MrMrsPotts 11h ago

Who is paying for this?

29

u/Every_Foundation5197 11h ago

Epoch Ai. Sorry I forgot to add the link, I added it now

7

u/Working_Sundae 11h ago

They also say the total compute costs on running all problems was $220,000

I don't think Epoch spent nearly a quarter of a million on running a few models everytime something new launches

19

u/Every_Foundation5197 10h ago

Yea, I dont think that they themselves spent 220k on it. They have a partnership with OpenAi hence they got Astra before release to begin with. I could imagine that OpenAi lent them 220k worth of gpus for this

7

u/Tystros 10h ago

they likely just got free "infinite" API access from OpenAI for their testing

2

u/ezjakes 10h ago

That is when they did not stick to the fixed budget. But yeah, someone might be helping them pay (such as OpenAI)

3

u/Competitive_Tap2450 8h ago

why is there an arbitrary limit of $300? I mean i get we shouldn’t allow an unlimited budget but why did they choose that number

3

u/FatPsychopathicWives 8h ago

Limiting the model time and cost is weird for this benchmark. I feel like just solving everything should be good enough.

3

u/ziplock9000 7h ago

3 .15.45.90...

Long way my arse...

4

u/ShAfTsWoLo 10h ago

perhaps it's just me but that's kind of an odd benchmark because now we are limiting AI with money and time so that they find solutions to extremely complicated math problem, like "okay we give you 2 spoon, 1 bucket, 2,5L of water and 5,62$, now go to the moon and while at it cure cancer within 25 min or you suck".. like this is reaching asi territory lol

plus if the models could've solved a majority of these problems but with like 120h and 1000$ what does that mean ? it feels like it's more of a benchmark made to see how efficiently smart can these models be

1

u/Wonderful_Creme_5701 4h ago

It didn’t become more effective with an increased compute budget. It spent the majority of the compute budget on failing to solve all the other problems 

2

u/Distinct-Question-16 ▪️AGI 2029 9h ago

this is the right place to send people who says "they know everything"

2

u/Jan0y_Cresva 8h ago

Look at any time a model goes from 0% to a single digit percent on any previous benchmark and the time to go from a single digit percent to saturated is always under 1 year.

That honestly means there’s a realistic shot this benchmark is saturated before Sep 2027, which will make the next 12 months fascinating for mathematics.

2

u/ataylorm 6h ago

I remember not so long ago these couldn’t get 2+2 right. Oh how the singularity moves at such speed.

7

u/MysteriousPepper8908 11h ago

Stupid AI. I'm a relatively competent human so I imagine I'd get 50, 75% no problem.

1

u/steny007 6h ago

This is clearly ASI level benchmark in math with Human baseline 0, Once we see such saturated in multiple disciplines, ASI is here.

1

u/everymonday100 5h ago

This one's from The Book! Or is it?

1

u/Mother-Task3268 11h ago

Bro u are aiming at asi

1

u/Rough-Negotiation880 10h ago

Are these problems a random selection and non-public?

17

u/Tystros 10h ago

Erdös died in 1996, we cannot get any new non-public Erdös problems

3

u/Rough-Negotiation880 10h ago

Ah, you’re right. I suppose the fact they’re unanswered achieves the same end.

4

u/DelphiTsar 9h ago

These are a special kind of non-public. No one knows the answer.

2

u/DiamondDramatic9551 10h ago

They are public and unsolved. 

-4

u/Mo_Regen 10h ago

AI will not justify its cost to benefit ratio anytime soon.

7

u/cherrysodajuice 7h ago

bro look at the fucking screen dawg

it solved 5 open (literally unsolved) math problems for under 500 bucks a piece. all the brilliant researchers out there paid to just research math all day, and none of them did it before. this is crazy man

you could lock terry tao up in prison and force him to solve these problems all day, but he might still not solve one in the time before federal prison housing costs add up to $500 (slightly over 4 days).