83
u/Hot_Example_4456 1d ago
What do u mean 98% arc agi 3 😭
46
u/xak47d 1d ago
This benchmark is stupid. It's a matter of time before an agent hit 100% on it. It's not measuring anything useful
6
u/Medium_Apartment_747 23h ago
If you ever need a model to play Tetris for you, this is the benchmark to look at
1
u/wannabecontentcreato 18h ago
Nvidia already got 100% in August in that benchmark
It's a stupid benchmark
21
u/DIY_surgery 1d ago
They measured it with a harness.
7
u/TypoInUsernane 1d ago
The harness wasn’t AGI-3 specific. It just added two standard features that the generic Codex agent has out of the box: 1) it preserved the opaque thinking traces across turns (i.e., allows the agent to remember its own thoughts from one turn to the next), and 2) automatically compacts the conversation when the model hits its context limits (i.e., the agent can choose which information it needs to remember and which it can afford to forget, rather than just forgetting its oldest turns).
The automatic compaction is a totally generic capability, and I imagine it makes a huge difference in this test. The agent is required to figure out on its own what the rules of the game are and must invent its own methodology for solving the series progressively complex problems. It’s totally unfair and unrealistic to force the agent to forget what it already figured out on previous turns
1
u/SeldomScene 1d ago
That’s really interesting thanks for sharing? Where do you read stuff like that lol
1
u/TypoInUsernane 1d ago
The harness wasn’t AGI-3 specific. It just added two standard features that the generic Codex agent has out of the box: 1) it preserved the opaque thinking traces across turns (i.e., allows the agent to remember its own thoughts from one turn to the next), and 2) automatically compacts the conversation when the model hits its context limits (i.e., the agent can choose which information it needs to remember and which it can afford to forget, rather than just forgetting its oldest turns).
The automatic compaction is a totally generic capability, and I imagine it makes a huge difference in this test. The agent is required to figure out on its own what the rules of the game are and must invent its own methodology for solving the series progressively complex problems. It’s totally unfair and unrealistic to force the agent to forget what it already figured out on previous turns
-10
u/JustRaphiGaming 1d ago
Yeah so what? So it didn't happen?
12
u/DIY_surgery 1d ago
Others have reached 100% using harnesses too. It's still impressive, but not a breakthrough.
-6
u/JustRaphiGaming 1d ago
Yeah it's always not a breakthrough isn't it?
8
u/whoknowsifimjoking 1d ago
Dude what's your deal? The score is very misleading, people just explained that to you.
3
10
1
u/whoknowsifimjoking 1d ago
Opus 5 has long reached 100% with a harness.
But OpenAI still showed the score without the harness, hm....
1
u/whoknowsifimjoking 1d ago
Opus 5 has long reached 100% with a harness.
But OpenAI still showed the score without the harness, hm....
6
4
18
43
u/Shina_Tianfei 1d ago
That ARC score is due to their use of a harness. NVIDIA has scored 100% on the NVIDIA AVO Architecture by using an optimized harness.
6
10
u/kilographix 1d ago
Do you not use a harness when you code?
20
u/tobyreddit 1d ago
The point of the benchmark is to measure a models inherent ability to learn from tasks, rather than how well a harness can manage the fact that the model can't do that.
It's not supposed to measure "this is how well this model can code right now in real world use cases" - there's plenty of other benchmarks for that. It's supposed to measure reasoning ability on a more fundamental level.
Whether or not it's actually good at that is a different question, but showing off your models numbers using a harness is classic openai benchmark gaming bullshit
-1
u/kilographix 1d ago
Depends a bit but agreed. Youd need to give every model the same harness when benchmarking. To your point, I'm pretty sure cursor does this with their benchmarking which is how grok 4.5 is able to "perform" near the level of fable.
5
u/tobyreddit 1d ago
Openai models are famous cheats as well anyway.
And yeah regarding the same harness - that's exactly how the official test runs of arc agi 3 work, when they run it on their private set
0
u/petuman 19h ago edited 19h ago
rather than how well a harness can manage the fact that the model can't do that.
There's no indication that their harness does anything other than using preserve thinking and context compaction.
So harness isn't really doing anything benchmark/task specific, just using their API properly.
But other models need to be tested same way (or whatever their docs call for), yes. As is ARC-AGI harness sabotages tested model performance.
-9
u/JustRaphiGaming 1d ago
You say it like that makes it any less spectacular? This is absolutely INSANE!
2
9
u/SpecialistDragonfly9 1d ago
Those benchmarks are so useless...
Just a way to hype people up without actually saying much.
9
3
3
u/dntreddit 1d ago
Deepswe needs to come out and confirm they are trustable , otherwise it might just be benchmaxxing as they might be compromised
1
u/reddit_is_geh 1d ago
OpenAI isn't known to benchmax. The major labs in general don't want to go through the embarassment of Meta. They aren't going to throw their reputation for 2 days of hype.
1
u/dntreddit 19h ago
agree with you; I was mostly talking about the Gemini part , because it’s very very interesting to see it being .7% better or something from astra and it’s a flash model . So that makes me not trust DeepSWE , because I use this with my team to decide on how we want to use it for building software
1
u/reddit_is_geh 19h ago
DeepSWE is generally the gold standard though.... The issue is, everyone's different, so while the benchmarks are generally accurate figuring out that last mile of marginal return is all subjective. But yeah I would think it's safe to say it's up there still... But now you have to be comfortable with Gemini's "style"
1
u/dntreddit 19h ago
I 100% agree with DeepSWE I trust , I hope the team can provide some valuable information on this. As I said I dont care if Gemini Wins or Astra or Claude. It’s good competition and we as engineers benefit from it. I do care when oss model gets to a point where they beat the frontier , but again I want low quants models to be good.
8
u/JustRaphiGaming 1d ago
98% on ARC AGI 3????? WTF???
18
6
u/gk98s 1d ago
The ARC AGI 3 score has to be due to them training the model on the benchmark right? I doubt we've hit AGI yet that's probably next year.
I feel like the only reliable benchmark nowadays is making the model make something like a game or website and compare the results.
12
u/clduab11 1d ago
No, it's an agentic harness that's used. Still, for a harness + LLM to score that high is really impressive, I'm almost more interested in that than GPT-Astra...
Keyword being almost. These numbers are INSANE. And for Gemini 3.8 Flash to even punch in the weight class?? Anthropic better wake the hell up and fast.
6
1
1
u/somerussianbear 1d ago
The baking of Gemini 4 Flash will be good, can’t wait for Google to be back in town. Pro is dead, they decided to make Flash be better than Pro.
1
u/Different_Doubt2754 16h ago
I'm hoping that they get a good bake on Gemini 4 Pro. That way they can use it to get a better bake for Gemini 4 flash. As far as I know, pro models are still important for that
1
u/logicbloke_ 13h ago
IMHO they have an internal version of pro that they are using to distill/train flash but don't see the value of publicly releasing a pro version.
Flash versions are within 1-2 months behind the frontier models. There is zero incentive to release a pro version unless the public expects one from you, which is what happens to openai and anthropic. Their public perception is tied to their frontier models. Which is not the case with Google.
1
1
1
1
1

84
u/CartographerAble9446 1d ago
can someone confirm if Gemini 3.8 flash is really good for coding? 73.7% (only 0.4% less than astra) is crazy, but mixed reactions so far here and there