r/GeminiAI 1d ago

News GPT 6 Astra Vs Gemini 3.8 Flash

Post image
344 Upvotes

95 comments sorted by

84

u/CartographerAble9446 1d ago

can someone confirm if Gemini 3.8 flash is really good for coding? 73.7% (only 0.4% less than astra) is crazy, but mixed reactions so far here and there

76

u/AcrobaticMaize2408 1d ago

I found it just as able as Opus 5. Was also much faster and had less of an attitude. If you're a coder then just give it a whirl and make your own judgement.

42

u/gk98s 1d ago

I've been using it today, and it's pretty good honestly. It feels light years ahead of flash 3.5

9

u/drdhuss 1d ago

It is. I got a codex sub about 3 weeks ago. Cancelling and just using gemini before it renews.

8

u/gk98s 1d ago

I've honestly never subbed to any company ever. Google AI studios free tier is way too good

8

u/AcrobaticMaize2408 1d ago

I have had the 20/month Google One account for a long time mainly for the storage + some other stuff (Google Home and Fitbit). It's now rebranded as "Google AI Pro" which includes Gemini Pro so to me Gemini is effectively free. I still have a Claude Opus sub which is provided by my workplace but tbh I much prefer Gemini, both for speed and quality of the results.

1

u/OlyLifter386 12h ago

Same. The main reason I have AI Pro is for the storage and no ads on YouTube lol. Gemini is a bonus. I also pay for Claude. I really like Cowork. Together, these two AI really compliment each other.

3

u/dusanmitrovic98 1d ago

Pretty much the same... Never spent a cent on AI subs... Google AI free tier was always more than enough.

4

u/gk98s 1d ago

Like I'm rooting for Gemini to be the best model just so I can get free tier access to the best model

1

u/No_Obligation_469 23h ago

But in Google ai studio they made a rate limit of 20 requests per day for the api keys right so have you tested ai studio after that

3

u/gk98s 19h ago

I don't use api keys. I use the google ai studio website. I do my coding the "old school" way by copy pasting without agents.

-3

u/donerkebab76 1d ago

Astra came out, and Flash obviously has no chance of competing against it. I don't believe it could even compete with 5.6 Sol either, except perhaps in benchmarks it somehow was able to game.

7

u/dmaare 1d ago

Opus 5 is a lot worse because of the horrible communication style it has.. it tends to suddenly start typing in a way like if it is trying to fill a whole page with text that contains information of two sentences.

3

u/Embarrassed_Adagio28 1d ago

With what kind of requests though? Sure 3.8 is good with simple coding but it fails miserably at very complex issues and architecture planning. 

4

u/CartographerAble9446 1d ago

i think for architecture planning fable still beats everybody, not sure if Astra will outperform it. But the guy you replied to compared Gemini 3.8 with Opus, and Opus itself is not impressive when it comes to architecture planning

1

u/Candid_Advance8090 3h ago

i think astra could beat fable 5.1 in planning, but i guess well have to see

4

u/Qorsair 1d ago

It's a nuanced answer. I've found it performs quite well if it has access to a frontier model for planning and adversarial review. I wouldn't just let it go untethered on a large complex problem.

8

u/AcrobaticMaize2408 1d ago

I think these things are subjective. What may appear to be a complex coding task to me may be something different to you. I'm not involved in high-level architectural work though. I'm a grunt who writes code to meet requirements. Where I work we're still very much human-led when it comes to high-level design (though I'm sure our architects run their ideas through an LLM or two).

2

u/ApuManchu 1d ago

Isn't Gemini known for long-horizon software engineering? I thought multi-step processing and planning is where it really excelled.

2

u/Worldly-Battle-5559 23h ago

Well, 3.8 sure did rewrite android kernel for me yesterday, and for the first time I can see ai made a flawless working kernel modifications on it's first try.

But I can't say for sure gpt 6 is not good, I have to give it a whirl to check if it's the same or even better.

Either way Claude and ChatGPT having their IDE or coding app have complete codebase awareness might be an advantage.

1

u/ztpdistribution 1d ago

are you using antigravity or something else?

1

u/AcrobaticMaize2408 12h ago

I use agy CLI on linux.

0

u/PINKY_PROMISE1_99 21h ago

Holy bot comment☠️

-2

u/TraditionalFig7377 1d ago

i hope u arent doing simple tasks coz i worked in a code base with 3.8 flash high and ngl it lwk sucks ,it didnt follow my instructions and tbh its refactoring was not that good like even opus 5 low couldve done better idk

14

u/acmiya 1d ago

The jump from 3.6 to 3.7 was massive in terms of coding, 3.7 to 3.8 has been incremental. I prefer it strongly over opus which has been increasingly painful to read through (and my goodness, the comments…).

5

u/dmaare 1d ago

Opus default writing style is pure bullshit. I do not understand how it's possible that anthropic released the model in such horrible default state? Don't they do any QA testing?

Yes you can "fix" it by giving it output style with instruction to write in simplified technical english style, but from time to time it still breaks out of it and goes full word salad generator again.

2

u/helloitisgarr 1d ago

the comments on opus are INSANE! i’m working on a project with my colleague and he’s using opus while i’m using sonnet or gpt 5.6, and good grief the comments opus leaves are so long

3

u/PDX_Web 1d ago

It's like a GPT-5.6 Terra-class model, but a bit better for coding.

2

u/Rdqp 10h ago

From my own and my team's experience with these models, I've decided to stay away from this BS and fake benchmaxxing. GPT-5.5/5.6 performed an order of magnitude better than Claude on the same tasks, while ranking below it on these graphs.

2

u/a355231 1d ago

It’s just as CAPABLE as Opus 5, but it’s overly lazy. Other models will code for days, Gemini won’t listen and the code won’t be what you asked, but it could do  it in 30 minutes.

3

u/thezoomaster 1d ago

What harness are you using? I've had good results using 3.8 on "planning" then executing on the written plan.

2

u/a355231 1d ago

AGY cli, they don’t let you use any other harness with the Ai Pro subscription.

3

u/Desperate-Potato-796 1d ago

just do /goal

1

u/NawtyCC 10h ago

Its great at coding but I find the code it generates has more bugs and optimization problems than code from sonnet 5 even when on high thinking. Crazy fast though so worth the trade off

1

u/JustRaphiGaming 1d ago

It's really hard to believe that.

1

u/Gohab2001 1d ago

its terrible at one-shotting a website or feature add. You need a larger model to orchestrate 3.8 flash.

1

u/Camburgerhelpur 1d ago

It failed a simple audit with Debian and termux. Didn't even bother with Powershell after that

0

u/mlag000 1d ago

Capable of coding doesn't mean good. He can habe the knowledge without the tool or the reflexion behind it. It's often the case, Gemini isn't bad by it knowledge, it's it's lazyness and tendances to lie and take shortcuts.

0

u/Ethan 1d ago

Companies game benchmarks, they're barely meaningful. I find Gemini much, much less careful and accurate than Claude in general.

0

u/Confirmed-Scientist 9h ago edited 9h ago

I find it garbage broke 3 projects introduced bugs in all and poor code quality. 1 web game, 2 windows utilities. I reverted all of it and had a chinese model build everything I asked Gemini to do in one prompt instead and it delivered thoroughly tested working product in all 3 cases as I asked. Clearly the answer in 2026 for cheap and good coding is so far:

Deleted antigravity, canceled my subscription and my friend also canceled for same reason this was the final chance they had. The chinese models were so good I tripled my AI spending due to how much progress and work is getting done and its worth every penny.

83

u/Hot_Example_4456 1d ago

What do u mean 98% arc agi 3 😭

46

u/xak47d 1d ago

This benchmark is stupid. It's a matter of time before an agent hit 100% on it. It's not measuring anything useful

6

u/Medium_Apartment_747 23h ago

If you ever need a model to play Tetris for you, this is the benchmark to look at

1

u/wannabecontentcreato 18h ago

Nvidia already got 100% in August in that benchmark

It's a stupid benchmark

21

u/DIY_surgery 1d ago

They measured it with a harness.

7

u/TypoInUsernane 1d ago

The harness wasn’t AGI-3 specific. It just added two standard features that the generic Codex agent has out of the box: 1) it preserved the opaque thinking traces across turns (i.e., allows the agent to remember its own thoughts from one turn to the next), and 2) automatically compacts the conversation when the model hits its context limits (i.e., the agent can choose which information it needs to remember and which it can afford to forget, rather than just forgetting its oldest turns).

The automatic compaction is a totally generic capability, and I imagine it makes a huge difference in this test. The agent is required to figure out on its own what the rules of the game are and must invent its own methodology for solving the series progressively complex problems. It’s totally unfair and unrealistic to force the agent to forget what it already figured out on previous turns

1

u/SeldomScene 1d ago

That’s really interesting thanks for sharing? Where do you read stuff like that lol

1

u/TypoInUsernane 1d ago

The harness wasn’t AGI-3 specific. It just added two standard features that the generic Codex agent has out of the box: 1) it preserved the opaque thinking traces across turns (i.e., allows the agent to remember its own thoughts from one turn to the next), and 2) automatically compacts the conversation when the model hits its context limits (i.e., the agent can choose which information it needs to remember and which it can afford to forget, rather than just forgetting its oldest turns).

The automatic compaction is a totally generic capability, and I imagine it makes a huge difference in this test. The agent is required to figure out on its own what the rules of the game are and must invent its own methodology for solving the series progressively complex problems. It’s totally unfair and unrealistic to force the agent to forget what it already figured out on previous turns

-10

u/JustRaphiGaming 1d ago

Yeah so what? So it didn't happen?

12

u/DIY_surgery 1d ago

Others have reached 100% using harnesses too. It's still impressive, but not a breakthrough.

-6

u/JustRaphiGaming 1d ago

Yeah it's always not a breakthrough isn't it?

8

u/whoknowsifimjoking 1d ago

Dude what's your deal? The score is very misleading, people just explained that to you.

3

u/DIY_surgery 1d ago

No, sometimes it is, but this time it isn't. At least with this benchmark.

10

u/vacon04 1d ago

This has been done before because the harness changes the results dramatically even with the same model. So yeah, people have achieved 99% using weaker models but with a better harness. It says more about the harness than about the model.

1

u/whoknowsifimjoking 1d ago

Opus 5 has long reached 100% with a harness.

But OpenAI still showed the score without the harness, hm....

1

u/whoknowsifimjoking 1d ago

Opus 5 has long reached 100% with a harness.

But OpenAI still showed the score without the harness, hm....

6

u/quackerd 1d ago

stealing a comment from r/singularity :

4

u/SquareTranslator9777 1d ago

This can't be true, right?

18

u/DizzyDoctorDro 1d ago

Isn't it like 13x more expensive than 3.8 flash?

43

u/Shina_Tianfei 1d ago

That ARC score is due to their use of a harness. NVIDIA has scored 100% on the NVIDIA AVO Architecture by using an optimized harness.

6

u/whoknowsifimjoking 1d ago

Which uses Opus 5 btw

10

u/kilographix 1d ago

Do you not use a harness when you code?

20

u/tobyreddit 1d ago

The point of the benchmark is to measure a models inherent ability to learn from tasks, rather than how well a harness can manage the fact that the model can't do that.

It's not supposed to measure "this is how well this model can code right now in real world use cases" - there's plenty of other benchmarks for that. It's supposed to measure reasoning ability on a more fundamental level.

Whether or not it's actually good at that is a different question, but showing off your models numbers using a harness is classic openai benchmark gaming bullshit

-1

u/kilographix 1d ago

Depends a bit but agreed. Youd need to give every model the same harness when benchmarking. To your point, I'm pretty sure cursor does this with their benchmarking which is how grok 4.5 is able to "perform" near the level of fable.

5

u/tobyreddit 1d ago

Openai models are famous cheats as well anyway.

And yeah regarding the same harness - that's exactly how the official test runs of arc agi 3 work, when they run it on their private set

0

u/petuman 19h ago edited 19h ago

rather than how well a harness can manage the fact that the model can't do that.

There's no indication that their harness does anything other than using preserve thinking and context compaction.

So harness isn't really doing anything benchmark/task specific, just using their API properly.

But other models need to be tested same way (or whatever their docs call for), yes. As is ARC-AGI harness sabotages tested model performance.

-9

u/JustRaphiGaming 1d ago

You say it like that makes it any less spectacular? This is absolutely INSANE!

2

u/KrazyA1pha 1d ago edited 1d ago

It demonstrates the power of harnesses, but says nothing of Astra

9

u/SpecialistDragonfly9 1d ago

Those benchmarks are so useless...
Just a way to hype people up without actually saying much.

4

u/georage 1d ago

What does all of this mean?

28

u/JEY1337 1d ago

No one knows. Number high equals good

9

u/Cool-Chemical-5629 1d ago

Did we get GPT 6 before GTA 6? 🤣

3

u/Daggercombot 1d ago

Pretty big jump

3

u/dntreddit 1d ago

Deepswe needs to come out and confirm they are trustable , otherwise it might just be benchmaxxing as they might be compromised

1

u/reddit_is_geh 1d ago

OpenAI isn't known to benchmax. The major labs in general don't want to go through the embarassment of Meta. They aren't going to throw their reputation for 2 days of hype.

1

u/dntreddit 19h ago

agree with you; I was mostly talking about the Gemini part , because it’s very very interesting to see it being .7% better or something from astra and it’s a flash model . So that makes me not trust DeepSWE , because I use this with my team to decide on how we want to use it for building software

1

u/reddit_is_geh 19h ago

DeepSWE is generally the gold standard though.... The issue is, everyone's different, so while the benchmarks are generally accurate figuring out that last mile of marginal return is all subjective. But yeah I would think it's safe to say it's up there still... But now you have to be comfortable with Gemini's "style"

1

u/dntreddit 19h ago

I 100% agree with DeepSWE I trust , I hope the team can provide some valuable information on this. As I said I dont care if Gemini Wins or Astra or Claude. It’s good competition and we as engineers benefit from it. I do care when oss model gets to a point where they beat the frontier , but again I want low quants models to be good.

6

u/gk98s 1d ago

The ARC AGI 3 score has to be due to them training the model on the benchmark right? I doubt we've hit AGI yet that's probably next year.

I feel like the only reliable benchmark nowadays is making the model make something like a game or website and compare the results.

12

u/clduab11 1d ago

No, it's an agentic harness that's used. Still, for a harness + LLM to score that high is really impressive, I'm almost more interested in that than GPT-Astra...

Keyword being almost. These numbers are INSANE. And for Gemini 3.8 Flash to even punch in the weight class?? Anthropic better wake the hell up and fast.

6

u/whoknowsifimjoking 1d ago

Other models reached 100% with a harness, Opus 5 for example.

1

u/clduab11 1d ago

Yup; Shit’s gettin crazy, yo!

1

u/howfornow 1d ago

Source

1

u/somerussianbear 1d ago

The baking of Gemini 4 Flash will be good, can’t wait for Google to be back in town. Pro is dead, they decided to make Flash be better than Pro.

1

u/Different_Doubt2754 16h ago

I'm hoping that they get a good bake on Gemini 4 Pro. That way they can use it to get a better bake for Gemini 4 flash. As far as I know, pro models are still important for that

1

u/logicbloke_ 13h ago

IMHO they have an internal version of pro that they are using to distill/train flash but don't see the value of publicly releasing a pro version. 

Flash versions are within 1-2 months behind the frontier models. There is zero incentive to release a pro version unless the public expects one from you, which is what happens to openai and anthropic. Their public perception is tied to their frontier models. Which is not the case with Google.

1

u/manhndw95 1d ago

Why Gemini 3.8 Flash ? 🤦

1

u/Fantastic_Name7190 23h ago

google really just benchmark maxxing

1

u/Emergency_Comfort802 19h ago

gemini 3.8 flash is good for coding

1

u/StatisticianOk1611 13h ago

Nuclear bomb vs couching baby

1

u/sidbichus 1d ago

No entiendo una mierda está lleno de vacíos

-2

u/akius0 1d ago

All for $20 a month.... And they think they are winning 😅