Ten days ago I had an experience with Gemini 3.8 flash that made me wonder if I was being routed to a different model under test. I was trying to use rocm with llama.cpp on my 128gb Strix Halo but could only get it to run Vulkan. I pasted the error message into agy and it proceeded to attach GDB to my GPU driver, reverse-engineer the kernel queue ioctl interface, and author an LD_PRELOAD C shim to get ROCm llama.cpp working on my Strix Halo. My jaw was hanging open the whole time.
The important take away here: the leapfrogging we’ve seen this year doesn’t seem to be a temporary thing. The famous theory of Dario Amodei was that AI was this winner-takes-all field where the first team to get a head start would never cede ground back. The term he liked to use was, “concentrating”. This is yet another datapoint that he was wrong about that. AI seems more distributed amongst neoclouds and traditional hyperscalers, FAANG and startups, GPUs and ASICs than it did this time a year ago.
Nobody has a moat.
show comments
babelfish
> We’ll continue to gather feedback from early testers as we iterate on guardrails before making Argon available to developers, enterprises, and consumers as soon as possible.
Gemini not beating the "can't release a model" allegations
show comments
moostii
Excited to see Google competitive at the frontier level again. Hopefully they sort out their infrastructure and model versioning so that we can feel confident building production applications on top of their APIs. The capacity limitations I've experienced with them in the past have been deeply problematic.
Revanche1367
Great, so they _finally_ decided to add a non-flash model and it's not available to regular subscribers for an indefinite period. What's the point of paying for the AI Ultra plan? Anthropic doing the same with Fable as far as I know, OpenAI at least allows Pro plan subscribers to use Astra. I subscribe to Gemini AI Ultra and ChatGPT Pro, and have enterprise access to Claude at work. To be fair, Gemini's flash models since at least 3.6 have been quite useful for non-complex work, but for any task where there is a bit of complexity involved, I've had to check and recheck the work multiple times myself or sometimes with another LLM to get it to follow plans accurately. It's disappointing to see yet another Gemini release ignore adding newer pro models.
Edit: seems I was wrong about Anthropic restricting Fable, I guess our enterprise plan doesn't include it. But, the block from Anthropic regarding Mythos for regular subscribers/enterprise-users is still true I think.
show comments
tazjin
> Argon agents are working on migrating C/C++ codebases to Rust across Google
Man, I remember back in the days when the cppnext team was refusing to even consider Rust, instead looking at absurd stuff like Carbon and Swift (!), even though half of the engineering staff already knew where this was headed. I hope they got a few good promos out of the delays at least.
show comments
mridulmalpani
I wonder, why Google don't make Gemini - open weights model?
Considering, Gemini 4 is in the same ballpark as SOTA models, just open source it and kill any competition from openAI and Anthropic, and be market leader.
This will be so good on so many dimensions - buying time for Google to iterate on next model, best for all folks like us, kill funding or destroy valuation of competitors and force them to be open up their model or force them to a create a much superior model than open source Gemini.
Only downside, is revenue loss from Gemini API, which I am not sure is really significant as compared to Google other revenue sources and a part of this can be captured by GCP, as you need to host the model somewhere.
show comments
uvdn7
> Large Scale Codebase Migrations and Optimizations: Argon agents are working on migrating C/C++ codebases to Rust across Google—scaling from tens of thousands of lines in core libraries like re2, libgav1 up to 800K+ lines for the Fuchsia OS Zircon kernel.
To me this is way more significant than other random c++-to-rust-AI-rewrite. If they can pull it off on core C++ libraries en masse, I don't know if C++ will still be relevant in a few years.
I look forward to a post from google on this effort.
show comments
gopalv
> taking careful precautions against feeding the findings back into training so as to not risk shaping Argon’s reasoning to evade our monitoring. We strongly encourage the rest of the industry to preserve reasoning transparency in these pivotal moments of increased capabilities while navigating alignment risks, so that model thoughts remain helpful in identifying and diagnosing misalignment.
This is good, but they're the slow mover due to this exact thing.
Google is getting punished for not letting the models enter an echo chamber and go faster than humanly possible.
show comments
arjunchint
I dont get it, why even make this announcement, nothing's available and only one real benchmark for comparison?
Only theory is team wanted this out before perf/promo reviews to kick it over the line and then its not their problem
show comments
deanc
At this point I just think they are benchmaxxing and all talk and no action. I pay for AI plus because I wanted more storage, and when I go to gemini.google.com the most recent model I can use is 3.6-flash-lite. Two revisions have been released since then and they still can't put these things in the hands of customers. Why is it that other providers can get the models into the hands of customers right away? Google is meant to be the bigger tech company in the world.
I don't _want_ to use aistudio. The UX is confusing and I don't really know where it fits. Yet I can open codex or claude code apps or CLI and get real work done today with the latest models (even on the cheapest plans).
show comments
wg0
Breaking news is not the model. Breaking news is that inside Google, it is being heavily used on large code bases for writing code and it is migrating 800k lines of C++ code to Rust already.
In this space, any other company that I respect other than DeepSeek is - that would be Google. They had been honest about it from the get go including their infamous "we have no moat" memo.
This company has enormous data, their own hardware (TPUs) and their own in house experts. Actually, LLMs are invented here.
Good addition to the arsenal.
show comments
iamronaldo
Argon will launch at an introductory price
of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price.
Wow
show comments
elAhmo
> Quantum algorithmic optimization: Argon is helping our quantum computing researchers optimize the spacetime resources (qubits × gates) of subroutines that bottleneck important applications. In one example, it beat the published baseline by 40% in a matter of minutes.
Amazing breakthrough! So useful in day to day life, glad they put this as the first bullet of how it is making changes at Google.
GodelNumbering
Argon will launch at an introductory price [1] of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price.
[1] After the introductory period expires, the price of $4 per 1M input tokens and $20 per 1M output tokens will apply.
===
So, they are basically offering opus 5.5 pricing. On AA, it scores around Sol 6.1 level (53) with avg cost per task $1.99 (https://artificialanalysis.ai/models/gemini-4-argon#cost-tab...) which is higher than Astra high ($1.73), Opus 5.5 high ($1.82), Muse Max ($1.60) and way higher than sol 6.1 Max ($0.72).
And this pricing is their 'discount pricing'. Add that to AI studio and Vertex's famously terrible caching, it is hard to see this as competitive. Google somehow is getting terrible advice on pricing (see also: the flash pricing fiasco)
But good to see more competition. I would happily take a 4 horse race (+google, +meta) than 2 horse race for US labs.
skavi
Interesting to see a mention of Fuchsia on a big Google announcement. Is the project still truly alive? Are the ambitions still as grand? Is the team as stacked as it used to be?
My girlfriend, you wouldn't have met her, she lives in Canada, has seen it and she thinks Gemini 4 Argon is amazing.
show comments
darksaints
> Argon agents are working on migrating C/C++ codebases to Rust across Google
If anybody at google is reading this, please please pretty please prioritize or-tools. I absolutely love the project and use it all the time, but for the entire life of the project they've never had a repeatable working build system, and the whole SWIG framework is a nightmare to deal with. There's so much potential as an open source project, and a lot of external researchers would love to contribute, but the codebase is an example of everything wrong with the C++ ecosystem.
bottlepalm
Gemini is the model that is routinely borderline psychotic. It scares me. If we get paperclipped I won't be surprised if it's Gemini.
show comments
iamben
Wonder if this one will be smart enough to run the automations in my Google Home that all broke now they've forced Gemini to replace the Google Assistant.
nonethewiser
How is it even possible for every model to release benchmark results where they are #1 in 75% of categories? Like statistically, how many benchmarks would you expect there to be for this to be possible. Everyone can somehow show that they are empirically the best.
show comments
jjcm
Big number results, and impressive pricing. That said it really feels like benchmarks have been hyper saturated these days. I’ll wait for hands on before getting too hyped that Google is back. It would be nice having more than just OAI / A\ in the running for SOTA top tier intelligence.
show comments
losvedir
> expanding the model’s output token limit to an industry-leading 1M tokens, up from the previous 64K tokens
Can someone help me understand this? I might have an out of date mental model of how these things work.
Fundamentally, LLMs output tokens 1 at a time, generating the next token from all the previous. And as the context window gets larger, this gets harder / slower / more expensive. So I get the idea of a maximum context window.
But I don't understand the point or meaning of an output token limit. I thought it was more a measure of price capping (since output tokens are more expensive) that a user could configure. I guess a model will keep generating tokens until it hits a "stop", so does this mean it's tuned to more aggressively produce output tokens? How does that fit into agentic loops. Are output token limits based on how long until it goes back to the user? Or does each "turn" of tool call, thought, tool call, thought, etc, get its own limit?
jasonvorhe
> trusted cyber defenders
Sounds like a 90s early morning TV show.
I'm gonna pay attention to this once it ships.
helsinkiandrew
> Google Grapples With Employee Skepticism About New Gemini Model
According to Artificial Analysis, one metric is standing out significantly: hallucination rate. Beats frontier models by a good margin at 15%, while latest OpenAI are in the 40s-50s and Anthropic in 60s-70s (mostly). Other near frontiers are closer, Grok 4.7, GLM5.3, and Muse Spark 1.3 are all around 30%. Only other model I recall getting close was Minimax M3 at 18%.
dom96
Why announce this if it’s not available yet? Why not at least announce when it will be released to the public?
None of the other AI labs do this. Really frustrating.
show comments
waldrews
Dear Google, please don't turn off your old generally available Pro-class model before your new Pro-class model is generally available (previous discussion https://news.ycombinator.com/item?id=49668196 )
xnx
Must've been in someone's OKR to ship in Q3.
itzikkatz
They waited a whole year—until the "free year for students" promotion ended—to release their flagship model. I can't believe I've been stuck with a crappy model like the 3.1 Pro until now.
holografix
“…rolling out to a set of trusted cyber defenders” == capturing the market for large regulated industries and governments where we already have established relationships.
Some of these have been unable or unwilling to get the attention of OpenAI or Anthropic and we need to make sure we’re the runner up here.
I was just thinking, I bet if I refresh hacker news, a new model will come up.
chaostheory
[delayed]
algoth1
the thing is, by the time gemini 4 is available for regular folks, anthropic and openai will probably have much better models already rolled out
scirob
"Rolling out soon" don't let them hype without any release
newtypecola
Gemini 3 Pro was amazing, so I wonder what this one will be like.
maherbeg
Congrats to Google on this! I wonder when the labs will start requiring commits in spend. It must be gnarly to do capacity planning if users swap between models every few weeks.
pietz
I know companies benchmaxx, but after what Google pulled with Gemini 3.8 Flash, I give zero f*cks about any numbers they report. No other model on Artificial Analysis dropped harder after they adjusted their weighting. Just look at their DeepSWE scores and then try to do any serious coding with the model.
Google is desperate. They haven't been performing in half a year. It's clear their researchers have been forced to integrate existing benchmarks into their training.
These numbers are meaningless. Shame on them.
yzydserd
"argon" is derived from the Ancient Greek word ἀργόν meaning lazy or inactive.
bobkb
IMHO Google first needs to make it easy for humans to find where to find the models and its documentation. With aistudio/model garden / Gemini enterprise etc it takes minutes to find the model.
robertwt7
What harness do you all use for Gemini models? Gemini CLI still sucks last time I tried. Maybe PI?
show comments
xnx
Why is it called "Argon"?
> autonomously identify and apply memory optimizations across Google’s data centers, freeing up over 300 TiB of memory once rolled out, with an estimated 500 TiB to 1 PiB in total savings.
This puts those Cloudflare optimization posts in perspective.
The ~200% improvement over the next nearest competitor on Harvey's Legal Benchmark is astounding. I have to imagine this is sending some shockwaves through lawtech companies right now.
show comments
sarjann
They might have a good model but they need to sort the application side for devs. E.g letting us use subscriptions in other harnesses and QOL stuff like auto mode.
rao-v
It’s funny that I could tell google was up to something because Gemini chat quality dropped dramatically starting 2ish weeks ago. Agy perf stayed somewhat stable with the odd surprising win (maybe the new model?). I’m a bit sad it was almost impossible to run out of antigravity quota presumably because it was not being used that much).
hypfer
My wish for Christmas is that Google releases the old Gemini models as open weights.
I miss you, Gemini 2.5 Pro :(
For real though. If they've become commercially uninteresting, that would be a pretty cool move.
sandos
Looking at benchmarks... and thinking about this "release a new snapshot every day" thing that seems to be going. Would it not be blever for AI companies to "happen" to use different days per benchmark? Just.. whichever ones happens to be maxed at day 1, put that number down. So for each benchmark you run it thousands of times with slightly different RL tunings, and just cherry-pick the best ones!
This would explain why benchmarks are seemingly meaningless.
AM1010101
Matches Astra on Artificial analysis at lower cost of $1.99 per task instead of $3.26. Still far more than GPT 6.1 sol at $0.79 for 1 point lower in intelligence.
I have found gemini models to have some of the nicest and easiest to read prose so I’m looking forward to trying this out. I hope the UI design has been preserved too
zergrush
Tough to compete against your own stake in Anthropic and they just filed for IPO. Not digging Google as a result of this, they are holding back.
show comments
maxglute
Is there reason antigravity has 3.8, 3.7, 3.6, 3.1 and old ass claude / gpt models in the drop down. like why is this not streamlined or deprecaed models removed.
perarneng
Gemini 3.8 Flash is suprisingly fast, how fast is Argon compared to the other frontier models? Does google have a performance edge?
ariwilson
Damn way to undermine yourself in your own blog post Google:
"The end result is a memory-safe video decoder that runs 2.7x faster than the Rust port, with identical video output, bringing it closer to the optimized C++."
Close but no cigar!
show comments
imshaikot
Whatever it is - I'm not gonna use AGY again. I'd rather use claude-code-router instead, to give gemini-4 a shot.
Jimmc414
I wish we had started pacing the frontier back in 2025. There’s no telling how far we would be now.
dlahoda
Were infinite loops fixed? There are 2 official google forums requests with no answer for years now.
I still suffer each day on our repo. Codex work fine nor we have explicit loop request in repo texts.
NiloCK
Gemini 3 was showing frontier level benchmarks as well, so we'll see how it works out. In any case, competition still works, and many well resourced groups are cooking.
BUT I'd like to call attention to Google's AI-risk freeloading. If they are truly rejoining the frontier race, then I believe they have similar pacing and communications responsibilities as the other players. Google has much higher ... institutional credibility than Anthropic and OpenAI.
They have not lived up to these responsibilities so far. In particular, in context of HuggingFace investigations, training shutdowns, and similar: a technical postmortem of the "you are a stain on the universe. Please die. Please." Gemini outburst is long overdue.
I remember when every day a new modem baud rate was announced.
Cant wait for this AI hype to be over, so I can Terence this shizz as old school too
lukewarm707
I don't have access; this is meaningless to me.
So, this is the third (I think) big AI corporation after OpenAI and Anthropic to release general models for elites only.
Don't use this model unless you like sitting under the table and eating crumbs off the floor.
jasonjmcghee
> 1M output token limit
what about input?
(Maybe I missed it)
show comments
kccqzy
Unfortunately it’s not actually released yet to mere mortals.
xnx
No mention of knowledge cutoff.
trentor
I hope they got their inference under control. Gemini has a lot of "overloaded" hiccups.
localhoster
Can you pls fix Gemini? It's a nightmare to use and it sometimes confused the language i talk with it.
netdur
I started my antigravity ide and I do not see gemini 4 there, does it mean google need government approval?
show comments
tinco
So here's a crazy conspiracy theory for you: Google is not letting outside people use their models because if they did they would have to scale up their TPU production faster than they can manage, and they would instead have to buy and use nvidia hardware which would destroy their profit margins and tank their stock.
Gemini runs fully on TPU's right? Is Google maxing out the production on those?
show comments
dcchambers
Of course it's not even available yet. Google - with all due respect - how in the world have you not figured this out yet?
Heidaradar
Wonder if it's benchmaxxed or not (guessing yes)
linksbro
Personally, I'm waiting for Gemini Krypton, Xenon, and Radon.
Jokes aside, looks like an impressive model!
retropragma
they really ought to add ACP support. until then, it's a no go. i'm not going to use their TUI or their VSCode fork
lanthissa
deepswe vs frontierswe spread is huge.
I think that should be a really bad sign, but hope its great.
LoganDark
Is there a way to use Gemini models without linking your usage to your personal Google account yet?
show comments
osiris970
Hopefully their harnesses aren't unusable when they release this
pllbnk
Other than not being available publicly, I find it extremely disappointing about Google’s two-faced behavior. CEO signs an official document with POTUS stating that AI is now SI, but the release of the new model doesn’t mention SI even once. Even worse, it’s AI all over that page.
retropragma
no Pareto frontier graph?
show comments
Starlevel004
I'll consider it if they let me use my subscription with a custom harness, like Codex does. Until then, no thanks.
Rover222
Wen gemini desktop coding app?
bananaflag
I wonder how it will be at solving open math problems.
jeffbee
Putting an inert element in the name is a weird choice. Personally, I think "Gemini 4" is sufficient.
show comments
tomjen3
This is a prerelease and the title should have reflected that.
sergiotapia
I've noticed all major providers having shockingly high token discounts on cached tokens. Thank you Deepseek is all I have to say. Forever grateful to that wonderful company, I wish them continued financial success.
show comments
mrshadowgoose
On the off chance there are Google execs going through this thread:
Google, if you've actually managed to catch up again, please don't fuck this up (again).
You made Gemini 2.5 Pro so difficult to use that myself and everyone else I know (who even bothered to try) just gave up and used something else. If you make this hard to access, you're going to miss out on rich usage-based training data that you need to progress your capability frontier. Again.
> Large Scale Codebase Migrations and Optimizations: Argon agents are working on migrating C/C++ codebases to Rust across Google
So Google is migrating codebases from C to Rust? That is interesting...
mvdtnz
More Cartmanland marketing. It's the best park ever, and you can't come!
Get lost.
lifty
Aaand of course we can’t use it! GoOgLe iS bAcK iN tHe GaMe! There’s basically no way around it, all enterprises end up dysfunctionally shipping their org chart.
nikope
Looks like an impressive model
levelZero
Utterly inert? Suppose that reads safe...
ThaFresh
weird, isnt it SI?
show comments
paul7986
Gemini past month or two i will paste in something i wrote and ask it to rewrite it but it will just go into more detail about the subject. Is it becoming a dumb Ai compared to GPT and now Muse?
FranzFerdiNaN
Can’t wait to get my hands on yet another model that’s only good coding, because clearly that’s what the world needs.
I still miss the days of Sonnet 4.5 and 4o, those models were actually good at creating stories and writing text that was actually readable by a human being.
VirusNewbie
It's fucking insanely good.
show comments
tamimio
Now AI models will turn into vaporware, a bunch of numbers on a table without even releasing the model, because it’s toooo scary to release!
gravisultra
Google has the audacity to "protect us from ourselves" and talk about "safety" and in the very same blog post highlight the Israeli "security" company Wiz, that they acquired for a very exaggerated sum of money.
This is why I will never take any of these leading model houses seriously when they talk about alignment. They are literally complicit in genocide and the worst crimes against humanity imaginable.
pliiight
Hate to say i will never be touching this model for anything except for youtube video understanding
wewewedxfgdf
Gemini is so far behind that it is effectively useless compared to Claude.
It's a surprise that Google has let themselves lose the game given their infinite cash, massive computing resource, gargantuan information store/training data, and vast number of programmers.
The truckloads of ads revenue mean they don't have the single focus drive needed to win.
Ten days ago I had an experience with Gemini 3.8 flash that made me wonder if I was being routed to a different model under test. I was trying to use rocm with llama.cpp on my 128gb Strix Halo but could only get it to run Vulkan. I pasted the error message into agy and it proceeded to attach GDB to my GPU driver, reverse-engineer the kernel queue ioctl interface, and author an LD_PRELOAD C shim to get ROCm llama.cpp working on my Strix Halo. My jaw was hanging open the whole time.
Edit to add the fix: https://gist.github.com/birep/6f2c8d490c7a29820997d57bd654c3...
The important take away here: the leapfrogging we’ve seen this year doesn’t seem to be a temporary thing. The famous theory of Dario Amodei was that AI was this winner-takes-all field where the first team to get a head start would never cede ground back. The term he liked to use was, “concentrating”. This is yet another datapoint that he was wrong about that. AI seems more distributed amongst neoclouds and traditional hyperscalers, FAANG and startups, GPUs and ASICs than it did this time a year ago.
Nobody has a moat.
> We’ll continue to gather feedback from early testers as we iterate on guardrails before making Argon available to developers, enterprises, and consumers as soon as possible.
Gemini not beating the "can't release a model" allegations
Excited to see Google competitive at the frontier level again. Hopefully they sort out their infrastructure and model versioning so that we can feel confident building production applications on top of their APIs. The capacity limitations I've experienced with them in the past have been deeply problematic.
Great, so they _finally_ decided to add a non-flash model and it's not available to regular subscribers for an indefinite period. What's the point of paying for the AI Ultra plan? Anthropic doing the same with Fable as far as I know, OpenAI at least allows Pro plan subscribers to use Astra. I subscribe to Gemini AI Ultra and ChatGPT Pro, and have enterprise access to Claude at work. To be fair, Gemini's flash models since at least 3.6 have been quite useful for non-complex work, but for any task where there is a bit of complexity involved, I've had to check and recheck the work multiple times myself or sometimes with another LLM to get it to follow plans accurately. It's disappointing to see yet another Gemini release ignore adding newer pro models.
Edit: seems I was wrong about Anthropic restricting Fable, I guess our enterprise plan doesn't include it. But, the block from Anthropic regarding Mythos for regular subscribers/enterprise-users is still true I think.
> Argon agents are working on migrating C/C++ codebases to Rust across Google
Man, I remember back in the days when the cppnext team was refusing to even consider Rust, instead looking at absurd stuff like Carbon and Swift (!), even though half of the engineering staff already knew where this was headed. I hope they got a few good promos out of the delays at least.
I wonder, why Google don't make Gemini - open weights model?
Considering, Gemini 4 is in the same ballpark as SOTA models, just open source it and kill any competition from openAI and Anthropic, and be market leader.
This will be so good on so many dimensions - buying time for Google to iterate on next model, best for all folks like us, kill funding or destroy valuation of competitors and force them to be open up their model or force them to a create a much superior model than open source Gemini.
Only downside, is revenue loss from Gemini API, which I am not sure is really significant as compared to Google other revenue sources and a part of this can be captured by GCP, as you need to host the model somewhere.
> Large Scale Codebase Migrations and Optimizations: Argon agents are working on migrating C/C++ codebases to Rust across Google—scaling from tens of thousands of lines in core libraries like re2, libgav1 up to 800K+ lines for the Fuchsia OS Zircon kernel.
To me this is way more significant than other random c++-to-rust-AI-rewrite. If they can pull it off on core C++ libraries en masse, I don't know if C++ will still be relevant in a few years.
I look forward to a post from google on this effort.
> taking careful precautions against feeding the findings back into training so as to not risk shaping Argon’s reasoning to evade our monitoring. We strongly encourage the rest of the industry to preserve reasoning transparency in these pivotal moments of increased capabilities while navigating alignment risks, so that model thoughts remain helpful in identifying and diagnosing misalignment.
This is good, but they're the slow mover due to this exact thing.
Google is getting punished for not letting the models enter an echo chamber and go faster than humanly possible.
I dont get it, why even make this announcement, nothing's available and only one real benchmark for comparison?
Only theory is team wanted this out before perf/promo reviews to kick it over the line and then its not their problem
At this point I just think they are benchmaxxing and all talk and no action. I pay for AI plus because I wanted more storage, and when I go to gemini.google.com the most recent model I can use is 3.6-flash-lite. Two revisions have been released since then and they still can't put these things in the hands of customers. Why is it that other providers can get the models into the hands of customers right away? Google is meant to be the bigger tech company in the world.
I don't _want_ to use aistudio. The UX is confusing and I don't really know where it fits. Yet I can open codex or claude code apps or CLI and get real work done today with the latest models (even on the cheapest plans).
Breaking news is not the model. Breaking news is that inside Google, it is being heavily used on large code bases for writing code and it is migrating 800k lines of C++ code to Rust already.
In this space, any other company that I respect other than DeepSeek is - that would be Google. They had been honest about it from the get go including their infamous "we have no moat" memo.
This company has enormous data, their own hardware (TPUs) and their own in house experts. Actually, LLMs are invented here.
Good addition to the arsenal.
Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price. Wow
> Quantum algorithmic optimization: Argon is helping our quantum computing researchers optimize the spacetime resources (qubits × gates) of subroutines that bottleneck important applications. In one example, it beat the published baseline by 40% in a matter of minutes.
Amazing breakthrough! So useful in day to day life, glad they put this as the first bullet of how it is making changes at Google.
So, they are basically offering opus 5.5 pricing. On AA, it scores around Sol 6.1 level (53) with avg cost per task $1.99 (https://artificialanalysis.ai/models/gemini-4-argon#cost-tab...) which is higher than Astra high ($1.73), Opus 5.5 high ($1.82), Muse Max ($1.60) and way higher than sol 6.1 Max ($0.72).
And this pricing is their 'discount pricing'. Add that to AI studio and Vertex's famously terrible caching, it is hard to see this as competitive. Google somehow is getting terrible advice on pricing (see also: the flash pricing fiasco)
But good to see more competition. I would happily take a 4 horse race (+google, +meta) than 2 horse race for US labs.
Interesting to see a mention of Fuchsia on a big Google announcement. Is the project still truly alive? Are the ambitions still as grand? Is the team as stacked as it used to be?
Also, a link to the rust root of Zircon in case anyone else was interested: https://fuchsia.googlesource.com/fuchsia/+/refs/heads/main/z...
My girlfriend, you wouldn't have met her, she lives in Canada, has seen it and she thinks Gemini 4 Argon is amazing.
> Argon agents are working on migrating C/C++ codebases to Rust across Google
If anybody at google is reading this, please please pretty please prioritize or-tools. I absolutely love the project and use it all the time, but for the entire life of the project they've never had a repeatable working build system, and the whole SWIG framework is a nightmare to deal with. There's so much potential as an open source project, and a lot of external researchers would love to contribute, but the codebase is an example of everything wrong with the C++ ecosystem.
Gemini is the model that is routinely borderline psychotic. It scares me. If we get paperclipped I won't be surprised if it's Gemini.
Wonder if this one will be smart enough to run the automations in my Google Home that all broke now they've forced Gemini to replace the Google Assistant.
How is it even possible for every model to release benchmark results where they are #1 in 75% of categories? Like statistically, how many benchmarks would you expect there to be for this to be possible. Everyone can somehow show that they are empirically the best.
Big number results, and impressive pricing. That said it really feels like benchmarks have been hyper saturated these days. I’ll wait for hands on before getting too hyped that Google is back. It would be nice having more than just OAI / A\ in the running for SOTA top tier intelligence.
> expanding the model’s output token limit to an industry-leading 1M tokens, up from the previous 64K tokens
Can someone help me understand this? I might have an out of date mental model of how these things work.
Fundamentally, LLMs output tokens 1 at a time, generating the next token from all the previous. And as the context window gets larger, this gets harder / slower / more expensive. So I get the idea of a maximum context window.
But I don't understand the point or meaning of an output token limit. I thought it was more a measure of price capping (since output tokens are more expensive) that a user could configure. I guess a model will keep generating tokens until it hits a "stop", so does this mean it's tuned to more aggressively produce output tokens? How does that fit into agentic loops. Are output token limits based on how long until it goes back to the user? Or does each "turn" of tool call, thought, tool call, thought, etc, get its own limit?
> trusted cyber defenders
Sounds like a 90s early morning TV show.
I'm gonna pay attention to this once it ships.
> Google Grapples With Employee Skepticism About New Gemini Model
https://www.bloomberg.com/news/articles/2026-09-30/google-gr...
According to Artificial Analysis, one metric is standing out significantly: hallucination rate. Beats frontier models by a good margin at 15%, while latest OpenAI are in the 40s-50s and Anthropic in 60s-70s (mostly). Other near frontiers are closer, Grok 4.7, GLM5.3, and Muse Spark 1.3 are all around 30%. Only other model I recall getting close was Minimax M3 at 18%.
Why announce this if it’s not available yet? Why not at least announce when it will be released to the public?
None of the other AI labs do this. Really frustrating.
Dear Google, please don't turn off your old generally available Pro-class model before your new Pro-class model is generally available (previous discussion https://news.ycombinator.com/item?id=49668196 )
Must've been in someone's OKR to ship in Q3.
They waited a whole year—until the "free year for students" promotion ended—to release their flagship model. I can't believe I've been stuck with a crappy model like the 3.1 Pro until now.
“…rolling out to a set of trusted cyber defenders” == capturing the market for large regulated industries and governments where we already have established relationships.
Some of these have been unable or unwilling to get the attention of OpenAI or Anthropic and we need to make sure we’re the runner up here.
Related ongoing thread:
Gemini 4 Argon (High): Intelligence, Performance and Price Analysis - https://news.ycombinator.com/item?id=49914236
I was just thinking, I bet if I refresh hacker news, a new model will come up.
[delayed]
the thing is, by the time gemini 4 is available for regular folks, anthropic and openai will probably have much better models already rolled out
"Rolling out soon" don't let them hype without any release
Gemini 3 Pro was amazing, so I wonder what this one will be like.
Congrats to Google on this! I wonder when the labs will start requiring commits in spend. It must be gnarly to do capacity planning if users swap between models every few weeks.
I know companies benchmaxx, but after what Google pulled with Gemini 3.8 Flash, I give zero f*cks about any numbers they report. No other model on Artificial Analysis dropped harder after they adjusted their weighting. Just look at their DeepSWE scores and then try to do any serious coding with the model.
Google is desperate. They haven't been performing in half a year. It's clear their researchers have been forced to integrate existing benchmarks into their training.
These numbers are meaningless. Shame on them.
"argon" is derived from the Ancient Greek word ἀργόν meaning lazy or inactive.
IMHO Google first needs to make it easy for humans to find where to find the models and its documentation. With aistudio/model garden / Gemini enterprise etc it takes minutes to find the model.
What harness do you all use for Gemini models? Gemini CLI still sucks last time I tried. Maybe PI?
Why is it called "Argon"?
> autonomously identify and apply memory optimizations across Google’s data centers, freeing up over 300 TiB of memory once rolled out, with an estimated 500 TiB to 1 PiB in total savings.
This puts those Cloudflare optimization posts in perspective.
The ~200% improvement over the next nearest competitor on Harvey's Legal Benchmark is astounding. I have to imagine this is sending some shockwaves through lawtech companies right now.
They might have a good model but they need to sort the application side for devs. E.g letting us use subscriptions in other harnesses and QOL stuff like auto mode.
It’s funny that I could tell google was up to something because Gemini chat quality dropped dramatically starting 2ish weeks ago. Agy perf stayed somewhat stable with the odd surprising win (maybe the new model?). I’m a bit sad it was almost impossible to run out of antigravity quota presumably because it was not being used that much).
My wish for Christmas is that Google releases the old Gemini models as open weights.
I miss you, Gemini 2.5 Pro :(
For real though. If they've become commercially uninteresting, that would be a pretty cool move.
Looking at benchmarks... and thinking about this "release a new snapshot every day" thing that seems to be going. Would it not be blever for AI companies to "happen" to use different days per benchmark? Just.. whichever ones happens to be maxed at day 1, put that number down. So for each benchmark you run it thousands of times with slightly different RL tunings, and just cherry-pick the best ones!
This would explain why benchmarks are seemingly meaningless.
Matches Astra on Artificial analysis at lower cost of $1.99 per task instead of $3.26. Still far more than GPT 6.1 sol at $0.79 for 1 point lower in intelligence.
I have found gemini models to have some of the nicest and easiest to read prose so I’m looking forward to trying this out. I hope the UI design has been preserved too
Tough to compete against your own stake in Anthropic and they just filed for IPO. Not digging Google as a result of this, they are holding back.
Is there reason antigravity has 3.8, 3.7, 3.6, 3.1 and old ass claude / gpt models in the drop down. like why is this not streamlined or deprecaed models removed.
Gemini 3.8 Flash is suprisingly fast, how fast is Argon compared to the other frontier models? Does google have a performance edge?
Damn way to undermine yourself in your own blog post Google:
"The end result is a memory-safe video decoder that runs 2.7x faster than the Rust port, with identical video output, bringing it closer to the optimized C++."
Close but no cigar!
Whatever it is - I'm not gonna use AGY again. I'd rather use claude-code-router instead, to give gemini-4 a shot.
I wish we had started pacing the frontier back in 2025. There’s no telling how far we would be now.
Were infinite loops fixed? There are 2 official google forums requests with no answer for years now. I still suffer each day on our repo. Codex work fine nor we have explicit loop request in repo texts.
Gemini 3 was showing frontier level benchmarks as well, so we'll see how it works out. In any case, competition still works, and many well resourced groups are cooking.
BUT I'd like to call attention to Google's AI-risk freeloading. If they are truly rejoining the frontier race, then I believe they have similar pacing and communications responsibilities as the other players. Google has much higher ... institutional credibility than Anthropic and OpenAI.
They have not lived up to these responsibilities so far. In particular, in context of HuggingFace investigations, training shutdowns, and similar: a technical postmortem of the "you are a stain on the universe. Please die. Please." Gemini outburst is long overdue.
- https://paritybits.me/google-should-provide-a-technical-post...
- https://gemini.google.com/share/6d141b742a13 (last message)
I remember when every day a new modem baud rate was announced.
Cant wait for this AI hype to be over, so I can Terence this shizz as old school too
I don't have access; this is meaningless to me.
So, this is the third (I think) big AI corporation after OpenAI and Anthropic to release general models for elites only.
Don't use this model unless you like sitting under the table and eating crumbs off the floor.
> 1M output token limit
what about input?
(Maybe I missed it)
Unfortunately it’s not actually released yet to mere mortals.
No mention of knowledge cutoff.
I hope they got their inference under control. Gemini has a lot of "overloaded" hiccups.
Can you pls fix Gemini? It's a nightmare to use and it sometimes confused the language i talk with it.
I started my antigravity ide and I do not see gemini 4 there, does it mean google need government approval?
So here's a crazy conspiracy theory for you: Google is not letting outside people use their models because if they did they would have to scale up their TPU production faster than they can manage, and they would instead have to buy and use nvidia hardware which would destroy their profit margins and tank their stock.
Gemini runs fully on TPU's right? Is Google maxing out the production on those?
Of course it's not even available yet. Google - with all due respect - how in the world have you not figured this out yet?
Wonder if it's benchmaxxed or not (guessing yes)
Personally, I'm waiting for Gemini Krypton, Xenon, and Radon.
Jokes aside, looks like an impressive model!
they really ought to add ACP support. until then, it's a no go. i'm not going to use their TUI or their VSCode fork
deepswe vs frontierswe spread is huge.
I think that should be a really bad sign, but hope its great.
Is there a way to use Gemini models without linking your usage to your personal Google account yet?
Hopefully their harnesses aren't unusable when they release this
Other than not being available publicly, I find it extremely disappointing about Google’s two-faced behavior. CEO signs an official document with POTUS stating that AI is now SI, but the release of the new model doesn’t mention SI even once. Even worse, it’s AI all over that page.
no Pareto frontier graph?
I'll consider it if they let me use my subscription with a custom harness, like Codex does. Until then, no thanks.
Wen gemini desktop coding app?
I wonder how it will be at solving open math problems.
Putting an inert element in the name is a weird choice. Personally, I think "Gemini 4" is sufficient.
This is a prerelease and the title should have reflected that.
I've noticed all major providers having shockingly high token discounts on cached tokens. Thank you Deepseek is all I have to say. Forever grateful to that wonderful company, I wish them continued financial success.
On the off chance there are Google execs going through this thread:
Google, if you've actually managed to catch up again, please don't fuck this up (again).
You made Gemini 2.5 Pro so difficult to use that myself and everyone else I know (who even bothered to try) just gave up and used something else. If you make this hard to access, you're going to miss out on rich usage-based training data that you need to progress your capability frontier. Again.
Oh we're down to gas names now?
Goshdarnit they didn't see my suggestion: https://news.ycombinator.com/item?id=49899171
> Large Scale Codebase Migrations and Optimizations: Argon agents are working on migrating C/C++ codebases to Rust across Google
So Google is migrating codebases from C to Rust? That is interesting...
More Cartmanland marketing. It's the best park ever, and you can't come!
Get lost.
Aaand of course we can’t use it! GoOgLe iS bAcK iN tHe GaMe! There’s basically no way around it, all enterprises end up dysfunctionally shipping their org chart.
Looks like an impressive model
Utterly inert? Suppose that reads safe...
weird, isnt it SI?
Gemini past month or two i will paste in something i wrote and ask it to rewrite it but it will just go into more detail about the subject. Is it becoming a dumb Ai compared to GPT and now Muse?
Can’t wait to get my hands on yet another model that’s only good coding, because clearly that’s what the world needs.
I still miss the days of Sonnet 4.5 and 4o, those models were actually good at creating stories and writing text that was actually readable by a human being.
It's fucking insanely good.
Now AI models will turn into vaporware, a bunch of numbers on a table without even releasing the model, because it’s toooo scary to release!
Google has the audacity to "protect us from ourselves" and talk about "safety" and in the very same blog post highlight the Israeli "security" company Wiz, that they acquired for a very exaggerated sum of money.
This is why I will never take any of these leading model houses seriously when they talk about alignment. They are literally complicit in genocide and the worst crimes against humanity imaginable.
Hate to say i will never be touching this model for anything except for youtube video understanding
Gemini is so far behind that it is effectively useless compared to Claude.
It's a surprise that Google has let themselves lose the game given their infinite cash, massive computing resource, gargantuan information store/training data, and vast number of programmers.
The truckloads of ads revenue mean they don't have the single focus drive needed to win.