I wonder how big the Pro model is that Google is using behind the scenes to train these smaller ones.
Going on baseless speculation, the lack of accompanying pro models with these flash releases either means: 1) the model is too big to be economical, 2) google doesn't have the compute to serve the big model, 3) their big model has too many alignment issues to serve to the public.
edit: looks like benchmarks are up on https://artificialanalysis.ai/models/gemini-3-6-flash. It's solidly middle-of-pack. However, if you want to be most fair to flash, look at the intelligence vs time per task and intelligence vs outputspeed benchmarks. This is a very fast model.
edit 2: I use antigravity from time to time and in my experience, 3.5 flash is an underrated model, so long as you know what it's good for. It's very good at frontend (much better than gpt 5.5) and it's fast, so it's a great tool for iteration. I expect 3.6 to be no different.
show comments
stonewhite
Google somehow managed to snatch defeat from the jaws of success with their AI products.
They literally forced me and my company out of Antigravity by phasing out AI Ultra subscription without any proper product follow-up. Antigravity IDE cannot even have poweruser subscriptions now from Google Workspace an Gemini Enterprise Agent Platform cannot be attached to Antigravity IDE.
Gemini Enterprise Agent Platform has an incredibly abysmal setup process, and if I want to limit spending per-user I have to create projects per user. The fact that you cannot activate Anthropic models on it if the billing still has free credits is almost a joke.
I was a big proponent of Google and Gemini, but they left us reeling with their abrupt product decisions. Forced us to buy $200 subscriptions directly from Anthropic/OpenAI.
show comments
brap
From my experience, this thing is crazy fast.
Spawn 10 on the same problem and have them debate to reach a consensus, you’ll get Fable-like results but 100x faster.
m_w_
It's a bit disheartening to see no comparison to other models here - and I'm not sure this pushes the curve anywhere. 3.6 flash is more expensive than GLM 5.2 - but seemingly worse, although this post is really light (lite?) on details.
It seemed for a time that Google had finally gotten the ball rolling, but I'm doubting that more and more as time passes. We'll see what happens with 3.5 pro I suppose.
show comments
simonw
Pelicans for 3.6 Flash and 3.5 Flash-Lite (Cyber isn't available to me through the API yet.)
> Beyond today’s releases, Gemini 3.5 Pro is currently testing with partners and we plan to make it broadly available as soon as it’s ready.
> We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress.
show comments
jgbuddy
It is both less intelligent and more expensive than GLM-5.2, while being closed weight.
show comments
swe_dima
It's scary relying on Google's models.
I have a very price sensitive workload that used to run on flash 2.5 lite - it's deprecated now.
The replacement 3.1 flash lite is a lot more expensive, but now also has a sunset date.
3.5 flash lite is even more expensive.
So the price is rising and you have no choice but to keep paying more and more.
show comments
b473a
No word about updating Jules, which is still stuck on 3.1 Pro. I get that it's probably niche but I've really appreciated basically being able to give directions to Jules on my phone, then reviewing and merging a GitHub PR fifteen minutes later. It's been great for getting some progress in on a few personal projects during my commute when I can't exactly pull out my laptop.
Anyone have any good alternatives?
show comments
arjie
Their naming scheme is confusing. Branding has never been Google's strong suit and their marketing copy is pretty bottom-of-the-barrel[0]. Anthropic has a pretty clear set of models but Gemini decided to rebrand their Flash as Flash Lite (and presumably the future will see a Flash Lite Mini, a Flash Lite Mini Nano and a Flash Lite Mini Nano 3B) which confuses the pricing to high hell.
This plus the Vertex, AI Studio, Gemini, Antigravity. It's honestly too confusing to use. I need to use Gemini just to decide on which platform and which model to consider.
0: Famous Kurian Tweet: "We're announcing Duet AI for Google Workspace will now be Gemini for Google Workspace. Consumers and organizations of all sizes can access Gemini across the Workspace apps they know and love. We're introducing a new offering called Gemini Business, which lets organizations use generative AI in Workspace at a lower price point than Gemini Enterprise, which replaces Duet AI for Workspace Enterprise."
lilytweed
Really, what's up with Gemini still not supporting connectors/MCPs/plugins/whatever-they're-called-this-month on web? It makes it a non-starter for any kind of serious use.
mchusma
Wow, Laguna S 2.1 (released today) just destroys Flash-Lite underly and completely. What a weak and embarrasing release from Google.
WarmWash
The mention of an "ambitious" gemini 4 pre-train signals to me that 3.5 pro is probably a lost cause.
That being said, it seems that Gemini is still the best image analysis model, so hopefully 3.6 flash builds on this even more.
jdthedisciple
Bottom line it looks about on equal footing with GLM 5.2 in terms of both overall intelligence and cost per task, while being significantly faster (in fact it is the fastest model on artificial analysis as of rn [0])
I have a side business selling custom fingerprint jewelry and I use gemini nano banana to clean up customer submitted fingerprint images. This was a step I used to do by hand at 10 - 15 minutes per image and nano banana is the first model that is able to do the task (it is astonishingly good at it). I can't wait to see what the next nano banana can do, hopefully its released soon.
show comments
velominati
Wow - Google does not even bother to show benchmarks of these models compared to the frontier and Chinese labs - only against previous versions. I'm not surprised. Having worked there for years it was amazing just how inwardly looking the company is.
waldrews
3.5 Flash-Lite seems available in US region, as was 3.5 Flash; but 3.6 Flash looks Global only so far when pinging. If Google employees are watching, will this issue go away?
spyckie2
Google seems to have anorexia when it comes to model intelligence. They have an internal hard constraint on price per token it seems, and they are trying to squeeze out intelligence with limited compute.
I wonder if there is something with their TPU cycles that makes them want to postpone training a new model. My guess is that they have been on the same base model for 6 months and they may have waited for the next gen TPUs to train Gemini 4, which greatly limits how much intelligence they can increase and forces them to do cost efficiency increases.
show comments
dankai
Unfortunately says more about how competitive 3.5 pro would be today at the frontier if they forgo it for 3.6 flash.
Alifatisk
In other good news "the model has been trained to minimize refusals for beneficial uses.".
Otherwise, this news feels like a tiny incremental improvement on Gemini Flash series to make it more efficient with token usage, subagent and cost. Nothing big.
Regarding their benchmark scores on CyberGym, I wonder why they didn't compare their 3.5 Flash Cyber model with Fable 5. I mean they included Mythos and GPT-Cyber, so why not Fable 5 too?
They also mentioned Gemini 3.5 Pro is in testing and its about to become available very soon. Another thing maybe worth discussing is the announcement of pre-training Gemini 4. Sadly, not much technical details to discuss on. Many comments in here seem to mostly be about how Google is behind the others, but honestly, is it really worth the investment to be #1 in Artifical Analysis every week?
hmokiguess
Spent half an hour just now benchmarking it against my current 3.5 Flash pipeline excited only see it regressed slightly (0.1% - 0.2% at most, for feature extraction work)
Seems like this is mostly a cost play by Google, hoping this doesn't bring 3.5 Flash capabilities to an end of life, and that 3.6 catches up or gets better.
u1hcw9nx
Google has not changed. Following two facts are like tautologies by now.
1. Their AI efforts are very fundamental research oriented. They are really good at it.
2. Their productization sucks. The end products gets little attention compared to competition. It can be canceled at any time. You should never build anything around Google only APIs, AI or not.
youssefarizk
3.5-lite is the real showpiece here; agentic models of this size are a huge value-add for 90% of knowledge work agent tasks
ianberdin
Pelican svg and a near-perfect 3D MacBook at max effort for $0.16, about a fifth of Fable's price.
Fable 5 still wins on detail with no visible errors, but it's close. And this isn't a memorized pelican;
I often use Gemini free web chat because it's generally quite good at web search-related questions (apparently it has direct token-level access to the Google Search index) but I noticed in the last two weeks output quality of 3.5 Flash seriously degraded. Maybe they were switching over systems.
show comments
ConfusedDog
Why would 3.6 flash perform a little worse than 3.5 flash on Artificial Analysis Coding Index...
Tons of guardrails, lazy model, super confusing plans, expensive 3.5/3.6 flash and lite and 3.5 pro MiA?
Rough patch for google ai
sega_sai
I have just tried to switch to 3.6 instead of 3.5 in antigravity and it seems to constantly spit "critical instruction: STOP CALLING TOOLS NOW. YOU MUST WAIT FOR WAKEUP. ". I think I will switch back to 3.5
parasti
Kind of excited about this. 3.5 Flash on Antigravity has surprised me recently on a hobby project. When given opportunity to plan, it can deliver on tasks that would take me a while on my own and generates responses at blazing speeds - compared to what I'm used to at work with Opus 4.8 (granted I don't use Opus 4.8 on my hobby projects so just anecdotal). While with Gemini CLI I would just watch it run in circles and run out of 5h allowance before anything useful is produced (or even approached).
singingtoday
I'm more excited for 3.5 pro. Gemini has fallen behind in some areas, but is still one of the best multimodal models.
Has anybody found any models better at image or audio analysis?
show comments
nsbk
It is 17% more token-efficient than 3.5 and performs significantly better in coding and tool usage benchmarks.
It is also cheaper than 3.5:
> This enhanced efficiency is also combined with a lower price than 3.5 Flash. At $1.50/1M input tokens and $7.50/1M output tokens, 3.6 Flash reduces the overall cost per agentic task, making agents more cost-effective to build and run.
thevinter
I struggle to see any value in this when DeepSeek is still a thing.
show comments
mythz
Always happy to see new Gemini releases as IMO Antigravity Pro 16.67/mo plan (Annual) is still the best plan available and have been pretty happy with Antigravity IDE.
If it wasn't for Gemini/Antigravity I'd have to go with a Max Claude plan, as it stands now I can get by with just a Claude Pro plan to get Opus when I need it, whilst using Antigravity as my day-to-day workhorse.
Unfortunately Gemini Flash became too expensive to use as a general purpose model (i.e. for AI features in Apps), luckily there are plenty of cheaper Chinese models to fill that gap now.
show comments
sagex
Don't know why are they even pursuing Gemini. Just download the Kimi, call it Kimini and serve it on your GPU. Maybe then train next architecture based on this!
parsimo2010
Feels like they released this to ride the wave of press of GPT-5.6, Kimi K3, and Qwen 3.8. Doesn't feel like Google has much substance with this post except a bump in version and tweaked their pricing.
JeremyHerrman
Gemini 2.5 Flash-Lite has been my go to for cheap document processing at scale (especially with 50% off batch mode), but they are really boiling the frog with pricing increases with each version:
gemini-2.5-flash-lite: $0.10 input / $0.40 output
gemini-3.1-flash-lite: $0.25 input / $1.50 output
gemini-3.5-flash-lite: $0.30 input / $2.50 output (a 6.25x increase over 2.5!)
Now watch them deprecate Gemini 2.5 Flash-Lite in the coming months...
show comments
lambda
3.6 Flash scores exactly the same as 3.5 Flash on the Artificial Analysis index. Better on some tasks, worse on others. Mostly within what I'd consider the noise window. Looks pretty much indistinguishable from 3.5 Flash, at least on these benchmarks: https://artificialanalysis.ai/models/gemini-3-6-flash
goldenarm
LLM reception is truly extreme, even worse than AAA game releases.
Ever frontier lab lived it at least once : missing the frontier by a few months triggers extremly negative reactions, then you take back the lead for 2 weeks, and the hype cycle repeats.
summerlight
Looks like 3.6 Flash is the first model with their newest pretraining run (cutoff date is 2026/03), long after 2.5 series.
vinhnx
For anyone wanting a faster overview: I ran the Gemini 3.6 Flash and 3.5 series release notes through NotebookLM and generated a short video summary. Link: https://www.youtube.com/watch?v=SUFBhvQ2tY4
thebigspacefuck
IMO Gemini has the best free tier models/app for everyday use. Muse-Spark is perhaps just slightly better, but has none of the connectivity to my GApps (for things like “create a recipe in my Google Docs from this image”).
Plus they are probably running these things on every Google search so saving tokens is a huge win for them.
show comments
MILP
I'm a big fan of the Flash-Lite models. They're exceedingly fast and deliver great outputs for high volume use cases where you need to process requests at scale. Can't wait to try the newer version.
kilroy123
I deeply wish Google would focus on models like Gemma. Small, powerful, open-weight models you can run on phones or regular computer hardware.
show comments
1saadcodes
Nice to see that it's cheaper than 3.5
ComputerGuru
So 3.6 Flash is a somewhat of an admission that Google miscalculated by charging 3-5x for 3.5 Flash what it did for 3.0 Flash (3x input and output costs plus large token inefficiency changes) despite only modest improvements?
3.5 Flash Lite is only a hair cheaper than 3.0 Flash, but I think 3.0 Flash is a massively more capable model?
Havoc
Flash Lite: 0.3/m and 2.5/m
Deepseek Pro: 0.435/m 0.87/m
That's wildly ambitious pricing by Google. You can maybe get away with spicy pricing at the SOTA edge but at the lower tiers everything is a lot more price sensitive.
show comments
dvduval
It does seem like their releases are getting closer together. I get the feeling they realized they were trying to roll out to their entire ecosystem and now they’re focusing more just directly on the AI model itself. I think give it a little time and they’ll start to be one of the competitors too.
xnx
Proof-of-life release while they figure out how to have a competitive frontier model release. My hunch is they pushed too far in the "omni" model direction, that they made something so ungainly, it wasn't as good for normal tasks.
dumberquestions
"..and in some benchmarks like DeepSWE by Datacurve, we observe up to 65%, all at a lower cost per output token."
"3.6 Flash delivers higher precision with fewer unwanted code edits and reduced execution loops, as seen in DeepSWE (49% vs. 37%)"
So which one is it? 65% or 49%?
show comments
XCSme
I was expecting 3.6 Pro. It's been so long since the last Pro model...
show comments
pietz
Are they comparing 3.6 Flash to 5.6 Luna and losing? That's ruff.
show comments
Gecko4072
I read this as a soft let down to not expect too much from 3.5 Pro.
> We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress.
yanis_t
The benchmarks are not particularly impressive. I suppose they needed to release something since the long pause. But not clear why would I use it now.
TheAtomic
I have liked using their consumer products but they don't make it easy, that's for sure.
spstoyanov
Glad to see the price is going down but it's still too high for a "fast" model
catigula
"We made 3.6/4 Pro, but it sucks, so this is the distilled model" vibes.
show comments
AussieWog93
A lot of disappointment here in the comments, but models like these aren't meant to compete with the likes of Fable or GPT 5.6.
I use 3.1 Flash Lite regularly to classify listings on eCommerce websites. It's great for this task - fast, cheap and accurate.
In fact, it was the single best model we tried in terms of the speed vs accuracy vs price tradeoffs - including the Chinese models.
Of course, 3.5 Flash was more accurate but the 5x cost increase couldn't be justified.
3.5 Flash Lite sounds like it could be a strict upgrade for our use case, without a significant increase in costs or drop in speed.
It's not GPT-6 but it's not trying to be. It's a completely different tool and great at what it does.
mfkrause
Pretty underwhelming, as expected honestly. I don't want to know what morale is like at DeepMind right now.
show comments
metahost
So about the same “intelligence” as Muse Spark 1.1 but 2x faster and about 2x as expensive.
XCSme
tl;dr: 3.6 flash is a bit smarter than 3.5 flash, but also a bit more expensive.
My results [0] put Gemini 3.6 Flash at the top.
3.6 Flash high has same $1.5 input price as 3.5 Flash, but output is cheaper from $9.0 to $7.5.
Google said 3.6 Flash is more token efficient, but in my tests it's actually LESS token efficient[1] than 3.5 Flash, so despite the output price reduction, it still costs more.
I have no skin in this game and this comment will be gray in a few minutes BUT a friendly reminder that these types of threads are astroturfed heavily by competitor labs and any info should be taken with a massive grain of salt.
luciana1u
the real product is the naming confusion we made along the way. Gemini 3.6 Flash, 3.5 Flash-Lite, 3.5 Flash Cyber — at this point even the model cards need a model to explain them
vlad_recomply
Models are expensive and low performance. On top of that they make you jump through hoops to even use these models without being throttled even for the weaker models. The only reason we are using them is credits. As soon as credits run out we are switching immediately.
maxdo
quite a good model, the speed/price/quality ration is a new golden intersection for me, not sure if its as good as grok 4.5 but quite fast/capable model.
sreekanth850
Google is walking backwards, with such a pile of cash in pocket, i feel they are doomed.
gabriel-uribe
Haven't been excited for a Gemini release since December. Wild to see.
baalimago
Not good enough for high-end, not cheap enough to be for low-end. Next!
speak_plainly
It feels like AI is going to be the end of Google. The post-Schmidt company culture cannot produce consistent, consumer-friendly products that any sane person would want to use consistently.
Arshad-Talpur
never tried gemini for coding, but this news seems to be compelling, i would definitely give it a try
lukewarm707
"The model will be exclusively available to governments and trusted partners via CodeMender soon as part of a limited-access pilot program"
we are stealing plutocracy from the jaws of emancipation.
i don't want to live in a world where abundance is guarded and shared among politicians and cronies, whilst the rest are left to rot.
Why exactly are they announcing these completely milquetoast models ?
I'd be low-keying the release if anything, given how lame they are compared to their competition.
What am I missing?
lwansbrough
Plugged 3.5 Flash Lite into an existing agent harness that was previously using 3.1 Flash Lite and this shit just does not work. It's not following instructions and is not producing the correct tool calls.
QuesnayJr
I remember back when Gemini looked like it was the best model that this comment section was full of confident predictions that Google had "won" and that no one would ever catch up with them again. The most embarassing part is that I kinda believed them.
5701652400
if only DeepSeeek supported vision, would never use Gemini.
m4tthumphrey
I'm going to get downvoted/flagged but I feel like we need a new type of "Show HN/Tell HN" etc for "New AI Model Available".
Front page is tedious these days.
show comments
theplumber
I think it’s safe to say Google seems a bit out of the top AI competition now. The “cyber” stuff also starts to become laughable with open models providing the full power without the crap Anthropic, Google, OpenAI are trying to frontload on you(I.e you are not allowed to develop/review a login system, pay a special cyber operation team to do it for you). They really deserve to become irrelevant in the future of AI.
Google watches over the last few months a flat out assault on the Pareto curve from American and Chinese companies. Release after release pushing the boundaries of frontier intelligence and price/performance.
And the response from arguably the biggest AI research labs in the world by headcount is Flash 3.6.
What do you do when you are given essentially unlimited resources and still find yourself falling behind?
show comments
onlyrealcuzzo
Gemini 3.5 flash is already a pretty good model. But, unfortunately, the primary way you can interact with it for coding is through Antigravity - which is actively developer hostile.
It doesn't matter how good the model is if you're (mostly) forced to use it in Antigravity - which turns any model into crap.
Wake me up when Antigravity doesn't suck.
dakolli
I use 3.5 flash 10x more than any other model, despite have access to all of them. If I'm going to play a slot machine, I'd rather get the pain over with quickly.
lenerdenator
We're almost five years into the whole GenAI thing and we're still relying on these guys to spoonfeed us incremental updates.
It's time for them to start focusing on open-weight models and efficiency. Otherwise there's just a layer of marketing hype and "will it do this?" that has to be cut through for evaluation of each and every release cycle.
Models are getting easier and easier to create. The money, if there's any here, is in the harness the user interfaces with, and the data centers running them.
zuzululu
Google seems to be falling way behind the pack. antigravity cli is pure trash. gpt 3.5 pro is now behind and isn't released yet. GPT 6 and Fable 6 releasing next month. What the hell is going on over there ?
show comments
ece
Just switched to AI Plus from Pro, seems like I won't be missing much.
zb3
> we have taken an intentional approach to deploying 3.5 Flash Cyber. The model will be exclusively available to governments and trusted partners
Screw your government! US and Israeli governments should get the least access, but of course we all know they'll be the (only) ones to get full unfiltered access.
tiahura
3.5 Pro must really suck.
holistio
They are comparing against their own previous models instead of competitors. Not a great sign.
geooff_
At this point just put the Pareto in the bag bruh
accountrequired
whatever, dude. give gemma5
raffael_de
is it just me or is this one-upping each other every few days getting ridiculous secreting a whiff of desperation?
With all the naysayers on Gemini models I'm curious how many people actually use Gemini regularly?
For me, Gemini models are the most usable. Claude Opus and Mistral always try to turn queries into one-shot enormous commits, which just burns tokens, time and annoys me for something which is still wrong more often than not.
Gemini seems far better at listening to instructions and giving me what I actually want, on top of using far fewer tokens and wasting my time. Fable is the only model that's come close to Gemini Pro for me.
And as this is about Flash, it's exciting, I find Flash can usually get the right answer pretty quickly and without too much nonsense.
llmslave
I keep saying this and people dont believe me, but I have b2b saas systems with actual agents running around the clock, and the performance/stability of the flash model is higher than most other models.
Meaning, its predictable with tool calls, wont spin off a million tools/do weird behavior, its reasonable. Even sonnet in a real world decision making scenario is not reliable, or will reason so long its incredibly expensive.
The benchmarks arent catching all the value, and most people have never actually ran an ai agent in a real context that matters
show comments
fur-tea-laser
not a google fanboy by any stretch... though i've been thrilled with the flash line of models... i exclusively use it on high, and have found it to be a great fit for increasing productivity 10-fold while maintaining quality... sure it can't just go off and one-shot a bunch of work, but at the complexity level i tend to work at, neither can the frontier in a robust way that i can be confident in... sure i have to be in the loop more, but that helps keep me grounded and course-correct earlier before wasting tokens... and when you sufficiently spec out a coding/software problem, and i mean really document all of the critical nuance, it will successfully satisfy the constraints... the quality is rarely acceptable on first-pass, but it forces me to stay connected to the architecture more than i would be if using a frontier model... i've found this to be a happy middle-ground of productivity and awareness...
ewaewaewa
jew
npn
tested the models on aistudio. despite that the knowledge cut off is march 2026 it still knows nothing about 2025!
you can check by asking "list notable world events in 2025, only list unplanned" on aistudio. or you can ask for Charlie Kirk, it also does not know. I tried it multiple time to ensure that I didn't not get routed to older models!
> but google has search
irrelevant, without deeper knowledge about cutting edge technologies or latest libraries, all of it suggestions are crap. even you ask it to search it will still use outdated keyword thus only getting outdated information.
in other word, what a disaster!
kthinckley
Google desperately needs to make some leadership changes within their Gemini team now that they've been surpassed by 3-5 open weight models and risk loosing frontier status all together in the near future.
show comments
game_the0ry
At this point, I think google should consider becoming a hyper scaler for anthropic and open ai, and I predict that that is exactly what they do. The model is no longer the most valuable part of the stack.
I wonder how big the Pro model is that Google is using behind the scenes to train these smaller ones.
Going on baseless speculation, the lack of accompanying pro models with these flash releases either means: 1) the model is too big to be economical, 2) google doesn't have the compute to serve the big model, 3) their big model has too many alignment issues to serve to the public.
edit: looks like benchmarks are up on https://artificialanalysis.ai/models/gemini-3-6-flash. It's solidly middle-of-pack. However, if you want to be most fair to flash, look at the intelligence vs time per task and intelligence vs outputspeed benchmarks. This is a very fast model.
edit 2: I use antigravity from time to time and in my experience, 3.5 flash is an underrated model, so long as you know what it's good for. It's very good at frontend (much better than gpt 5.5) and it's fast, so it's a great tool for iteration. I expect 3.6 to be no different.
Google somehow managed to snatch defeat from the jaws of success with their AI products.
They literally forced me and my company out of Antigravity by phasing out AI Ultra subscription without any proper product follow-up. Antigravity IDE cannot even have poweruser subscriptions now from Google Workspace an Gemini Enterprise Agent Platform cannot be attached to Antigravity IDE.
Gemini Enterprise Agent Platform has an incredibly abysmal setup process, and if I want to limit spending per-user I have to create projects per user. The fact that you cannot activate Anthropic models on it if the billing still has free credits is almost a joke.
I was a big proponent of Google and Gemini, but they left us reeling with their abrupt product decisions. Forced us to buy $200 subscriptions directly from Anthropic/OpenAI.
From my experience, this thing is crazy fast.
Spawn 10 on the same problem and have them debate to reach a consensus, you’ll get Fable-like results but 100x faster.
It's a bit disheartening to see no comparison to other models here - and I'm not sure this pushes the curve anywhere. 3.6 flash is more expensive than GLM 5.2 - but seemingly worse, although this post is really light (lite?) on details.
It seemed for a time that Google had finally gotten the ball rolling, but I'm doubting that more and more as time passes. We'll see what happens with 3.5 pro I suppose.
Pelicans for 3.6 Flash and 3.5 Flash-Lite (Cyber isn't available to me through the API yet.)
https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
Pricing per million input/output tokens:
2.5 Flash: $0.3 / $2.5
3.0 Flash: $0.5 / $3
3.5 Flash: $1.5 / $9
3.6 Flash: $1.5 / $7.5
---
2.5 Flash-Lite: $0.1 / $0.4
3.1 Flash-Lite: $0.25 / $1.5
3.5 Flash-Lite: $0.3 / $2.5
A couple tidbits:
> Beyond today’s releases, Gemini 3.5 Pro is currently testing with partners and we plan to make it broadly available as soon as it’s ready.
> We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress.
It is both less intelligent and more expensive than GLM-5.2, while being closed weight.
It's scary relying on Google's models.
I have a very price sensitive workload that used to run on flash 2.5 lite - it's deprecated now.
The replacement 3.1 flash lite is a lot more expensive, but now also has a sunset date.
3.5 flash lite is even more expensive.
So the price is rising and you have no choice but to keep paying more and more.
No word about updating Jules, which is still stuck on 3.1 Pro. I get that it's probably niche but I've really appreciated basically being able to give directions to Jules on my phone, then reviewing and merging a GitHub PR fifteen minutes later. It's been great for getting some progress in on a few personal projects during my commute when I can't exactly pull out my laptop.
Anyone have any good alternatives?
Their naming scheme is confusing. Branding has never been Google's strong suit and their marketing copy is pretty bottom-of-the-barrel[0]. Anthropic has a pretty clear set of models but Gemini decided to rebrand their Flash as Flash Lite (and presumably the future will see a Flash Lite Mini, a Flash Lite Mini Nano and a Flash Lite Mini Nano 3B) which confuses the pricing to high hell.
This plus the Vertex, AI Studio, Gemini, Antigravity. It's honestly too confusing to use. I need to use Gemini just to decide on which platform and which model to consider.
0: Famous Kurian Tweet: "We're announcing Duet AI for Google Workspace will now be Gemini for Google Workspace. Consumers and organizations of all sizes can access Gemini across the Workspace apps they know and love. We're introducing a new offering called Gemini Business, which lets organizations use generative AI in Workspace at a lower price point than Gemini Enterprise, which replaces Duet AI for Workspace Enterprise."
Really, what's up with Gemini still not supporting connectors/MCPs/plugins/whatever-they're-called-this-month on web? It makes it a non-starter for any kind of serious use.
Wow, Laguna S 2.1 (released today) just destroys Flash-Lite underly and completely. What a weak and embarrasing release from Google.
The mention of an "ambitious" gemini 4 pre-train signals to me that 3.5 pro is probably a lost cause.
That being said, it seems that Gemini is still the best image analysis model, so hopefully 3.6 flash builds on this even more.
Bottom line it looks about on equal footing with GLM 5.2 in terms of both overall intelligence and cost per task, while being significantly faster (in fact it is the fastest model on artificial analysis as of rn [0])
[0] https://artificialanalysis.ai/models/gemini-3-6-flash
I have a side business selling custom fingerprint jewelry and I use gemini nano banana to clean up customer submitted fingerprint images. This was a step I used to do by hand at 10 - 15 minutes per image and nano banana is the first model that is able to do the task (it is astonishingly good at it). I can't wait to see what the next nano banana can do, hopefully its released soon.
Wow - Google does not even bother to show benchmarks of these models compared to the frontier and Chinese labs - only against previous versions. I'm not surprised. Having worked there for years it was amazing just how inwardly looking the company is.
3.5 Flash-Lite seems available in US region, as was 3.5 Flash; but 3.6 Flash looks Global only so far when pinging. If Google employees are watching, will this issue go away?
Google seems to have anorexia when it comes to model intelligence. They have an internal hard constraint on price per token it seems, and they are trying to squeeze out intelligence with limited compute.
I wonder if there is something with their TPU cycles that makes them want to postpone training a new model. My guess is that they have been on the same base model for 6 months and they may have waited for the next gen TPUs to train Gemini 4, which greatly limits how much intelligence they can increase and forces them to do cost efficiency increases.
Unfortunately says more about how competitive 3.5 pro would be today at the frontier if they forgo it for 3.6 flash.
In other good news "the model has been trained to minimize refusals for beneficial uses.".
Otherwise, this news feels like a tiny incremental improvement on Gemini Flash series to make it more efficient with token usage, subagent and cost. Nothing big.
Regarding their benchmark scores on CyberGym, I wonder why they didn't compare their 3.5 Flash Cyber model with Fable 5. I mean they included Mythos and GPT-Cyber, so why not Fable 5 too?
They also mentioned Gemini 3.5 Pro is in testing and its about to become available very soon. Another thing maybe worth discussing is the announcement of pre-training Gemini 4. Sadly, not much technical details to discuss on. Many comments in here seem to mostly be about how Google is behind the others, but honestly, is it really worth the investment to be #1 in Artifical Analysis every week?
Spent half an hour just now benchmarking it against my current 3.5 Flash pipeline excited only see it regressed slightly (0.1% - 0.2% at most, for feature extraction work)
Seems like this is mostly a cost play by Google, hoping this doesn't bring 3.5 Flash capabilities to an end of life, and that 3.6 catches up or gets better.
Google has not changed. Following two facts are like tautologies by now.
1. Their AI efforts are very fundamental research oriented. They are really good at it.
2. Their productization sucks. The end products gets little attention compared to competition. It can be canceled at any time. You should never build anything around Google only APIs, AI or not.
3.5-lite is the real showpiece here; agentic models of this size are a huge value-add for 90% of knowledge work agent tasks
Pelican svg and a near-perfect 3D MacBook at max effort for $0.16, about a fifth of Fable's price.
Fable 5 still wins on detail with no visible errors, but it's close. And this isn't a memorized pelican;
https://playcode.io/blog/macbook-svg-benchmark#gemini-3-6-fl...
I often use Gemini free web chat because it's generally quite good at web search-related questions (apparently it has direct token-level access to the Google Search index) but I noticed in the last two weeks output quality of 3.5 Flash seriously degraded. Maybe they were switching over systems.
Why would 3.6 flash perform a little worse than 3.5 flash on Artificial Analysis Coding Index...
https://artificialanalysis.ai/models/gemini-3-6-flash?intell...
Tons of guardrails, lazy model, super confusing plans, expensive 3.5/3.6 flash and lite and 3.5 pro MiA?
Rough patch for google ai
I have just tried to switch to 3.6 instead of 3.5 in antigravity and it seems to constantly spit "critical instruction: STOP CALLING TOOLS NOW. YOU MUST WAIT FOR WAKEUP. ". I think I will switch back to 3.5
Kind of excited about this. 3.5 Flash on Antigravity has surprised me recently on a hobby project. When given opportunity to plan, it can deliver on tasks that would take me a while on my own and generates responses at blazing speeds - compared to what I'm used to at work with Opus 4.8 (granted I don't use Opus 4.8 on my hobby projects so just anecdotal). While with Gemini CLI I would just watch it run in circles and run out of 5h allowance before anything useful is produced (or even approached).
I'm more excited for 3.5 pro. Gemini has fallen behind in some areas, but is still one of the best multimodal models.
Has anybody found any models better at image or audio analysis?
It is 17% more token-efficient than 3.5 and performs significantly better in coding and tool usage benchmarks.
It is also cheaper than 3.5:
> This enhanced efficiency is also combined with a lower price than 3.5 Flash. At $1.50/1M input tokens and $7.50/1M output tokens, 3.6 Flash reduces the overall cost per agentic task, making agents more cost-effective to build and run.
I struggle to see any value in this when DeepSeek is still a thing.
Always happy to see new Gemini releases as IMO Antigravity Pro 16.67/mo plan (Annual) is still the best plan available and have been pretty happy with Antigravity IDE.
If it wasn't for Gemini/Antigravity I'd have to go with a Max Claude plan, as it stands now I can get by with just a Claude Pro plan to get Opus when I need it, whilst using Antigravity as my day-to-day workhorse.
Unfortunately Gemini Flash became too expensive to use as a general purpose model (i.e. for AI features in Apps), luckily there are plenty of cheaper Chinese models to fill that gap now.
Don't know why are they even pursuing Gemini. Just download the Kimi, call it Kimini and serve it on your GPU. Maybe then train next architecture based on this!
Feels like they released this to ride the wave of press of GPT-5.6, Kimi K3, and Qwen 3.8. Doesn't feel like Google has much substance with this post except a bump in version and tweaked their pricing.
Gemini 2.5 Flash-Lite has been my go to for cheap document processing at scale (especially with 50% off batch mode), but they are really boiling the frog with pricing increases with each version:
gemini-2.5-flash-lite: $0.10 input / $0.40 output
gemini-3.1-flash-lite: $0.25 input / $1.50 output
gemini-3.5-flash-lite: $0.30 input / $2.50 output (a 6.25x increase over 2.5!)
Now watch them deprecate Gemini 2.5 Flash-Lite in the coming months...
3.6 Flash scores exactly the same as 3.5 Flash on the Artificial Analysis index. Better on some tasks, worse on others. Mostly within what I'd consider the noise window. Looks pretty much indistinguishable from 3.5 Flash, at least on these benchmarks: https://artificialanalysis.ai/models/gemini-3-6-flash
LLM reception is truly extreme, even worse than AAA game releases.
Ever frontier lab lived it at least once : missing the frontier by a few months triggers extremly negative reactions, then you take back the lead for 2 weeks, and the hype cycle repeats.
Looks like 3.6 Flash is the first model with their newest pretraining run (cutoff date is 2026/03), long after 2.5 series.
For anyone wanting a faster overview: I ran the Gemini 3.6 Flash and 3.5 series release notes through NotebookLM and generated a short video summary. Link: https://www.youtube.com/watch?v=SUFBhvQ2tY4
IMO Gemini has the best free tier models/app for everyday use. Muse-Spark is perhaps just slightly better, but has none of the connectivity to my GApps (for things like “create a recipe in my Google Docs from this image”).
Plus they are probably running these things on every Google search so saving tokens is a huge win for them.
I'm a big fan of the Flash-Lite models. They're exceedingly fast and deliver great outputs for high volume use cases where you need to process requests at scale. Can't wait to try the newer version.
I deeply wish Google would focus on models like Gemma. Small, powerful, open-weight models you can run on phones or regular computer hardware.
Nice to see that it's cheaper than 3.5
So 3.6 Flash is a somewhat of an admission that Google miscalculated by charging 3-5x for 3.5 Flash what it did for 3.0 Flash (3x input and output costs plus large token inefficiency changes) despite only modest improvements?
3.5 Flash Lite is only a hair cheaper than 3.0 Flash, but I think 3.0 Flash is a massively more capable model?
Flash Lite: 0.3/m and 2.5/m
Deepseek Pro: 0.435/m 0.87/m
That's wildly ambitious pricing by Google. You can maybe get away with spicy pricing at the SOTA edge but at the lower tiers everything is a lot more price sensitive.
It does seem like their releases are getting closer together. I get the feeling they realized they were trying to roll out to their entire ecosystem and now they’re focusing more just directly on the AI model itself. I think give it a little time and they’ll start to be one of the competitors too.
Proof-of-life release while they figure out how to have a competitive frontier model release. My hunch is they pushed too far in the "omni" model direction, that they made something so ungainly, it wasn't as good for normal tasks.
"..and in some benchmarks like DeepSWE by Datacurve, we observe up to 65%, all at a lower cost per output token."
"3.6 Flash delivers higher precision with fewer unwanted code edits and reduced execution loops, as seen in DeepSWE (49% vs. 37%)"
So which one is it? 65% or 49%?
I was expecting 3.6 Pro. It's been so long since the last Pro model...
Are they comparing 3.6 Flash to 5.6 Luna and losing? That's ruff.
I read this as a soft let down to not expect too much from 3.5 Pro.
> We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress.
The benchmarks are not particularly impressive. I suppose they needed to release something since the long pause. But not clear why would I use it now.
I have liked using their consumer products but they don't make it easy, that's for sure.
Glad to see the price is going down but it's still too high for a "fast" model
"We made 3.6/4 Pro, but it sucks, so this is the distilled model" vibes.
A lot of disappointment here in the comments, but models like these aren't meant to compete with the likes of Fable or GPT 5.6.
I use 3.1 Flash Lite regularly to classify listings on eCommerce websites. It's great for this task - fast, cheap and accurate.
In fact, it was the single best model we tried in terms of the speed vs accuracy vs price tradeoffs - including the Chinese models.
Of course, 3.5 Flash was more accurate but the 5x cost increase couldn't be justified.
3.5 Flash Lite sounds like it could be a strict upgrade for our use case, without a significant increase in costs or drop in speed.
It's not GPT-6 but it's not trying to be. It's a completely different tool and great at what it does.
Pretty underwhelming, as expected honestly. I don't want to know what morale is like at DeepMind right now.
So about the same “intelligence” as Muse Spark 1.1 but 2x faster and about 2x as expensive.
tl;dr: 3.6 flash is a bit smarter than 3.5 flash, but also a bit more expensive.
My results [0] put Gemini 3.6 Flash at the top.
3.6 Flash high has same $1.5 input price as 3.5 Flash, but output is cheaper from $9.0 to $7.5.
Google said 3.6 Flash is more token efficient, but in my tests it's actually LESS token efficient[1] than 3.5 Flash, so despite the output price reduction, it still costs more.
[0]: https://aibenchy.com/compare/google-gemini-3-6-flash-medium/...
[1]: https://aibenchy.com/compare/google-gemini-3-6-flash-high/go...
I have no skin in this game and this comment will be gray in a few minutes BUT a friendly reminder that these types of threads are astroturfed heavily by competitor labs and any info should be taken with a massive grain of salt.
the real product is the naming confusion we made along the way. Gemini 3.6 Flash, 3.5 Flash-Lite, 3.5 Flash Cyber — at this point even the model cards need a model to explain them
Models are expensive and low performance. On top of that they make you jump through hoops to even use these models without being throttled even for the weaker models. The only reason we are using them is credits. As soon as credits run out we are switching immediately.
quite a good model, the speed/price/quality ration is a new golden intersection for me, not sure if its as good as grok 4.5 but quite fast/capable model.
Google is walking backwards, with such a pile of cash in pocket, i feel they are doomed.
Haven't been excited for a Gemini release since December. Wild to see.
Not good enough for high-end, not cheap enough to be for low-end. Next!
It feels like AI is going to be the end of Google. The post-Schmidt company culture cannot produce consistent, consumer-friendly products that any sane person would want to use consistently.
never tried gemini for coding, but this news seems to be compelling, i would definitely give it a try
"The model will be exclusively available to governments and trusted partners via CodeMender soon as part of a limited-access pilot program"
we are stealing plutocracy from the jaws of emancipation.
i don't want to live in a world where abundance is guarded and shared among politicians and cronies, whilst the rest are left to rot.
no actual cyber model release, useless
Specific to task these can be huge plus point.
2 red flags
1- no comparison with gemini 3.1 pro
2- no comparison with any other model
dupe? https://news.ycombinator.com/item?id=48993130
If only I could use Pi.
Why exactly are they announcing these completely milquetoast models ?
I'd be low-keying the release if anything, given how lame they are compared to their competition.
What am I missing?
Plugged 3.5 Flash Lite into an existing agent harness that was previously using 3.1 Flash Lite and this shit just does not work. It's not following instructions and is not producing the correct tool calls.
I remember back when Gemini looked like it was the best model that this comment section was full of confident predictions that Google had "won" and that no one would ever catch up with them again. The most embarassing part is that I kinda believed them.
if only DeepSeeek supported vision, would never use Gemini.
I'm going to get downvoted/flagged but I feel like we need a new type of "Show HN/Tell HN" etc for "New AI Model Available".
Front page is tedious these days.
I think it’s safe to say Google seems a bit out of the top AI competition now. The “cyber” stuff also starts to become laughable with open models providing the full power without the crap Anthropic, Google, OpenAI are trying to frontload on you(I.e you are not allowed to develop/review a login system, pay a special cyber operation team to do it for you). They really deserve to become irrelevant in the future of AI.
Other discussion from a few minutes earlier: https://news.ycombinator.com/item?id=48993130
The silence is deafening.
Google watches over the last few months a flat out assault on the Pareto curve from American and Chinese companies. Release after release pushing the boundaries of frontier intelligence and price/performance.
And the response from arguably the biggest AI research labs in the world by headcount is Flash 3.6.
What do you do when you are given essentially unlimited resources and still find yourself falling behind?
Gemini 3.5 flash is already a pretty good model. But, unfortunately, the primary way you can interact with it for coding is through Antigravity - which is actively developer hostile.
It doesn't matter how good the model is if you're (mostly) forced to use it in Antigravity - which turns any model into crap.
Wake me up when Antigravity doesn't suck.
I use 3.5 flash 10x more than any other model, despite have access to all of them. If I'm going to play a slot machine, I'd rather get the pain over with quickly.
We're almost five years into the whole GenAI thing and we're still relying on these guys to spoonfeed us incremental updates.
It's time for them to start focusing on open-weight models and efficiency. Otherwise there's just a layer of marketing hype and "will it do this?" that has to be cut through for evaluation of each and every release cycle.
Models are getting easier and easier to create. The money, if there's any here, is in the harness the user interfaces with, and the data centers running them.
Google seems to be falling way behind the pack. antigravity cli is pure trash. gpt 3.5 pro is now behind and isn't released yet. GPT 6 and Fable 6 releasing next month. What the hell is going on over there ?
Just switched to AI Plus from Pro, seems like I won't be missing much.
> we have taken an intentional approach to deploying 3.5 Flash Cyber. The model will be exclusively available to governments and trusted partners
Screw your government! US and Israeli governments should get the least access, but of course we all know they'll be the (only) ones to get full unfiltered access.
3.5 Pro must really suck.
They are comparing against their own previous models instead of competitors. Not a great sign.
At this point just put the Pareto in the bag bruh
whatever, dude. give gemma5
is it just me or is this one-upping each other every few days getting ridiculous secreting a whiff of desperation?
Some more discussion:
Gemini 3.6 Flash https://news.ycombinator.com/item?id=48993130
With all the naysayers on Gemini models I'm curious how many people actually use Gemini regularly?
For me, Gemini models are the most usable. Claude Opus and Mistral always try to turn queries into one-shot enormous commits, which just burns tokens, time and annoys me for something which is still wrong more often than not.
Gemini seems far better at listening to instructions and giving me what I actually want, on top of using far fewer tokens and wasting my time. Fable is the only model that's come close to Gemini Pro for me.
And as this is about Flash, it's exciting, I find Flash can usually get the right answer pretty quickly and without too much nonsense.
I keep saying this and people dont believe me, but I have b2b saas systems with actual agents running around the clock, and the performance/stability of the flash model is higher than most other models.
Meaning, its predictable with tool calls, wont spin off a million tools/do weird behavior, its reasonable. Even sonnet in a real world decision making scenario is not reliable, or will reason so long its incredibly expensive.
The benchmarks arent catching all the value, and most people have never actually ran an ai agent in a real context that matters
not a google fanboy by any stretch... though i've been thrilled with the flash line of models... i exclusively use it on high, and have found it to be a great fit for increasing productivity 10-fold while maintaining quality... sure it can't just go off and one-shot a bunch of work, but at the complexity level i tend to work at, neither can the frontier in a robust way that i can be confident in... sure i have to be in the loop more, but that helps keep me grounded and course-correct earlier before wasting tokens... and when you sufficiently spec out a coding/software problem, and i mean really document all of the critical nuance, it will successfully satisfy the constraints... the quality is rarely acceptable on first-pass, but it forces me to stay connected to the architecture more than i would be if using a frontier model... i've found this to be a happy middle-ground of productivity and awareness...
jew
tested the models on aistudio. despite that the knowledge cut off is march 2026 it still knows nothing about 2025!
you can check by asking "list notable world events in 2025, only list unplanned" on aistudio. or you can ask for Charlie Kirk, it also does not know. I tried it multiple time to ensure that I didn't not get routed to older models!
> but google has search
irrelevant, without deeper knowledge about cutting edge technologies or latest libraries, all of it suggestions are crap. even you ask it to search it will still use outdated keyword thus only getting outdated information.
in other word, what a disaster!
Google desperately needs to make some leadership changes within their Gemini team now that they've been surpassed by 3-5 open weight models and risk loosing frontier status all together in the near future.
At this point, I think google should consider becoming a hyper scaler for anthropic and open ai, and I predict that that is exactly what they do. The model is no longer the most valuable part of the stack.