Interesting to see that peak hours are work hours in China, night in the US and Europe, and also morning in Europe. So Deepseek's customers are mostly domestic.
show comments
javier123454321
Ever since I started using flash, it has slowly crept up to be my default for everything. It is at the good enough state for a fraction of everything else that's out there.
show comments
alkonaut
There is no relative/percentage increases noted (understandably). Just because i'm lazy: roughly how much more expensive is it to work with v4 flash and v4 pro through the API, compared to before the price increases? Is it 2x, 5x, 10x higher?
show comments
alexpotato
I'm no expert in pricing economics but once peak/off-peak pricing arrives, it seems like tokens are going to be like electricity or long distance phone minutes where it just becomes a commodity/race to the bottom.
show comments
roenxi
This is somewhat funny when you realise the data centres are now going to start a process that looks very so slightly like daydreaming. Depending on the time of day they're going to be thinking about different things in a cyclic manner. They're going to be doing things like finishing a hard days work then kicking back to think about tricky math problems.
show comments
PeterStuer
Not yet, but in the end high quality tokens are a commodity market. Every optimization to increase inference efficiency will be universally rolled out. The 'hyperspenders' will run into demishing returns unless regulatory capture succeeds.
hopfenspergerj
Does the API response include a "service tier" response to indicate whether you paid peak/off-peak for a given request? I like to compute cost for each request, and save it with my results.
j1elo
So many changes in so little time, that it all makes no sense. Continuous churning. Reminds me of the experience of trying to be on top of the dependencies in a medium-large JS project.
I am a person that buys into a tool or a process and expects it to be part of the life with no major changes through the years (or as long as the need exists). But AI? You buy into something today, not 2 weeks have passed and there's already a large "update" introduced to the conditions or the optimal usage patterns you should be adopting.
It's tiring. Makes all prices and offers feel so unreliable and gets me a bit more disinterested each time they change.
show comments
declan_roberts
This actually works out favorably for US customers since the peak hours are Chinese working hours and cheap hours are US working hours.
Some of the US companies do the same, but rather than "off-peak" hours they price lower for "batch" jobs with non-committal response times.
The same motivation of course - the GPUs have a finite service lifetime, so to maximize revenue you need to keep them busy 24x7.
xbmcuser
They benefit from a strong captive market because Chinese firms cannot use Nvidia chips and are legally barred from processing data abroad, forcing them to rely on domestic infrastructure.
show comments
poly2it
That's a hefty increase. Flash pricing during peak is now 1.32/M out, compared to the current 0.28/M, which in turn is a quite a bit above the cheapest provider at 0.16/M.
Are we gonna see "we work those unusual hours because that's when LLMs are cheap"?
sebastiennight
With proprietary labs lowering their prices and Deepseek raising theirs over time, wouldn't it possible to extrapolate a graph to look at where the terminal frontier-model million-token-cost asymptotes to?
show comments
mateenah
This is good for other competitors I guess. People rarely calculate the bump in price but the fact that price is increasing might bring them to other vendors.
cheesecakegood
I wonder if this is enough to push people back onto Luna with their comparative price drop
show comments
flakiness
Now Baseten's pricing is cheaper than the official one? Probably won't last, but still interesting.
Interesting to see that peak hours are work hours in China, night in the US and Europe, and also morning in Europe. So Deepseek's customers are mostly domestic.
Ever since I started using flash, it has slowly crept up to be my default for everything. It is at the good enough state for a fraction of everything else that's out there.
There is no relative/percentage increases noted (understandably). Just because i'm lazy: roughly how much more expensive is it to work with v4 flash and v4 pro through the API, compared to before the price increases? Is it 2x, 5x, 10x higher?
I'm no expert in pricing economics but once peak/off-peak pricing arrives, it seems like tokens are going to be like electricity or long distance phone minutes where it just becomes a commodity/race to the bottom.
This is somewhat funny when you realise the data centres are now going to start a process that looks very so slightly like daydreaming. Depending on the time of day they're going to be thinking about different things in a cyclic manner. They're going to be doing things like finishing a hard days work then kicking back to think about tricky math problems.
Not yet, but in the end high quality tokens are a commodity market. Every optimization to increase inference efficiency will be universally rolled out. The 'hyperspenders' will run into demishing returns unless regulatory capture succeeds.
Does the API response include a "service tier" response to indicate whether you paid peak/off-peak for a given request? I like to compute cost for each request, and save it with my results.
So many changes in so little time, that it all makes no sense. Continuous churning. Reminds me of the experience of trying to be on top of the dependencies in a medium-large JS project.
I am a person that buys into a tool or a process and expects it to be part of the life with no major changes through the years (or as long as the need exists). But AI? You buy into something today, not 2 weeks have passed and there's already a large "update" introduced to the conditions or the optimal usage patterns you should be adopting.
It's tiring. Makes all prices and offers feel so unreliable and gets me a bit more disinterested each time they change.
This actually works out favorably for US customers since the peak hours are Chinese working hours and cheap hours are US working hours.
Duplicate:
https://news.ycombinator.com/item?id=49287881
https://news.ycombinator.com/item?id=49285160
Some of the US companies do the same, but rather than "off-peak" hours they price lower for "batch" jobs with non-committal response times.
The same motivation of course - the GPUs have a finite service lifetime, so to maximize revenue you need to keep them busy 24x7.
They benefit from a strong captive market because Chinese firms cannot use Nvidia chips and are legally barred from processing data abroad, forcing them to rely on domestic infrastructure.
That's a hefty increase. Flash pricing during peak is now 1.32/M out, compared to the current 0.28/M, which in turn is a quite a bit above the cheapest provider at 0.16/M.
https://openrouter.ai/deepseek/deepseek-v4-flash-0731#provid...
Are we gonna see "we work those unusual hours because that's when LLMs are cheap"?
With proprietary labs lowering their prices and Deepseek raising theirs over time, wouldn't it possible to extrapolate a graph to look at where the terminal frontier-model million-token-cost asymptotes to?
This is good for other competitors I guess. People rarely calculate the bump in price but the fact that price is increasing might bring them to other vendors.
I wonder if this is enough to push people back onto Luna with their comparative price drop
Now Baseten's pricing is cheaper than the official one? Probably won't last, but still interesting.
https://www.baseten.co/pricing/
If anyone has tried Baseten versions of these Chinese frontier models, let me know what you found.
[dupe] https://news.ycombinator.com/item?id=49285160
Now if there could be a bot that defers queries until when it’s cheap…
Full table with multipliers from previous prices:
DeepSeek-V4-Flash (off-peak, x2 for peak)
* Cache Hit $0.007 (x2.5)
* Cache Miss $0.22 (x1.5)
* Output $0.66 (x2.25)
DeepSeek-V4-Pro (off-peak, x2 for peak)
* Cache Hit $0.022 (x6)
* Cache Miss $0.66 (x1.5)
* Output $1.98 (x2.25)
Peak Hours: 01:00–04:00 and 06:00–10:00 UTC
Effective from: 16:00, August 16, 2026 (UTC)