The reason people aren’t freaking out is because most people are using heavily subsidized subscriptions.
I tried the cheapest provider on openrouter and burned through $50 in a few days. Quality was ok, seems slightly above Luna quality perhaps? But that $50 is 1/4 of my codex subscription where I could have burned that many tokens or more using Astra within my weekly reset.
This won’t last forever but as long as the frontier labs are subsidizing this heavily the open models won’t matter.
show comments
mlinsey
I'm paying for the heavily-discounted subscriptions, not the API rates. There isn't really a cost gap for me. DeepSeek doesn't have a subscription to compare to, but when I compared the GLM 5.3 usage I got from a $100/mo Z.ai subscription compared to Opus 5.5 on a $100/mo Claude subscription, there wasn't a big gap. And GLM 5.3 is very clearly not a frontier model (deepseek v4 seemed a lot
closer, but I didn't use it enough to really say for my workloads).
I don't think those subscriptions nave negative contribution margins, either. I think we're seeing a lot of price discrimination by the big labs, and huge margins on their frontier models. The fact that they have been cutting prices to their second-biggest tier of models (Opus/Sol).
Open models catching up and collapsing these margins would worry me if I were a shareholder in the big labs, but as a user, I really doubt that the western labs have bigger environmental impact just because they have higher API costs, I think they have a ton of efficiencies they aren't sharing with customers yet because demand is so high.
show comments
damowangcy
Marketing.
Have you guys see how aggressive is the push for enterprise use by both OpenAI and Anthropic? I had friend from a non-tech industry in Asia telling me that their company was offered free trial of the enterprise version of Claude, with trainings and such.
On the other hand, DS and Z.ai, have zero to none marketing outside China. There is friction to use DS/GlM models and the ZDR is unclear, so most enterprise that has heavy AI usage hasn't move over yet. They would rather spent $200 for the peace of mind than to take the risk of being slam as a national traitor down the road (which again is another form of marketing by Big AI, trying to frame Chinese models as thiefs).
So, I don't think they are not freaking out, it's just that they are addressing different market segments and reacting to the situation differently.
show comments
jarbus
Thank you for posting this, I didn't realize just how good 4.1 flash was. it far surpases gpt-luna in the limited testing I did: luna would take so many tokens that implementing a new feature would be nearly as expensive as if i just did it with sol 6.1, but 4.1 flash can do it for basically 1/6th the price, if not less. Obviously depends on the feature, but very very impressed.
gregwebs
I have been using DeepSeek 4.1 flash intensively for over a month. If I run it all day long it costs $1-2. Its fast. Previously I was always quickly running up to my Claude/Codex 5 hour window (on the $20/month plan). The cost savings of DeepSeek is real as shown in this article and I am using subsidized plans.
DeepSeek is horrible at grilling sessions (the /grill* skills to make technical decisions). It doesn't know how to explain things. Maybe the skill could be adjusted. It also doesn't come up with as good solutions as Opus/Sol.
What I use it for is
* the orchestator of my coding workflows
* the tester/verifier of code changes
* the sub agent that explores code or does web searches
* putting together code base research reports
Previously I planned with Opus/Sol/Astra and then I used DeepSeek for coding, and then reviewed with Opus/Sol/Astra. With the cost improvements to Opus/Sol I am trying to use them for coding instead now so there will be less back and forth review needed.
They are all working together in Pi using the extension @tintinweb/pi-subagents where my workflow skill is calling different subagents that use different models.
Luna is cost competitive, but doesn't score as well on intelligence. I do need the intelligence for most of what I use it for, so I am not motivated to use Luna. Haiku also doesn't seem like a competitive price/performance mix.
show comments
lmf4lol
Oh man. v4.1-flash has been an sbolute game changer for us. We run all our Personal Assistants now on flash (thinking high) by default and it works incredibly well. There is really no need for basic agentic tasks that might require Kimi K.3 or GLM-5.3 levels.
Once its gets juicier, we let flash launch specialized subagents with specific models. GLM-5.3 for coding or Kimi K.3 for research and critique.
But as a main driver. I love flash. And it brought our bill down by A LOT :D
show comments
user43928
Because DeepSeek is not "a month or two" behind as claimed in the article.
These open models still did not beat February's Mythos / Fable 5.
DeepSeek 4.1 Flash is behind GPT 5.6 Sol, and that one is left in the dust by the excellent Opus 5.5.
Rumors say Anthropic is holding in reserve the big improvement, Fable 5.5, for the IPO.
It's plausible that open models are 6 - 12 months behind, and there is no "good enough". As long as progress doesn't slow down, leading labs have nothing to fear.
show comments
p1necone
I have a pretty large, complex project I've been building with heavy AI use (new language + compiler). I was following a 'strong model as orchestrator launching cheap models as implementers' pattern, but I recently trialled just using Deepseek-V4.1-Flash as the model for both layers because of the cost savings (with mimo v2.6 flash on code review agents for some decorrelation).
I was previously using GLM-5.3 as the orchestrator, after switching to DS anecdotally there was an unnacceptable quality loss, mostly around not taking all the relevant context into account when making decisions, pulling new design out of thin air without discussion too often, and being way too wordy and rambly in documentation despite prompting to avoid it. There's a lot of docs, rulings, core concepts, design philosophy to uphold and DS was just not cutting it.
However, it's perfectly capable of being the sole agent for all of my well specced implementation tasks. I've gone back to GLM as the orchestrator.
show comments
joshstrange
I’ve played with DS4.1 Flash. It’s neat, no doubt. In my tests it does decently well against Sonnet 5.5 but uses way more tokens. That works out since it costs ~1/3rd Sonnet per task/PR according to my tracking.
However, I can easily burn ~$5/day if I use it as my “worker” (still using opus for planning and review) so call it ~$150/mo.
I could, if I were so inclined, get a second Claude Sub and have even more headroom (though I’m able to stay under my limits most of the time with my current setup). Also Claude gives me Artifacts, Web Search, and now even some API Credits.
I have no doubt the future is open weights and I can’t wait, literally, I can’t wait for them to catch up on intelligence or for hardware to run a decent model to be within my grasp. But until that comes to pass, I’ll keep using Anthropic.
show comments
827a
Its simple. My company is very willing and able to pay ~$200/engineer/month for the best version of these tools. My company is even willing and able to pay as much as $500/engineer/month, but does not need to at this moment.
My company is not willing to pay $50/engineer/month for a cheaper version that is nearly as good. My company is also not willing to pay any amount for a product produced by China, even if it is hosted in the United States.
show comments
hmontazeri
I had the same experience using ds 4.1 last couple of weeks. It’s insanely good for the price. I’m doing mostly web dev it excels at everything I throw at it. The pricing is ridiculous. I canceled my gpt subscription and haven’t looked back hope the pricing stays like that. I almost never need a better model. I still keep my Claude 20$ sub for now but I feel like one more iteration and I won’t need even that anymore I hardly use it
show comments
kraig911
Everyone I talk to about Deepseek has a bad taste it feels like for CHinese models? I have a Kimi sub and use deepseek a lot at home. I've hit my limit on K3 many a time for Deepseek to come in and finish a job. So far no complaints. I feel it makes a lot of round trips by design but over all a good option. It's hard to compete though with an open sub out on codex/claude and just do everything in that subscription. Deepseek for me is when I run out of my subs. It's usually always coming behind a great session and fixing things so I don't give it a chance to try something novel on it's own.
weknowbetter
People who have not used DeepSeek massively under estimate it's capability and how little it costs to run.
show comments
simpaticoder
The question seems rhetorical but I think there are two reasons in some combination. First is it there is some awareness lag here. That lag can be on the producer and consumer side. Software enterprises are pretty slow to adopt new things and slow to try new things so they might only be aware of openai and Claude as options. Plus there are some scariness because deep seek is a Chinese model and therefore export restricted - never mind that there are American in European providers.
The other reason is more interesting. Maybe the frontier providers think that price performance is irrelevant in light of very powerful frontier models that can start the RSI loop and or a huge displacement of work and a winner take all economic situation. After all if frontier providers earn everyone's money then you won't have any money to spend on any model 100x cheaper or not.
show comments
aguilaair
What about MiMo v2.6 Pro? It’s throughput is slower by default (UltraSpeed is faster than DS4.1F) but is above the pareto line, and cheaper.
While using DeepSeek v4.1 Flash I was architecting a system and I made a mistake of drawing the RPC boundaries at a wrong place that did cost me in so many ways.
I realized that mistake and guided DeepSeek where it should be.
Next I fired Fabble 5.5 set to high to check if the hype is real about Fabble. It exhausted 89% of quota and came up with NOTHING that DeepSeek hadn't flagged itself already in its notes.
show comments
arush15june
I am 4.1 maxxing on commandcode GOAT Plan + api rates with oh my pi for the last 4 weeks, it's absolutely amazing and crazy fast, it's alright if it makes a mistake, I have enough time to iterate again, I have also added an advisor layer of mimo 2.6 pro which does make it a notch smarter. Getting haiku 5.5/sonnet5.5 to work on plans and letting 4.1 flash work through it is helping a ton too.
I am a big ChatGPT fan, all our team has ChatGPT Subs, but the TPS across all models including luna is just so damn slow.
Commandcode giving 60$ worth of Deepseek for 10$ is just genuinely goat.
And it never says no for cyber tasks so that's a big win
show comments
pierreb-aiva
Because of several factors:
- Models are still very jagged. A superior model (e.g. Opus 5.5, Astra) may not be materially better in some tasks, but usually is for completely new modalities: computer use, game development, etc. People will always prefer using less jagged intelligence, because it allows them to do so much more
- Distilling intelligence from frontier models means that unless chinese labs manage to replicate the training regimes of OpenAI & Anthropic, they are always going to be behind a few months. That's a feature of training by distillation.
- China doesn't yet have the compute to compete at the frontier. Because they don't have access to top tier chips, their GWs are not equal to US GWs. Until they close the hardware gap, I think they will always be focused on competing on efficiency, as opposed to intelligence.
- OpenAI and Anthropic subscriptions are valuable of quality of intelligence + generous compute allowances
xzjis
The problem with this low listed price per token is that, in reality, DeepSeek 4.1 Flash uses 10× more tokens than GPT-6.1 Sol for an equivalent task and delivers a lower-quality result. So there’s no real benefit to paying 10× less per token. Also, as someone else pointed out, OpenAI and Anthropic currently offer subsidized subscriptions for $100 or $200 a month that provide far more tokens than the API, so we should take advantage of them while that lasts.
show comments
RGS1811
This model finally got me off my Claude Max subscription. I’ve found it superior to Opus 5.5 in certain use cases, and certainly faster.
I’m convinced that I’ll have good enough inference on my laptop at reasonable speeds within the next year.
ojr
It's like the boy crying wolf with these models, Deepseek is not as good as Claude, I like using Gemini Flash Lite even, it is good enough for most of my crud app tasks, brownfield projects you don't need the strongest models, but greenfield projects with lack of context from indexed code and lack of domain expertise, I think the Opus models have been working for people
apitman
> With my OpenCode Go sub of $10/month, DeepSeek is basically unlimited
My OpenCode Go monthly window was scheduled to reset this morning. It was sitting at 22% used despite me using DeepSeek V4.1 Flash heavily as my implementation agent the past couple weeks (I use gpt-6.1-sol high for planning/orchestration).
I had 1.5 hours left so I fired up first 10, then 20, and finally 50 concurrent subagents all working on reverse engineering C code from an old PC game. They found over 100 new functions.
This is the first workload I've found that could make a dent in my sub. It got my 5 hour window to 85% used, but sadly my monthly was still only at about 35% when it reset. So that cost maybe $2.
show comments
alex-moon
I think because we're all just using it thinking we have found the "model for me" and never mentioning it to anyone because what would we say? It's good. It's a bit like the Logitech MX Master, as more and more people assumed they had found the ideal mouse for their purposes, it quietly became the professional standard through sheer adoption.
show comments
Primer81
I've been waiting for something like this article for a month. The other AI companies are being left in the dust. I rip about 200 million tokens at least a day and spend maybe $3 with deepseek flash v4.1 for self hosted related coding tasks for ~20 projects i work on / maintain in parallel. Its a daily driver for sure. not even worth considering anything else, but maybe a locally run model on my GPU at the moment.
show comments
gutchapa
Actually OPUS sucks... Deepseek has its own flaw, but fares lot better in many instances. And yes it's far subsidised. If you are mindful of its peak and off peak hrs, you could make best use of its cost surge.
K0IN
I really really love deepseek v4(and 4.1), but for everything I thrown at it, it felt like gpt 5.6 + 6 luna or terra do it faster, in way less tokens and in a way I like more. So even tho the price is (very) cheap, i found myself using it less just cause I don't want to wait on something I have to itteratee on with the model.
erans
I've been using DS 4.1 Flash for a while now as my main workhorse. It's amazing. It's fast and it is also a GREAT debugger - it might not always find the best solution but it will definitely find the problem quickly.
I'm also the CEO of a new company - LunaRoute - so if you want private, fixed cost, we run our own server in the US type of DeepSeek 4.1 Flash check out https://www.lunaroute.com
If you want a trial, hit the contact us and write that you saw this post.
We also have GLM 5.3 (with vision!) and GLM 5.3 Flash - all included.
swiftcoder
I think the interesting provider to cross-check this assumption with here is Meta, who is clearly freaking out, and is currently providing Muse 1.3 even cheaper so long as you are willing to share data with them
show comments
nightpool
> This is like comparing big pharma with generic manufacturers who can skip the R&D. While these drugs are not 1 to 1 copies, we are still comparing apples to apples, but with a 90% price cut.
Weird! It's almost like distillation is bad for the long-term growth of the industry, just like generic manufacturers would be if they could release the generic versions of drugs 2 weeks after the original R&D completes.
Very strange sentence to include in an article after saying "I don't care at all about distillation" up at the top. Can the author not hear themselves?
show comments
zug_zug
I did a test on this a couple weeks ago. What I found was that the chinese models were far better than API rates, but about comparable on price vs the subsidized subcription model (chatgpt). Also it was my experience that codex completed tasks quicker.
That said, it's my best understanding that these american companies aren't profitable and will eventually raise rates (the old uber trick) so I'm keeping myself ready to switch when that day comes.
show comments
poulpy123
Why ?
Because I don't have the time and money to do an extensive benchmark of all major LLM, so when I had to select a LLM for my usage (which was not coding at first), I went to the most used one, chatgpt, because I knew if would be one of the best at the task.
I suspect it's the case for many if not most people.
rbnafo
Because most of the people use it through enterprise agreements and don't pay the bill?
I run it for my own use cases and its pricing plus caching capabilities are hard to beat, cents for millions of tokens.
https://substack.com/@rubenafo/note/c-332218129?r=26y5kn&utm...
giancarlostoro
Call me crazy but:
VRAM & Memory Requirements by Precision
• FP16 (Full Precision): Requires ~1,664 GB of VRAM (e.g., an 8x B300 288GB cluster).
VRAM aint cheap, Sam Altman ruined the cost of memory, Nvidia doesnt make enough consumer GPUs letting the market go insane over them, I still have friends on 1070s or 1070 TIs because GPUs have been severely overpriced for too long. I remember when a gaming PC was only $1000.
Even so why would anyone not sleep on a model they cannot run?
show comments
jokers132
> It's like asking a math PhD to organize the files on your desktop.
I actually do use an agent harness to organize files on my desktop. They make a great fuzzy file renamer. Point it at a directory of disorganized files with names all over the place, give the directory layout and file name pattern you want it to have and it makes it happen.
_jayhack_
Enterprise is not freaking out because DeepSeek 4.1 Flash does not actually occupy a spot on the Pareto frontier for non-coding enterprise workflows. We see this at my employer, focused on non-technical knowledge work. Luna 6 and now Haiku 5.5 are both very competitive if not better on all axes that we care about
Kuyawa
> China is going to eat their lunch
No doubt about it, that's why their push for international regulation to the levels of nuclear inspections using the narrative of annihilation and apocalypse
james2doyle
Been using Flash 4.1 via the ante harness to blast through a GBA recomp. The ante team has pushed hard to make Flash 4.1 perform well under it. So far, I've maybe spent $10 over the last 3 days. Its a real workhorse and works much better in this harness
atleastoptimal
Deepseek and many other models are heavily benchmaxxed. They aren't genuinely as good as the best frontier models.
wasfgwp
Because it’s not even that cheap? The author chose to only include Claude in their chart and ignored the fact that 6.1-sol and even more so luna can easily beat Deepseek on cost. Of course almost free cache used to be the main differentiator, raw token cost is deceptive since 4.1 just uses way more tokens than most other models
peteforde
I actually have a fairly simple answer to that: if it doesn't come up in the list of LLMs that Cursor supports, it effectively doesn't exist.
I'm well aware that there's nearly infinite opportunities to yak shave "perfect" OpenRouter setups and some people appear to enjoy bouncing from IDE to IDE as though change costs aren't a thing, but I discovered that I genuinely like Cursor and at least right now it's insanely subsidized by Auto clearly defaulting to whatever Grok's most powerful model is.
I dropped my $200/month subscription to $20/month and stick to Auto for all but really important Plan tasks, and I have basically zero chance of using up my monthly credits even using it 6-10 hours some days.
show comments
WiSaGaN
Deepseek v4.1 flash is my baseline model to use at original provider's api. The issue is, for my personal use, it's cheap enough that i don't need it to go cheaper compare to the time I spent using it. And it's already a very capable model in dealing everyday simple things. For sure, for research level questions or large scale coding projects, I would want to use frontier model. But more and more daily tasks can be done now just using pi with deepseek api directly without thinking about much.
ElProlactin
> Sure, they stole Claude's training, and Anthropic stole it from other people. I'm not getting into the whole who-owns-whose-data debate, because most developers aren't thinking like that. They're just trying to get the most bang for their buck.
Until they get laid off and suddenly discover their moral compass.
Roark66
Honestly, having benchmarked 3 models Qwen3.8-Flash-Next, GLM5.3 Flash and DeepSeek 4.1 Flash. I absolutely do not understand the hype about Glm and DeepSeek. It's an improvement over older models, but Qwen is an actual opus replacement for me. It has been for last month.
But I run it locally. When I tried it on open router when my gpus were busy I must have gotten routed to some crappy providers, because it was pretty bad.
For me Glm and DeepSeek are nowhere near this Qwen model. I tried various harnesses including omp which I heard supposedly "makes DeepSeek 20 points better". The difference was in the noise (1 point). I run a bunch of benchmarks Terminal World 40, terminal bench 2.1,SWE Pro, GSO. Before those 3 there was no open model that scored more than 1 point on my subset of GSO. Glm scored 5, DeepSeek 3, but Qwen did 17 and opus 19.
Qwen is a small model so it fails on factual recall. But if you give it most of the info it needs it us amazing.
Anoian
If you use open models you can also read the thinking tokens, I think this makes a huge difference because one tends to read them leading to better understanding of the end results and also giving opportunities to correct course if its hung up about ambiguity and stuff like that. It also helps a lot with not being distracted while the clanker is working, instead of going to youtube or reading up on whatever, you spend more time with the code and the thoughts behind it.
It stops the constant context switching.
liuliu
DeepSeek 4.1 Flash 0910 is perfect for M5 Ultra 256GiB. Running it fully resident in RAM, prefill at ~2500 tok/s and decode at ~40 tok/s. Probably tons of room to improve from there.
show comments
balboer4487
we do. we moved 100% of our production traffic from Gemini to Deepseek. best decision we ever made.
robertheadley
DeepSeek Flash 4.1 is generally what I use as as my workhorse. I use ChatGPT web to create the outline, then have Deep Seek built it out. Works generally pretty great.
maxdo
The hype is almost reverse now . With opus , sonnet and haiku 5.5 , why should I care about Chinese models that are so behind ? Also Jev seems like killed a big chunk of deepseek market too .
I personally used it for secondary research agents and classification . Now classification part is gone .
show comments
jerieljan
I've been on the API-only mentality for months and all the open models were definitely the stuff I loved the most. Kimi K2.5, 2.7 and Deepseek v4 were among my favorites, while sparingly using Opus or whatever OpenAI had for specific situations.
But ever since I've switched to one of the $100-tier subs, I can see why a lot of the people on it don't really discuss the open models often. I'd still use it especially when it comes to sensitive inputs, but for most work, what you get on OpenAI or Anthropic is really more than enough.
It really got even better when they also made their cheaper models up to par if not better than the open models.
I do think the crowd for open models are out there, especially when you see trillions of tokens running for them on OpenCode or OpenRouter leaderboards.
seanmcdirmid
It’s still cheaper to get a subscription to an agent harness with frontier model backing for most people. DeepSeek really only becomes appealing to me when I want to do something the monthly subscription harnesses don’t support (API access) or more harshly charge quota for (like running coordinated sandboxed subagents). DeepSeek is then great because of the low price and price transparency, it just can’t compete with subsidized monthly access.
taf2
it's ok but not great compared to the commerical LLM's it's really not very good. it's benchmarks are clearly fake. even still we run it internally for a ton of workloads on our rack of gpu's
pants2
Probably because Luna is faster, cheaper, and approximately as smart
chicco4life
Thanks for sharing this.
I've been focusing on deep research related tasks for biotch and life science applications. The problem with this sort of task is that we need subagents to reason through multiple (potentially 100s or more) chains of knowledge/concept/evidence, so the token usage really explodes as complexity of the task and data expands. A typical task can cost me nearly a $1k overnight...
I've been testing out GLM5.3, but now I'm really tempted to try to Deepseek 4.1 flash too. Any chance you've benchmarked / compared the two?
elmer2
DeepSeek isn't even on my mind. I use the frontier models and can get the best in the industry for a relatively cheap price.
show comments
anon373839
I think the industry is freaking out about open weights models in general, if not specifically DeepSeek. That is why we're now on the ~4th call for pacing the frontier from the very people who, if they wanted to pace the frontier, would simply do it rather than asking Washington to get involved.
And it's why people like Hillary Clinton have been trotted out to talk about the dangers of open weights models -- I mean, does she even know what that phrase means? (I know HRC is a controversial figure and I'm not bringing her up for that purpose; I just note that she and other prominent retired politicians are now doing the circuit on Anthropic's behalf.)
show comments
pietz
Aren't they? Not about 4.1 Flash but open weight models in general.
The fight between Anthropic and OpenAI is just about who leads the duopoly. Chinese open weight models are threatening a duopoly in its entirety. That's why Ant and OAI are begging for regulation. A regulation that will hit Chinese models way harder than US models, finally giving them a moat.
ctolsen
Not sure "freaking out" is the word I would use, but it’s fairly obvious looking at OpenRouter usage that the price cuts on Luna a while back were in response to intense competition from dsv4.
So the industry is responding, where it matters. Which is on heavy API usage, not coding subs.
habosa
It's quite good and the labs are definitely scared, that's why they are lowering API prices and continuing to subsidize subscription plans aggressively to keep anyone from using this stuff.
That said, has anyone else found DS models to be unpolished? They seem to "lose their mind" a lot more often than Claude/GPT. I have tried all of the top open source models that came out over the past ~4 months or so and the GLM models (5.2, 5.3, and 5.3-Flash) have been much more usable for me. They feel like Opus but X months ago, DS feels like something else.
santiagobasulto
Comparing it with Anthropic, ANY model is cheaper and more effective. Don't get me wrong, Anthropic models are good, but they're always more expensive for the same task, even compared to other closed source models. At this point, I honestly think Anthropic has played the nasty trick to fine tune the models to be too verbose and charge us for more tokens.
throwdbaaway
Finally an article that gets the maths. Following the release of DSV4 preview where 1M context can fit in single digit GB of VRAM, frontier labs pricing just became stupidly expensive. Other Chinese labs and r/LocalLLaMA also can't compete on pricing.
And since then, there has been so many articles that made it to HN front page, and all of them didn't get it. They just went on and on about tokens generation. Most HN commenters didn't get it either, find-in-page for "cach" typically yield 2~3 responses. If I had a dime for every time this happened, I could have .. paid for 1B cached input tokens?
Anyway, DeepSeek still has to come up with a frontier model, and they almost did it with DSV4 Pro 0813, which is just slightly below GLM 5.3, but 30x cheaper. Unfortunately, the massive price hike happened just 3 days later.
DSV4.1 Flash is good, but not quite the same level. Much easy to self-host though, especially for serving a team of developers. Let's see what the next one can do.
elorant
People aren't freaking out because self-hosting isn't a solve issue and it requires a lot of capital. The average company won't go and build an $1M GPU cluster just to self-host any model. If that gets commoditized then they'll start freaking out.
show comments
zackangelo
I've been loving DS4.1 Flash, it's been one of my daily drivers since we started testing it internally.
We just launched it on our platform today (Mixlayer, https://mixlayer.com), promo code LAUNCH-DSV41F gets you some free credits if anyone wants to check it out.
AYBABTME
Not sure why the author thinks Anthropic's models spend more power and water than DeepSeek, there's no evidence of that. Their pricing has more to do with premium perception and less to do with COGS.
Everyone does optimization of model serving because it's good for every player in there.
(Also the water consumption thing is not a real issue.)
rw2
Because the free api is mispriced, opus 5.5 on subscription is 10x cheaper. Also, the use case for flash models are
For coding, I rather spend 10x more than have even 1 bug but I'm only spending 2x 3x more if you count subscription cost.
This is a killer use case for something like customer support though.
Shekelphile
Luna and Haiku 5.5 are just as cheap and much better.
Don't really understand people who say DS4 or 4.1 have frontier level performance. Anyone who has used it will tell you that it's a hallucination factory. The only thing it has going for it is deepseek's unique infrastructure that allows better cache retention, but the cost savings from that obviously come nowhere near how much subsidized usage you get out of even a $20 subscription with openai or anthropic.
show comments
thefourthchime
For non-coding tasks it may be fine. But for coding, Opus 5.5 is just a completely another level than something like Deepseek 4.1 Flash.
Personally I have not used anything but 4.1 since it came out. I have a dataset that turns any model into pure hallucination machine, not only DeepSeek does not hallucinate, it builds new insights by combining its insights. It's not only cheap, its far better (at least for me)
show comments
leoyoung2026
I’m also a heavy DS 4.1 Flash user—especially when it’s available at those off‑peak prices, which is an awesome deal. And, like you said, it’s genuinely powerful and very snappy. I’m planning to evaluate the differences between `reasoning_effort` settings today.
flying_sheep
No one will freak out until anyone can run frontier model in their own laptop ;)
LeFantome
Not only the model but the hardware it is running on. The Huawei chips they are using instead of NVIDIA are vastly less expensive. China is going to scale past the west. The idea that they are “only a few months behind” is today and many of us cannot even bring ourselves to admit it. The future is even more dramatic.
codeprimate
Mimo 2.6 Pro is even better.
I switched from DeepSeek 4.1 flash about 2 weeks ago for my Hermes sysadmin/coding agents and I am seeing better intelligence and lower overall spend.
Because we are still in the phase where we can expect there to be something else to freak out to next week.
ApolloFortyNine
I have no idea if anthropic can actually make money at their subsidized subscription rates (you can easily hit your monthly cost in one 5 hour session if you price out the tokens through the api), but if subscriptions didn't exist, I do think everyone would be on deepseek 4.1 and just not look back.
skc
Eventually the party will be over but for now it should be a no brainer first choice tool for some 90% of dev tasks
LeBit
I have subscriptions to OpenAI and Claude but use DeepSeek 4.1 Flash for my coding agents.
It costs pennies and you got really great output.
The author is spot on.
joshrw
I am rooting for the Chinese models to win so they can save us from our capitalist overlords.
However, the quality of the open-source Chinese models is terrible. They game or fake their benchmarks because in real-world usage they suck.
airtnp
Because good enough in the writer's context is a pretty low standard. While many people regards GPT 6.1 Sol or Opus 5.5 as "incapable" in some cases.
Just try Opus 5.5 reminds me how Opus 4.5/4.6 astonishes me. Completely different, and GLM-5.3/Kimi3/DS-4.1 are still like Opus4.8 levels.
browningstreet
What would freaking out look like, or is this just a stupid bloggish title flourish?
Is OpenAI coming in $20B under a sign of "freaking out"?
show comments
smallmancontrov
They might be. They would delay public admission as long as possible, because public admission would make stocks go down.
rurban
That's what I thought until August. Used it for half a year almost exclusively. But after the API price increase I'm back at the Claude Pro and Kimi subscriptions.
It's a solid little model, and I appreciate DeepSeek's commitment to the bit in releasing a brand new pretrain, double the size, numerous architectural innovations as a ".1" release over the excellent DeepSeek V4 Flash.
cyberrock
Maybe this is a minor issue, but it seems like different providers on OpenRouter etc. have different quant settings. I imagine that that affects the perception of the model quite a bit.
abhinavsharma
I just keep getting more ambitious with what I use AI for; and that type of work needs surfing the frontier at all times. Because ultimately, many capabilities are not yet saturated
shikck200
Companies pay for claude, devs not. IF i had to use AI from my own purse, i would never pay for a claude sub.
shadyr
I've been using DeepSeek's API and have been happy with it, but I might look into OpenCode as well. Does OpenCode run a quantised version or use different providers from the official one?
ne01
Deepseek V4.1 Flash is a hidden gem, really. Not to mention, you can easily get it through many providers that offer zero data retention and consistent speeds above 200 tokens per second!
scosman
> Chasing the latest and greatest is silly
Said every week by someone who would never go back to using the model they had 6 months ago
yuhmahp
We use Claude Fable to plan, and DeepSeek 4.1 Flash (hosted on DeepInfra) for everything else. Very cost effective.
npodbielski
Because I am using local Qwen. Flash Next is awesome!
When hardware will not be a problem, everybody will be running their own models.
brunooliv
It’s obvious: they train on prompts and store data when using through their official API.
And for third party it’s just… not good. That’s it.
minton
> So why aren't the frontier labs freaking out right now?
They are. Isn’t this why they’re trying to get regulatory capture?
s0ulf3re
I’m partially assuming that there’s a bit of burnout.
jbellis
I built mjolnir in large part so I could have Opus manage DeepSeek Flash subagents. It's phenomenal and extremely light on the Claude tokens. https://github.com/BrokkAi/mjolnir/
And yes, Opus is enough smarter than DSF that it's worth the extra steps. This ranking is from live tickets, no contamination: https://slopcop.com/power-ranking
show comments
profsummergig
Why isn't the author worried about sending her/his ideas to DeepSeek online (instead of hosting it and using it locally)?
Iolaum
One could argue that all this whole "Pacing the Frontier" bullshit is the industry freaking out regarding the danger of open models.
P.S. That's not to mean there aren't dangers regarding AI. I just don't trust the people making money from selling AI to manage those risks ethically instead of "protecting" us from those risks like pimps but with suits and good manners.
pianopatrick
I was just using a bunch of models in Cursor to review a project. I went looking for DeepSeek and it was not one of the options.
Would be cool if they added it.
show comments
xyzsparetimexyz
There was a moment 3 months back where the sentiment was that cheaper models were the way to go. Since then the pendulum has swung back.
linzhangrun
There are too many commoditized models to count: GLM5.3 Flash, Kimi K2.8, Mimo V2.6, MiniMax M3.1...
sotander
Because it hallucinates a lot. There's no free lunch. Although the new architecture is a genuine move forward. The DeepSeek guys are really top notch researchers and devs.
f6v
My anecdotal experience is that I can’t even trust DS4Pro let alone Flash. I always have to have Sol reviewing the code.
aus10d
Very interesting post!
kristianp
> shrank the KV cache by roughly 437X
Can't you just say "shrank to 1/437th the size"? It's not that hard.
wildster
I like GLM 5.3 Flash, it seems good enough for coding features if you have a good structure and a good AGENTS.md
show comments
tengbretson
I don't know about "freaking out", but I'd say I'm having a good time here with DS 4.1 flash.
andxor
Because Haiku is more performant and costs less, even at API prices.
aszen
Because subscription plans are cheaper, only enterprises paying per tok pricing should be freaking out
gsky
America bans Chinese models sooner or later just the China banned American big tech
jokethrowaway
Once companies start realizing how much they're spending in AI to OpenAI and Anthropic for not much more employee output (the bottleneck is always the initiatives, not the code), they will look for cheaper options and they will eventually discover the chinese models.
My AI pilled clients who were early AI adopters are already there and they are looking for solutions to spend less.
aussieguy1234
What blows me away about this model is it's speed.
It's way faster than Opus or any of the GPT models.
I have a coding harness which is opencode plus a few skills relevant to my workflow. Deepseek 4.1 Flash does very well in this environment. I haven't noticed much difference quality wise compared to Opus 5, which I use in my day job as my employer pays for it (although I'm considering using DeepSeek here too given how cheap it is).
hypfer
Is it known why unsloth seems to not have touched DeepSeek 4.1 Flash?
show comments
anguralbanish2
I would love to get them more better, it's good not a bad thing.
Frannky
I mean, they kinda tried to regulatory-capture the market after trying to scare the public, possibly because those models will be a cheap option that gets the job done?
For now, I think everyone is still using Anthropic and OpenAI because if you use a subscription you pay 1/40–1/50 of the API prices, and the models are good when they don’t nerf them, and they are also way cheaper than open models’ API prices.
The interesting thing will happen when they pull the plug and become economically smarter to stop using them. I regularly try alternatives to avoid being locked in and found GLM-5.3 as an orchestrator and GLM5.3 Flash + OMP and DeepSeek Flash as advisor to be able to get jobs done just fine. Space Bunny too was pretty great, which was probably MiniMax’s new model.
I think they are using an Uber like strategy but without the network effects that justify losing money for so long
dizhn
"With my OpenCode Go sub of $10/month, DeepSeek is basically unlimited. "
This article might have sat as draft for a few months. With the current deepseek pricing, the same membership lasts me a week at most even though I am using the free middle too and have Gemini pro+ultra.
It used to be that you could dump pocket change into the deepseek api and forget about it. Nowadays it'll make you notice real fast as the dollars pile up.
lmeyerov
GLM 5.3 Flash even more so... But yes :)
vietvu
This artcile is like 2 months too late?
option
Why freak out about good model? Better ones are coming too
robertlane0
Honestly for me the intelligence gap between DS 4.1 Flash and Muse Spark 1.3 makes Muse more worth it for me, especially on a $10 OpenCode Go sub, with the caveat that everything I use it on is open source which makes the fact that I'm sharing it with Meta a little moot because it's already published permissively on GitHub anyways.
pizza234
People have been raving since forever about Deepseek, but if one looks at the CoT, it's evident that it's way way stupider than frontier models (there's a reason why it's cheap). It's laughable to compare Deepseek 4.1 with Opus 5.5.
I've benchmarked, rigorously, deepseek-v4-flash for programming and personal use, and it is definitely less smart than Qwen3.8-flash-next (which in turn, is not terribly smart).
Local models are also really slow, unless one spends insane amounts of money.
Having said that, Qwen3.8-flash-next is an impressive evolution; it reaches the small versions of the frontier models (like Sonnet) - but again, it's massively slower and not 100% reliable (including: stability).
show comments
epolanski
Opencode Go for 8$/month feels like getting more intelligence and tokens than Claude in June 26 at the 200$ mark.
patchg
With two big players thinking of IPOs there is a lot of reasons to whistle on by.
potsandpans
I'm using it quite extensively in my PlayStation decompilation harness
MisterMunchkin
I had it make 25 different things today and it cost $0.70
It’s disgustingly good value. I find it capable of doing anything I want.
Obviously can’t use it at work, but for home projects it’s awesome.
Isn't Mimo 2.6 pro smarter and cheaper? Haiku 5.5 is smarter and cheaper. Luna is basically as smart and much cheaper.
nurettin
I didn't know it was so good. That was a gut punch. But I'm pretty sure the market already priced this in. And people are rightfully concerned about the owner of the data. Anything that is concerned with social, political or financial data goes out of your country to another one and is maybe even kept as a potential weapon.
ur-whale
> And to the self-hosters out there, the economics of 4.1 Flash mean self-hosting is not worth it. If saving money is your goal, you will never recoup the costs.
Self-hosting is, for most enterprises, absolutely not about economics but rather about data confidentiality.
And in that regard, yes, the open-source weight models, especially the chinese ones will eat the fat closed US model's lunch big time.
athrael-soju
Because it will be replaced within weeks?
ulfw
Because it should be obvious to anyone with a brain now that AI is s commodity product.
Today this leads a bit, tomorrow that. They're all interchangeable if we are being honest.
PunchyHamster
They are. That's what the push for regulations is
criley2
I feel like whoever wrote this doesn't use these models regularly. Deepseek v4.1 Flash is far from the pareto line. You can get the same performance for half the cost from Luna or Haiku 5.5 now, or you can get substantially improved performance at the same price with Sol 6.1 ~medium.
It did correctly make waves when it launched, but was quickly eclipsed by the deluge of american model releases, especially those competing on cost.
ltbarcly3
DeepSeek 4.1 Flash kindof sucks. I used it a bunch and it kindof sucks. I don't know if they are gaming benchmarks or what.
Luna is on par in benchmarks and my personal experience is Luna is better for what I do, and Luna is cheaper.
Comparing Deepseek 4.1 flash to Opus is just ludicrous.
Also, it is willing to do legitimate work I need done which other models flag as dangerous and refuse to do. (Software testing of a DHCP server to survive bad inputs.)
bitfilped
Because in two weeks someone will be asking why I'm not freaking out about AlphaDolphins 0.3 Zip and then in a month FrozenMonkey 2.5 Artic.
try-working
I have used over 40B tokens and spent over $800 on DeepSeek API over the past 30 days, mostly on V4.1 Flash.
It's good, and you can do most work with this. For complex software implementation you need to split your runs into various phases, build in verification, and use subagents so that work gets another audit and repair pass from the lead agent. You can do pretty much everything then. Frontier models can do without compelx workflows, that's the difference.
cactusplant7374
Because engineers are lusting for 1000 tokens per second. You can only achieve something like that with OpenAI.
pessimizer
I'm no expert, but it think that it's the pricing on GPT-6 Luna. I'm also guessing that it's been underpriced just for this reason. I also don't think it's all that great, but it's definitely very cheap.
If it's underpriced, it's a loss leader to sell the other models, so it actually can't be too good.
I really put these things through their paces because I use them to review and work with new abstract game rules and models, so they're always flying blind. Luna misses the obvious (and more importantly, the clearly explained) consistently. My second prompt is listing all of the points in its first response, and saying "No, it doesn't work like that." The third prompt is picking out the two or three suggestions it made after correcting itself on all of the original points and saying "That's how it already works." The fourth prompt is "Now that we're done going over the rules, can we start?"
I actually feel like 5.6 Luna seemed better.
sergiotapia
In my experience it just takes so much longer to arrive at "done" state for me. It thinks for soooooo long. I guess if you're running 12 sessions at once you don't really notice.
AIblemblio
No they can't.
And as long as I pay as little for claude opus 5.5 i do right now, i'm using it.
But yes i'm glad that we have alternatives.
samyar
it's good but not good enough
tonyhart7
it literally hallucinating a lot
I dont get why people says D4.1 flash is good
m3kw9
i thought 6.1sol copied the caching architecture so this isn't such a big deal no more
doctorpangloss
because it doesn't work very well?
if you have a legitimate coding application, it isn't very good. if you have some kind of inauthentic activity, which could be what it is trained for for all sorts of reasons...
show comments
xiaodai
cos it's shit.
yipinwong
to acctually answer the question, because it sucks to put it bluntly.
verdverm
Why would we freak out? The systems we use have always gotten better, faster, cheaper with time
cbeach
Honestly, I just don't trust the Chinese Communist Party having agentic access to my computer.
Every company in China has to abide by the 2017 National Intelligence Law: "supporting, assisting and cooperating" with state intelligence work, and keeping that cooperation secret. They have to hand prior knowledge of vulnerabilities to the state before public disclosure, in order that the state always has an exploit pipeline. No matter how ethical the company staff may be, they'll always be bound by law into being an arm of the Communist Party.
Agentic access is infinitely worse than chatbots. They can exfiltrate silently, target users, plant persistent malware, and be run by third parties through you.
You don't have to be a tin foil hat sinophobe to understand the dangers of being a Westerner granting CCP access to your files and network.
ByteDance staff accessed US journalists' TikTok data to hunt leakers (admitted in 2022). Volt Typhoon and Salt Typhoon were state operations pre-positioned in Western infrastructure and telecoms. Regulators in Italy and South Korea blocked DeepSeek's app over data handling, and analysts found its web client sending data to a China Mobile domain.
Please don't sacrifice security for cost and convenience.
show comments
oh_no
AA shows Luna at 1/4 the price, 1 point behind on intelligence matrix with a 38.
Haiku 5.5 is 23% cheaper with a 4 point intelligence lead.
I'm on subscription usage so I can't compare Flash 4.1 to them directly but the OP has his head up his ass if he thinks Opus 5.5 is the best point of comparison. Why is anyone using Opus if the new Haiku is indistinguishable /s
Just absolutely terrible post, admits to using Opus for review but claims its intelligence isn't needed, why aren't you using Haiku or Sonnet then?
distantsounds
because we've all figured out that AI is just a huge grift?
sroussey
Not comparing to gpt-6-luna which seems comparable and priced well.
wewewedxfgdf
You might also choose to pay money for a service that provides real value instead of actively choosing to support the Chinese deliberate effort to undermine this country.
The reason people aren’t freaking out is because most people are using heavily subsidized subscriptions.
I tried the cheapest provider on openrouter and burned through $50 in a few days. Quality was ok, seems slightly above Luna quality perhaps? But that $50 is 1/4 of my codex subscription where I could have burned that many tokens or more using Astra within my weekly reset.
This won’t last forever but as long as the frontier labs are subsidizing this heavily the open models won’t matter.
I'm paying for the heavily-discounted subscriptions, not the API rates. There isn't really a cost gap for me. DeepSeek doesn't have a subscription to compare to, but when I compared the GLM 5.3 usage I got from a $100/mo Z.ai subscription compared to Opus 5.5 on a $100/mo Claude subscription, there wasn't a big gap. And GLM 5.3 is very clearly not a frontier model (deepseek v4 seemed a lot closer, but I didn't use it enough to really say for my workloads).
I don't think those subscriptions nave negative contribution margins, either. I think we're seeing a lot of price discrimination by the big labs, and huge margins on their frontier models. The fact that they have been cutting prices to their second-biggest tier of models (Opus/Sol).
Open models catching up and collapsing these margins would worry me if I were a shareholder in the big labs, but as a user, I really doubt that the western labs have bigger environmental impact just because they have higher API costs, I think they have a ton of efficiencies they aren't sharing with customers yet because demand is so high.
Marketing.
Have you guys see how aggressive is the push for enterprise use by both OpenAI and Anthropic? I had friend from a non-tech industry in Asia telling me that their company was offered free trial of the enterprise version of Claude, with trainings and such.
On the other hand, DS and Z.ai, have zero to none marketing outside China. There is friction to use DS/GlM models and the ZDR is unclear, so most enterprise that has heavy AI usage hasn't move over yet. They would rather spent $200 for the peace of mind than to take the risk of being slam as a national traitor down the road (which again is another form of marketing by Big AI, trying to frame Chinese models as thiefs).
So, I don't think they are not freaking out, it's just that they are addressing different market segments and reacting to the situation differently.
Thank you for posting this, I didn't realize just how good 4.1 flash was. it far surpases gpt-luna in the limited testing I did: luna would take so many tokens that implementing a new feature would be nearly as expensive as if i just did it with sol 6.1, but 4.1 flash can do it for basically 1/6th the price, if not less. Obviously depends on the feature, but very very impressed.
I have been using DeepSeek 4.1 flash intensively for over a month. If I run it all day long it costs $1-2. Its fast. Previously I was always quickly running up to my Claude/Codex 5 hour window (on the $20/month plan). The cost savings of DeepSeek is real as shown in this article and I am using subsidized plans.
DeepSeek is horrible at grilling sessions (the /grill* skills to make technical decisions). It doesn't know how to explain things. Maybe the skill could be adjusted. It also doesn't come up with as good solutions as Opus/Sol.
What I use it for is
Previously I planned with Opus/Sol/Astra and then I used DeepSeek for coding, and then reviewed with Opus/Sol/Astra. With the cost improvements to Opus/Sol I am trying to use them for coding instead now so there will be less back and forth review needed.They are all working together in Pi using the extension @tintinweb/pi-subagents where my workflow skill is calling different subagents that use different models.
Luna is cost competitive, but doesn't score as well on intelligence. I do need the intelligence for most of what I use it for, so I am not motivated to use Luna. Haiku also doesn't seem like a competitive price/performance mix.
Oh man. v4.1-flash has been an sbolute game changer for us. We run all our Personal Assistants now on flash (thinking high) by default and it works incredibly well. There is really no need for basic agentic tasks that might require Kimi K.3 or GLM-5.3 levels.
Once its gets juicier, we let flash launch specialized subagents with specific models. GLM-5.3 for coding or Kimi K.3 for research and critique.
But as a main driver. I love flash. And it brought our bill down by A LOT :D
Because DeepSeek is not "a month or two" behind as claimed in the article.
These open models still did not beat February's Mythos / Fable 5.
DeepSeek 4.1 Flash is behind GPT 5.6 Sol, and that one is left in the dust by the excellent Opus 5.5.
Rumors say Anthropic is holding in reserve the big improvement, Fable 5.5, for the IPO.
It's plausible that open models are 6 - 12 months behind, and there is no "good enough". As long as progress doesn't slow down, leading labs have nothing to fear.
I have a pretty large, complex project I've been building with heavy AI use (new language + compiler). I was following a 'strong model as orchestrator launching cheap models as implementers' pattern, but I recently trialled just using Deepseek-V4.1-Flash as the model for both layers because of the cost savings (with mimo v2.6 flash on code review agents for some decorrelation).
I was previously using GLM-5.3 as the orchestrator, after switching to DS anecdotally there was an unnacceptable quality loss, mostly around not taking all the relevant context into account when making decisions, pulling new design out of thin air without discussion too often, and being way too wordy and rambly in documentation despite prompting to avoid it. There's a lot of docs, rulings, core concepts, design philosophy to uphold and DS was just not cutting it.
However, it's perfectly capable of being the sole agent for all of my well specced implementation tasks. I've gone back to GLM as the orchestrator.
I’ve played with DS4.1 Flash. It’s neat, no doubt. In my tests it does decently well against Sonnet 5.5 but uses way more tokens. That works out since it costs ~1/3rd Sonnet per task/PR according to my tracking.
However, I can easily burn ~$5/day if I use it as my “worker” (still using opus for planning and review) so call it ~$150/mo.
I could, if I were so inclined, get a second Claude Sub and have even more headroom (though I’m able to stay under my limits most of the time with my current setup). Also Claude gives me Artifacts, Web Search, and now even some API Credits.
I have no doubt the future is open weights and I can’t wait, literally, I can’t wait for them to catch up on intelligence or for hardware to run a decent model to be within my grasp. But until that comes to pass, I’ll keep using Anthropic.
Its simple. My company is very willing and able to pay ~$200/engineer/month for the best version of these tools. My company is even willing and able to pay as much as $500/engineer/month, but does not need to at this moment.
My company is not willing to pay $50/engineer/month for a cheaper version that is nearly as good. My company is also not willing to pay any amount for a product produced by China, even if it is hosted in the United States.
I had the same experience using ds 4.1 last couple of weeks. It’s insanely good for the price. I’m doing mostly web dev it excels at everything I throw at it. The pricing is ridiculous. I canceled my gpt subscription and haven’t looked back hope the pricing stays like that. I almost never need a better model. I still keep my Claude 20$ sub for now but I feel like one more iteration and I won’t need even that anymore I hardly use it
Everyone I talk to about Deepseek has a bad taste it feels like for CHinese models? I have a Kimi sub and use deepseek a lot at home. I've hit my limit on K3 many a time for Deepseek to come in and finish a job. So far no complaints. I feel it makes a lot of round trips by design but over all a good option. It's hard to compete though with an open sub out on codex/claude and just do everything in that subscription. Deepseek for me is when I run out of my subs. It's usually always coming behind a great session and fixing things so I don't give it a chance to try something novel on it's own.
People who have not used DeepSeek massively under estimate it's capability and how little it costs to run.
The question seems rhetorical but I think there are two reasons in some combination. First is it there is some awareness lag here. That lag can be on the producer and consumer side. Software enterprises are pretty slow to adopt new things and slow to try new things so they might only be aware of openai and Claude as options. Plus there are some scariness because deep seek is a Chinese model and therefore export restricted - never mind that there are American in European providers.
The other reason is more interesting. Maybe the frontier providers think that price performance is irrelevant in light of very powerful frontier models that can start the RSI loop and or a huge displacement of work and a winner take all economic situation. After all if frontier providers earn everyone's money then you won't have any money to spend on any model 100x cheaper or not.
What about MiMo v2.6 Pro? It’s throughput is slower by default (UltraSpeed is faster than DS4.1F) but is above the pareto line, and cheaper.
see https://artificialanalysis.ai/models/releases/comparisons?co...
While using DeepSeek v4.1 Flash I was architecting a system and I made a mistake of drawing the RPC boundaries at a wrong place that did cost me in so many ways.
I realized that mistake and guided DeepSeek where it should be.
Next I fired Fabble 5.5 set to high to check if the hype is real about Fabble. It exhausted 89% of quota and came up with NOTHING that DeepSeek hadn't flagged itself already in its notes.
I am 4.1 maxxing on commandcode GOAT Plan + api rates with oh my pi for the last 4 weeks, it's absolutely amazing and crazy fast, it's alright if it makes a mistake, I have enough time to iterate again, I have also added an advisor layer of mimo 2.6 pro which does make it a notch smarter. Getting haiku 5.5/sonnet5.5 to work on plans and letting 4.1 flash work through it is helping a ton too.
I am a big ChatGPT fan, all our team has ChatGPT Subs, but the TPS across all models including luna is just so damn slow.
Commandcode giving 60$ worth of Deepseek for 10$ is just genuinely goat.
And it never says no for cyber tasks so that's a big win
Because of several factors:
- Models are still very jagged. A superior model (e.g. Opus 5.5, Astra) may not be materially better in some tasks, but usually is for completely new modalities: computer use, game development, etc. People will always prefer using less jagged intelligence, because it allows them to do so much more
- Distilling intelligence from frontier models means that unless chinese labs manage to replicate the training regimes of OpenAI & Anthropic, they are always going to be behind a few months. That's a feature of training by distillation.
- China doesn't yet have the compute to compete at the frontier. Because they don't have access to top tier chips, their GWs are not equal to US GWs. Until they close the hardware gap, I think they will always be focused on competing on efficiency, as opposed to intelligence.
- OpenAI and Anthropic subscriptions are valuable of quality of intelligence + generous compute allowances
The problem with this low listed price per token is that, in reality, DeepSeek 4.1 Flash uses 10× more tokens than GPT-6.1 Sol for an equivalent task and delivers a lower-quality result. So there’s no real benefit to paying 10× less per token. Also, as someone else pointed out, OpenAI and Anthropic currently offer subsidized subscriptions for $100 or $200 a month that provide far more tokens than the API, so we should take advantage of them while that lasts.
This model finally got me off my Claude Max subscription. I’ve found it superior to Opus 5.5 in certain use cases, and certainly faster.
I’m convinced that I’ll have good enough inference on my laptop at reasonable speeds within the next year.
It's like the boy crying wolf with these models, Deepseek is not as good as Claude, I like using Gemini Flash Lite even, it is good enough for most of my crud app tasks, brownfield projects you don't need the strongest models, but greenfield projects with lack of context from indexed code and lack of domain expertise, I think the Opus models have been working for people
> With my OpenCode Go sub of $10/month, DeepSeek is basically unlimited
My OpenCode Go monthly window was scheduled to reset this morning. It was sitting at 22% used despite me using DeepSeek V4.1 Flash heavily as my implementation agent the past couple weeks (I use gpt-6.1-sol high for planning/orchestration).
I had 1.5 hours left so I fired up first 10, then 20, and finally 50 concurrent subagents all working on reverse engineering C code from an old PC game. They found over 100 new functions.
This is the first workload I've found that could make a dent in my sub. It got my 5 hour window to 85% used, but sadly my monthly was still only at about 35% when it reset. So that cost maybe $2.
I think because we're all just using it thinking we have found the "model for me" and never mentioning it to anyone because what would we say? It's good. It's a bit like the Logitech MX Master, as more and more people assumed they had found the ideal mouse for their purposes, it quietly became the professional standard through sheer adoption.
I've been waiting for something like this article for a month. The other AI companies are being left in the dust. I rip about 200 million tokens at least a day and spend maybe $3 with deepseek flash v4.1 for self hosted related coding tasks for ~20 projects i work on / maintain in parallel. Its a daily driver for sure. not even worth considering anything else, but maybe a locally run model on my GPU at the moment.
Actually OPUS sucks... Deepseek has its own flaw, but fares lot better in many instances. And yes it's far subsidised. If you are mindful of its peak and off peak hrs, you could make best use of its cost surge.
I really really love deepseek v4(and 4.1), but for everything I thrown at it, it felt like gpt 5.6 + 6 luna or terra do it faster, in way less tokens and in a way I like more. So even tho the price is (very) cheap, i found myself using it less just cause I don't want to wait on something I have to itteratee on with the model.
I've been using DS 4.1 Flash for a while now as my main workhorse. It's amazing. It's fast and it is also a GREAT debugger - it might not always find the best solution but it will definitely find the problem quickly.
I'm also the CEO of a new company - LunaRoute - so if you want private, fixed cost, we run our own server in the US type of DeepSeek 4.1 Flash check out https://www.lunaroute.com
If you want a trial, hit the contact us and write that you saw this post.
We also have GLM 5.3 (with vision!) and GLM 5.3 Flash - all included.
I think the interesting provider to cross-check this assumption with here is Meta, who is clearly freaking out, and is currently providing Muse 1.3 even cheaper so long as you are willing to share data with them
> This is like comparing big pharma with generic manufacturers who can skip the R&D. While these drugs are not 1 to 1 copies, we are still comparing apples to apples, but with a 90% price cut.
Weird! It's almost like distillation is bad for the long-term growth of the industry, just like generic manufacturers would be if they could release the generic versions of drugs 2 weeks after the original R&D completes.
Very strange sentence to include in an article after saying "I don't care at all about distillation" up at the top. Can the author not hear themselves?
I did a test on this a couple weeks ago. What I found was that the chinese models were far better than API rates, but about comparable on price vs the subsidized subcription model (chatgpt). Also it was my experience that codex completed tasks quicker.
That said, it's my best understanding that these american companies aren't profitable and will eventually raise rates (the old uber trick) so I'm keeping myself ready to switch when that day comes.
Why ?
Because I don't have the time and money to do an extensive benchmark of all major LLM, so when I had to select a LLM for my usage (which was not coding at first), I went to the most used one, chatgpt, because I knew if would be one of the best at the task.
I suspect it's the case for many if not most people.
Because most of the people use it through enterprise agreements and don't pay the bill? I run it for my own use cases and its pricing plus caching capabilities are hard to beat, cents for millions of tokens. https://substack.com/@rubenafo/note/c-332218129?r=26y5kn&utm...
Call me crazy but:
VRAM & Memory Requirements by Precision
• FP16 (Full Precision): Requires ~1,664 GB of VRAM (e.g., an 8x B300 288GB cluster).
• INT8 Quantization: Requires ~832 GB of VRAM (e.g., 8x H200 141GB).
• INT4 Quantization: Requires ~416 GB of VRAM (e.g., 8x A100 80GB)
VRAM aint cheap, Sam Altman ruined the cost of memory, Nvidia doesnt make enough consumer GPUs letting the market go insane over them, I still have friends on 1070s or 1070 TIs because GPUs have been severely overpriced for too long. I remember when a gaming PC was only $1000.
Even so why would anyone not sleep on a model they cannot run?
> It's like asking a math PhD to organize the files on your desktop.
I actually do use an agent harness to organize files on my desktop. They make a great fuzzy file renamer. Point it at a directory of disorganized files with names all over the place, give the directory layout and file name pattern you want it to have and it makes it happen.
Enterprise is not freaking out because DeepSeek 4.1 Flash does not actually occupy a spot on the Pareto frontier for non-coding enterprise workflows. We see this at my employer, focused on non-technical knowledge work. Luna 6 and now Haiku 5.5 are both very competitive if not better on all axes that we care about
> China is going to eat their lunch
No doubt about it, that's why their push for international regulation to the levels of nuclear inspections using the narrative of annihilation and apocalypse
Been using Flash 4.1 via the ante harness to blast through a GBA recomp. The ante team has pushed hard to make Flash 4.1 perform well under it. So far, I've maybe spent $10 over the last 3 days. Its a real workhorse and works much better in this harness
Deepseek and many other models are heavily benchmaxxed. They aren't genuinely as good as the best frontier models.
Because it’s not even that cheap? The author chose to only include Claude in their chart and ignored the fact that 6.1-sol and even more so luna can easily beat Deepseek on cost. Of course almost free cache used to be the main differentiator, raw token cost is deceptive since 4.1 just uses way more tokens than most other models
I actually have a fairly simple answer to that: if it doesn't come up in the list of LLMs that Cursor supports, it effectively doesn't exist.
I'm well aware that there's nearly infinite opportunities to yak shave "perfect" OpenRouter setups and some people appear to enjoy bouncing from IDE to IDE as though change costs aren't a thing, but I discovered that I genuinely like Cursor and at least right now it's insanely subsidized by Auto clearly defaulting to whatever Grok's most powerful model is.
I dropped my $200/month subscription to $20/month and stick to Auto for all but really important Plan tasks, and I have basically zero chance of using up my monthly credits even using it 6-10 hours some days.
Deepseek v4.1 flash is my baseline model to use at original provider's api. The issue is, for my personal use, it's cheap enough that i don't need it to go cheaper compare to the time I spent using it. And it's already a very capable model in dealing everyday simple things. For sure, for research level questions or large scale coding projects, I would want to use frontier model. But more and more daily tasks can be done now just using pi with deepseek api directly without thinking about much.
> Sure, they stole Claude's training, and Anthropic stole it from other people. I'm not getting into the whole who-owns-whose-data debate, because most developers aren't thinking like that. They're just trying to get the most bang for their buck.
Until they get laid off and suddenly discover their moral compass.
Honestly, having benchmarked 3 models Qwen3.8-Flash-Next, GLM5.3 Flash and DeepSeek 4.1 Flash. I absolutely do not understand the hype about Glm and DeepSeek. It's an improvement over older models, but Qwen is an actual opus replacement for me. It has been for last month.
But I run it locally. When I tried it on open router when my gpus were busy I must have gotten routed to some crappy providers, because it was pretty bad.
For me Glm and DeepSeek are nowhere near this Qwen model. I tried various harnesses including omp which I heard supposedly "makes DeepSeek 20 points better". The difference was in the noise (1 point). I run a bunch of benchmarks Terminal World 40, terminal bench 2.1,SWE Pro, GSO. Before those 3 there was no open model that scored more than 1 point on my subset of GSO. Glm scored 5, DeepSeek 3, but Qwen did 17 and opus 19.
Qwen is a small model so it fails on factual recall. But if you give it most of the info it needs it us amazing.
If you use open models you can also read the thinking tokens, I think this makes a huge difference because one tends to read them leading to better understanding of the end results and also giving opportunities to correct course if its hung up about ambiguity and stuff like that. It also helps a lot with not being distracted while the clanker is working, instead of going to youtube or reading up on whatever, you spend more time with the code and the thoughts behind it.
It stops the constant context switching.
DeepSeek 4.1 Flash 0910 is perfect for M5 Ultra 256GiB. Running it fully resident in RAM, prefill at ~2500 tok/s and decode at ~40 tok/s. Probably tons of room to improve from there.
we do. we moved 100% of our production traffic from Gemini to Deepseek. best decision we ever made.
DeepSeek Flash 4.1 is generally what I use as as my workhorse. I use ChatGPT web to create the outline, then have Deep Seek built it out. Works generally pretty great.
The hype is almost reverse now . With opus , sonnet and haiku 5.5 , why should I care about Chinese models that are so behind ? Also Jev seems like killed a big chunk of deepseek market too .
I personally used it for secondary research agents and classification . Now classification part is gone .
I've been on the API-only mentality for months and all the open models were definitely the stuff I loved the most. Kimi K2.5, 2.7 and Deepseek v4 were among my favorites, while sparingly using Opus or whatever OpenAI had for specific situations.
But ever since I've switched to one of the $100-tier subs, I can see why a lot of the people on it don't really discuss the open models often. I'd still use it especially when it comes to sensitive inputs, but for most work, what you get on OpenAI or Anthropic is really more than enough.
It really got even better when they also made their cheaper models up to par if not better than the open models.
I do think the crowd for open models are out there, especially when you see trillions of tokens running for them on OpenCode or OpenRouter leaderboards.
It’s still cheaper to get a subscription to an agent harness with frontier model backing for most people. DeepSeek really only becomes appealing to me when I want to do something the monthly subscription harnesses don’t support (API access) or more harshly charge quota for (like running coordinated sandboxed subagents). DeepSeek is then great because of the low price and price transparency, it just can’t compete with subsidized monthly access.
it's ok but not great compared to the commerical LLM's it's really not very good. it's benchmarks are clearly fake. even still we run it internally for a ton of workloads on our rack of gpu's
Probably because Luna is faster, cheaper, and approximately as smart
Thanks for sharing this.
I've been focusing on deep research related tasks for biotch and life science applications. The problem with this sort of task is that we need subagents to reason through multiple (potentially 100s or more) chains of knowledge/concept/evidence, so the token usage really explodes as complexity of the task and data expands. A typical task can cost me nearly a $1k overnight...
I've been testing out GLM5.3, but now I'm really tempted to try to Deepseek 4.1 flash too. Any chance you've benchmarked / compared the two?
DeepSeek isn't even on my mind. I use the frontier models and can get the best in the industry for a relatively cheap price.
I think the industry is freaking out about open weights models in general, if not specifically DeepSeek. That is why we're now on the ~4th call for pacing the frontier from the very people who, if they wanted to pace the frontier, would simply do it rather than asking Washington to get involved.
And it's why people like Hillary Clinton have been trotted out to talk about the dangers of open weights models -- I mean, does she even know what that phrase means? (I know HRC is a controversial figure and I'm not bringing her up for that purpose; I just note that she and other prominent retired politicians are now doing the circuit on Anthropic's behalf.)
Aren't they? Not about 4.1 Flash but open weight models in general.
The fight between Anthropic and OpenAI is just about who leads the duopoly. Chinese open weight models are threatening a duopoly in its entirety. That's why Ant and OAI are begging for regulation. A regulation that will hit Chinese models way harder than US models, finally giving them a moat.
Not sure "freaking out" is the word I would use, but it’s fairly obvious looking at OpenRouter usage that the price cuts on Luna a while back were in response to intense competition from dsv4.
So the industry is responding, where it matters. Which is on heavy API usage, not coding subs.
It's quite good and the labs are definitely scared, that's why they are lowering API prices and continuing to subsidize subscription plans aggressively to keep anyone from using this stuff.
That said, has anyone else found DS models to be unpolished? They seem to "lose their mind" a lot more often than Claude/GPT. I have tried all of the top open source models that came out over the past ~4 months or so and the GLM models (5.2, 5.3, and 5.3-Flash) have been much more usable for me. They feel like Opus but X months ago, DS feels like something else.
Comparing it with Anthropic, ANY model is cheaper and more effective. Don't get me wrong, Anthropic models are good, but they're always more expensive for the same task, even compared to other closed source models. At this point, I honestly think Anthropic has played the nasty trick to fine tune the models to be too verbose and charge us for more tokens.
Finally an article that gets the maths. Following the release of DSV4 preview where 1M context can fit in single digit GB of VRAM, frontier labs pricing just became stupidly expensive. Other Chinese labs and r/LocalLLaMA also can't compete on pricing.
And since then, there has been so many articles that made it to HN front page, and all of them didn't get it. They just went on and on about tokens generation. Most HN commenters didn't get it either, find-in-page for "cach" typically yield 2~3 responses. If I had a dime for every time this happened, I could have .. paid for 1B cached input tokens?
Anyway, DeepSeek still has to come up with a frontier model, and they almost did it with DSV4 Pro 0813, which is just slightly below GLM 5.3, but 30x cheaper. Unfortunately, the massive price hike happened just 3 days later.
DSV4.1 Flash is good, but not quite the same level. Much easy to self-host though, especially for serving a team of developers. Let's see what the next one can do.
People aren't freaking out because self-hosting isn't a solve issue and it requires a lot of capital. The average company won't go and build an $1M GPU cluster just to self-host any model. If that gets commoditized then they'll start freaking out.
I've been loving DS4.1 Flash, it's been one of my daily drivers since we started testing it internally.
We just launched it on our platform today (Mixlayer, https://mixlayer.com), promo code LAUNCH-DSV41F gets you some free credits if anyone wants to check it out.
Not sure why the author thinks Anthropic's models spend more power and water than DeepSeek, there's no evidence of that. Their pricing has more to do with premium perception and less to do with COGS.
Everyone does optimization of model serving because it's good for every player in there.
(Also the water consumption thing is not a real issue.)
Because the free api is mispriced, opus 5.5 on subscription is 10x cheaper. Also, the use case for flash models are
For coding, I rather spend 10x more than have even 1 bug but I'm only spending 2x 3x more if you count subscription cost.
This is a killer use case for something like customer support though.
Luna and Haiku 5.5 are just as cheap and much better.
Don't really understand people who say DS4 or 4.1 have frontier level performance. Anyone who has used it will tell you that it's a hallucination factory. The only thing it has going for it is deepseek's unique infrastructure that allows better cache retention, but the cost savings from that obviously come nowhere near how much subsidized usage you get out of even a $20 subscription with openai or anthropic.
For non-coding tasks it may be fine. But for coding, Opus 5.5 is just a completely another level than something like Deepseek 4.1 Flash.
Opus 5.5: TIME 9.3m COST / $1.99 / SCORE 99/100 https://jonclegg.github.io/pacman-bakeoff/#claude-opus-5-5
Deepseek 4.1 Flash: TIME 2.8m / COST $1.89 / SCORE 72/100 https://jonclegg.github.io/pacman-bakeoff/dev/#deepseek-v4.1...
Personally I have not used anything but 4.1 since it came out. I have a dataset that turns any model into pure hallucination machine, not only DeepSeek does not hallucinate, it builds new insights by combining its insights. It's not only cheap, its far better (at least for me)
I’m also a heavy DS 4.1 Flash user—especially when it’s available at those off‑peak prices, which is an awesome deal. And, like you said, it’s genuinely powerful and very snappy. I’m planning to evaluate the differences between `reasoning_effort` settings today.
No one will freak out until anyone can run frontier model in their own laptop ;)
Not only the model but the hardware it is running on. The Huawei chips they are using instead of NVIDIA are vastly less expensive. China is going to scale past the west. The idea that they are “only a few months behind” is today and many of us cannot even bring ourselves to admit it. The future is even more dramatic.
Mimo 2.6 Pro is even better.
I switched from DeepSeek 4.1 flash about 2 weeks ago for my Hermes sysadmin/coding agents and I am seeing better intelligence and lower overall spend.
https://artificialanalysis.ai/models/mimo-v2-6-pro
Because we are still in the phase where we can expect there to be something else to freak out to next week.
I have no idea if anthropic can actually make money at their subsidized subscription rates (you can easily hit your monthly cost in one 5 hour session if you price out the tokens through the api), but if subscriptions didn't exist, I do think everyone would be on deepseek 4.1 and just not look back.
Eventually the party will be over but for now it should be a no brainer first choice tool for some 90% of dev tasks
I have subscriptions to OpenAI and Claude but use DeepSeek 4.1 Flash for my coding agents.
It costs pennies and you got really great output.
The author is spot on.
I am rooting for the Chinese models to win so they can save us from our capitalist overlords.
However, the quality of the open-source Chinese models is terrible. They game or fake their benchmarks because in real-world usage they suck.
Because good enough in the writer's context is a pretty low standard. While many people regards GPT 6.1 Sol or Opus 5.5 as "incapable" in some cases.
Just try Opus 5.5 reminds me how Opus 4.5/4.6 astonishes me. Completely different, and GLM-5.3/Kimi3/DS-4.1 are still like Opus4.8 levels.
What would freaking out look like, or is this just a stupid bloggish title flourish?
Is OpenAI coming in $20B under a sign of "freaking out"?
They might be. They would delay public admission as long as possible, because public admission would make stocks go down.
That's what I thought until August. Used it for half a year almost exclusively. But after the API price increase I'm back at the Claude Pro and Kimi subscriptions.
this was beautiful to read thanks
Because GLM 5.3 Flash is even cheaper?
https://tangled.org/astrra.space/ds4-recipe is an incredibly cool writeup on making deepseek v4.1 flash run really fast.
It's a solid little model, and I appreciate DeepSeek's commitment to the bit in releasing a brand new pretrain, double the size, numerous architectural innovations as a ".1" release over the excellent DeepSeek V4 Flash.
Maybe this is a minor issue, but it seems like different providers on OpenRouter etc. have different quant settings. I imagine that that affects the perception of the model quite a bit.
I just keep getting more ambitious with what I use AI for; and that type of work needs surfing the frontier at all times. Because ultimately, many capabilities are not yet saturated
Companies pay for claude, devs not. IF i had to use AI from my own purse, i would never pay for a claude sub.
I've been using DeepSeek's API and have been happy with it, but I might look into OpenCode as well. Does OpenCode run a quantised version or use different providers from the official one?
Deepseek V4.1 Flash is a hidden gem, really. Not to mention, you can easily get it through many providers that offer zero data retention and consistent speeds above 200 tokens per second!
> Chasing the latest and greatest is silly
Said every week by someone who would never go back to using the model they had 6 months ago
We use Claude Fable to plan, and DeepSeek 4.1 Flash (hosted on DeepInfra) for everything else. Very cost effective.
Because I am using local Qwen. Flash Next is awesome! When hardware will not be a problem, everybody will be running their own models.
It’s obvious: they train on prompts and store data when using through their official API. And for third party it’s just… not good. That’s it.
> So why aren't the frontier labs freaking out right now?
They are. Isn’t this why they’re trying to get regulatory capture?
I’m partially assuming that there’s a bit of burnout.
I built mjolnir in large part so I could have Opus manage DeepSeek Flash subagents. It's phenomenal and extremely light on the Claude tokens. https://github.com/BrokkAi/mjolnir/
And yes, Opus is enough smarter than DSF that it's worth the extra steps. This ranking is from live tickets, no contamination: https://slopcop.com/power-ranking
Why isn't the author worried about sending her/his ideas to DeepSeek online (instead of hosting it and using it locally)?
One could argue that all this whole "Pacing the Frontier" bullshit is the industry freaking out regarding the danger of open models.
P.S. That's not to mean there aren't dangers regarding AI. I just don't trust the people making money from selling AI to manage those risks ethically instead of "protecting" us from those risks like pimps but with suits and good manners.
I was just using a bunch of models in Cursor to review a project. I went looking for DeepSeek and it was not one of the options.
Would be cool if they added it.
There was a moment 3 months back where the sentiment was that cheaper models were the way to go. Since then the pendulum has swung back.
There are too many commoditized models to count: GLM5.3 Flash, Kimi K2.8, Mimo V2.6, MiniMax M3.1...
Because it hallucinates a lot. There's no free lunch. Although the new architecture is a genuine move forward. The DeepSeek guys are really top notch researchers and devs.
My anecdotal experience is that I can’t even trust DS4Pro let alone Flash. I always have to have Sol reviewing the code.
Very interesting post!
> shrank the KV cache by roughly 437X
Can't you just say "shrank to 1/437th the size"? It's not that hard.
I like GLM 5.3 Flash, it seems good enough for coding features if you have a good structure and a good AGENTS.md
I don't know about "freaking out", but I'd say I'm having a good time here with DS 4.1 flash.
Because Haiku is more performant and costs less, even at API prices.
Because subscription plans are cheaper, only enterprises paying per tok pricing should be freaking out
America bans Chinese models sooner or later just the China banned American big tech
Once companies start realizing how much they're spending in AI to OpenAI and Anthropic for not much more employee output (the bottleneck is always the initiatives, not the code), they will look for cheaper options and they will eventually discover the chinese models.
My AI pilled clients who were early AI adopters are already there and they are looking for solutions to spend less.
What blows me away about this model is it's speed.
It's way faster than Opus or any of the GPT models.
I have a coding harness which is opencode plus a few skills relevant to my workflow. Deepseek 4.1 Flash does very well in this environment. I haven't noticed much difference quality wise compared to Opus 5, which I use in my day job as my employer pays for it (although I'm considering using DeepSeek here too given how cheap it is).
Is it known why unsloth seems to not have touched DeepSeek 4.1 Flash?
I would love to get them more better, it's good not a bad thing.
I mean, they kinda tried to regulatory-capture the market after trying to scare the public, possibly because those models will be a cheap option that gets the job done?
For now, I think everyone is still using Anthropic and OpenAI because if you use a subscription you pay 1/40–1/50 of the API prices, and the models are good when they don’t nerf them, and they are also way cheaper than open models’ API prices.
The interesting thing will happen when they pull the plug and become economically smarter to stop using them. I regularly try alternatives to avoid being locked in and found GLM-5.3 as an orchestrator and GLM5.3 Flash + OMP and DeepSeek Flash as advisor to be able to get jobs done just fine. Space Bunny too was pretty great, which was probably MiniMax’s new model.
I think they are using an Uber like strategy but without the network effects that justify losing money for so long
"With my OpenCode Go sub of $10/month, DeepSeek is basically unlimited. "
This article might have sat as draft for a few months. With the current deepseek pricing, the same membership lasts me a week at most even though I am using the free middle too and have Gemini pro+ultra.
It used to be that you could dump pocket change into the deepseek api and forget about it. Nowadays it'll make you notice real fast as the dollars pile up.
GLM 5.3 Flash even more so... But yes :)
This artcile is like 2 months too late?
Why freak out about good model? Better ones are coming too
Honestly for me the intelligence gap between DS 4.1 Flash and Muse Spark 1.3 makes Muse more worth it for me, especially on a $10 OpenCode Go sub, with the caveat that everything I use it on is open source which makes the fact that I'm sharing it with Meta a little moot because it's already published permissively on GitHub anyways.
People have been raving since forever about Deepseek, but if one looks at the CoT, it's evident that it's way way stupider than frontier models (there's a reason why it's cheap). It's laughable to compare Deepseek 4.1 with Opus 5.5.
I've benchmarked, rigorously, deepseek-v4-flash for programming and personal use, and it is definitely less smart than Qwen3.8-flash-next (which in turn, is not terribly smart).
Local models are also really slow, unless one spends insane amounts of money.
Having said that, Qwen3.8-flash-next is an impressive evolution; it reaches the small versions of the frontier models (like Sonnet) - but again, it's massively slower and not 100% reliable (including: stability).
Opencode Go for 8$/month feels like getting more intelligence and tokens than Claude in June 26 at the 200$ mark.
With two big players thinking of IPOs there is a lot of reasons to whistle on by.
I'm using it quite extensively in my PlayStation decompilation harness
I had it make 25 different things today and it cost $0.70
It’s disgustingly good value. I find it capable of doing anything I want.
Obviously can’t use it at work, but for home projects it’s awesome.
https://artificialanalysis.ai/#intelligence-comparison-tabs
Isn't Mimo 2.6 pro smarter and cheaper? Haiku 5.5 is smarter and cheaper. Luna is basically as smart and much cheaper.
I didn't know it was so good. That was a gut punch. But I'm pretty sure the market already priced this in. And people are rightfully concerned about the owner of the data. Anything that is concerned with social, political or financial data goes out of your country to another one and is maybe even kept as a potential weapon.
> And to the self-hosters out there, the economics of 4.1 Flash mean self-hosting is not worth it. If saving money is your goal, you will never recoup the costs.
Self-hosting is, for most enterprises, absolutely not about economics but rather about data confidentiality.
And in that regard, yes, the open-source weight models, especially the chinese ones will eat the fat closed US model's lunch big time.
Because it will be replaced within weeks?
Because it should be obvious to anyone with a brain now that AI is s commodity product. Today this leads a bit, tomorrow that. They're all interchangeable if we are being honest.
They are. That's what the push for regulations is
I feel like whoever wrote this doesn't use these models regularly. Deepseek v4.1 Flash is far from the pareto line. You can get the same performance for half the cost from Luna or Haiku 5.5 now, or you can get substantially improved performance at the same price with Sol 6.1 ~medium.
It did correctly make waves when it launched, but was quickly eclipsed by the deluge of american model releases, especially those competing on cost.
DeepSeek 4.1 Flash kindof sucks. I used it a bunch and it kindof sucks. I don't know if they are gaming benchmarks or what.
Luna is on par in benchmarks and my personal experience is Luna is better for what I do, and Luna is cheaper.
Comparing Deepseek 4.1 flash to Opus is just ludicrous.
https://artificialanalysis.ai/models/releases/comparisons?co...
Because GLM-5.3-Flash is both cheaper and better?
Also, it is willing to do legitimate work I need done which other models flag as dangerous and refuse to do. (Software testing of a DHCP server to survive bad inputs.)
Because in two weeks someone will be asking why I'm not freaking out about AlphaDolphins 0.3 Zip and then in a month FrozenMonkey 2.5 Artic.
I have used over 40B tokens and spent over $800 on DeepSeek API over the past 30 days, mostly on V4.1 Flash.
It's good, and you can do most work with this. For complex software implementation you need to split your runs into various phases, build in verification, and use subagents so that work gets another audit and repair pass from the lead agent. You can do pretty much everything then. Frontier models can do without compelx workflows, that's the difference.
Because engineers are lusting for 1000 tokens per second. You can only achieve something like that with OpenAI.
I'm no expert, but it think that it's the pricing on GPT-6 Luna. I'm also guessing that it's been underpriced just for this reason. I also don't think it's all that great, but it's definitely very cheap.
If it's underpriced, it's a loss leader to sell the other models, so it actually can't be too good.
I really put these things through their paces because I use them to review and work with new abstract game rules and models, so they're always flying blind. Luna misses the obvious (and more importantly, the clearly explained) consistently. My second prompt is listing all of the points in its first response, and saying "No, it doesn't work like that." The third prompt is picking out the two or three suggestions it made after correcting itself on all of the original points and saying "That's how it already works." The fourth prompt is "Now that we're done going over the rules, can we start?"
I actually feel like 5.6 Luna seemed better.
In my experience it just takes so much longer to arrive at "done" state for me. It thinks for soooooo long. I guess if you're running 12 sessions at once you don't really notice.
No they can't.
And as long as I pay as little for claude opus 5.5 i do right now, i'm using it.
But yes i'm glad that we have alternatives.
it's good but not good enough
it literally hallucinating a lot
I dont get why people says D4.1 flash is good
i thought 6.1sol copied the caching architecture so this isn't such a big deal no more
because it doesn't work very well?
if you have a legitimate coding application, it isn't very good. if you have some kind of inauthentic activity, which could be what it is trained for for all sorts of reasons...
cos it's shit.
to acctually answer the question, because it sucks to put it bluntly.
Why would we freak out? The systems we use have always gotten better, faster, cheaper with time
Honestly, I just don't trust the Chinese Communist Party having agentic access to my computer.
Every company in China has to abide by the 2017 National Intelligence Law: "supporting, assisting and cooperating" with state intelligence work, and keeping that cooperation secret. They have to hand prior knowledge of vulnerabilities to the state before public disclosure, in order that the state always has an exploit pipeline. No matter how ethical the company staff may be, they'll always be bound by law into being an arm of the Communist Party.
Agentic access is infinitely worse than chatbots. They can exfiltrate silently, target users, plant persistent malware, and be run by third parties through you.
You don't have to be a tin foil hat sinophobe to understand the dangers of being a Westerner granting CCP access to your files and network.
ByteDance staff accessed US journalists' TikTok data to hunt leakers (admitted in 2022). Volt Typhoon and Salt Typhoon were state operations pre-positioned in Western infrastructure and telecoms. Regulators in Italy and South Korea blocked DeepSeek's app over data handling, and analysts found its web client sending data to a China Mobile domain.
Please don't sacrifice security for cost and convenience.
AA shows Luna at 1/4 the price, 1 point behind on intelligence matrix with a 38.
Haiku 5.5 is 23% cheaper with a 4 point intelligence lead.
I'm on subscription usage so I can't compare Flash 4.1 to them directly but the OP has his head up his ass if he thinks Opus 5.5 is the best point of comparison. Why is anyone using Opus if the new Haiku is indistinguishable /s
Just absolutely terrible post, admits to using Opus for review but claims its intelligence isn't needed, why aren't you using Haiku or Sonnet then?
because we've all figured out that AI is just a huge grift?
Not comparing to gpt-6-luna which seems comparable and priced well.
You might also choose to pay money for a service that provides real value instead of actively choosing to support the Chinese deliberate effort to undermine this country.