This jumps around a lot based on the top throughput and latency of whatever provider happens to be best at the moment.
theropost
Anthropic is a bit nuts, I had $260 of credits on my max account for the extra usage the other night. It was expiring, so I figured I'll fire up an agentic swarm to deep dive and make some deep changes to some old cold bases.. literally 25 minutes or less, $260 burnt, it didn't get get into the implementation, just wrote a ton of useless plans for the most part. It really opened my eyes to what they expect to charge people.. wayyyy overpriced.
show comments
d2p
I clicked through and it showed Qwen at the top at 55.4 compared to 55.3 for Opus Max. I have a screenshot.
Then I clicked away and back, and now it goes Qwen second, with 58.4, to Opus Max at top with 59.2.
I have screenshots of both. The description above the chart is the same in boh cases:
> Artificial Analysis Agentic Index
> Represents the weighted average of agentic capabilities benchmarks in the Artificial Analysis Intelligence Index (GDPval-AA v2, ³-Banking)
What happened? How can the scores change so much in a few seconds?
show comments
eli
I believe it. It's extremely good at troubleshooting. I gave Qwen and Kimi K3 the same annoying, complicated, intermittent bug to track down. Kimi did a bit better in understanding the existing code, but Qwen built some diagnostic tools and did an excellent statistical analysis on the log data. Qwen got way closer to the truth.
I'm very much looking forward to their forthcoming smaller model Qwen 3.8 releases. A version that can easily run locally would be great.
Any benchmark showing Opus 5 as the best just loses credibility for me. Anyone who's actually used Opus 5 daily knows what I'm talking about.
show comments
petercooper
Hopefully this boils down to the smaller versions they've teased. In my experience, Qwen models are the closest to the "less knowledge, more intelligence" (yes, the two are hugely correlated!) ideal some tool-dependent tasks need. Even the 3.5 2B can be easily prompted to always lean on tools and not jump to false conclusions (although its actual coding skills are abysmal, as you'd expect).
show comments
esafak
It is also the most expensive open source frontier model, per task. After seeing this table I see no reason to use it.
drnick1
Why does an open weights model cost nearly the same as GPT5.6? $1.14 vs $1.23 on the cost index. Since you can't presumably run this on your own hardware given the model size and hence gain other things like privacy, I don't see any reason to move away from GPT at this rate.
show comments
syntaxing
I am so excited for Qwen 3.8 27B. It’s a shame how slow prefill (~3-400) is on a strix halo but it’s such a good model for agentic tasks.
show comments
camnora
Qwen is just crushing it overall. I regularly use 3.7-flash for everyday coding needs and it gets the job done.
SwellJoe
I find that surprising.
I've been trying it on several projects and have found it's pretty sloppy. It leaves stuff broken, doesn't reliably write tests to check its own work unless explicitly prompted, misunderstands the assignment, etc.
It is smart and reasonably quick but not reliable.
show comments
h14h
This has me hopeful for Qwen3.8-27B!
quirino
A couple days ago they had published an overall score of 53 for this model, but that was removed and today it returned with a score of 56.
I wasn't able to find an explanation from them. Anyone knows what happened?
show comments
aliljet
Is there a path to distill this model to do very specific things? Like a RAG strategy for a small (or even large) corpus?
show comments
looksjjhg
That took what 2 years? I love how the chip ban made them more efficient
steve-atx-7600
curious about methodology. ive seen them post results for claude/codex when they only ran over benchmarks 3 times per model...
atemerev
Well, that's the bad index then. It is barely usable in my opinion compared to other Chinese frontier models.
brcmthrowaway
Could someone like Apple be playing the long game - Good Enough(tm) intelligence will eventually fit in our pocket and homes?
show comments
sirbor
Qwen is the way to go
dyauspitr
It doesn’t even show up in the raw intelligence index, so how could it possibly be the best?
delduca
Go China!
ben8bit
Haven't tried this yet, but going to soon! I have to wonder what happened at Anthropic. We've cancelled our subscription in favor of OpenCode & Codex. Sol is just so good & OC goes so far for every $ spent. Claude's become a pain to work with - average output with an annoying personality. Who knew this would be an issue even a year ago? In any case, loving the stuff from the Chinese models!
Opus is still first in Intelligence Index followed by Fable, GPT 5.6, Kimi K3 then Qwen 3.8 max. https://artificialanalysis.ai/#intelligence
Our leaderboard combines Arena ELO, AA Intelligence index, latency and speed and goes: #1 Opus 5 #2 Kimi K3 #3 Qwen3.8 Max #4 GPT 5.6 Sol
Source: http://pellmell.ai/leaderboard.
This jumps around a lot based on the top throughput and latency of whatever provider happens to be best at the moment.
Anthropic is a bit nuts, I had $260 of credits on my max account for the extra usage the other night. It was expiring, so I figured I'll fire up an agentic swarm to deep dive and make some deep changes to some old cold bases.. literally 25 minutes or less, $260 burnt, it didn't get get into the implementation, just wrote a ton of useless plans for the most part. It really opened my eyes to what they expect to charge people.. wayyyy overpriced.
I clicked through and it showed Qwen at the top at 55.4 compared to 55.3 for Opus Max. I have a screenshot.
Then I clicked away and back, and now it goes Qwen second, with 58.4, to Opus Max at top with 59.2.
I have screenshots of both. The description above the chart is the same in boh cases:
> Artificial Analysis Agentic Index > Represents the weighted average of agentic capabilities benchmarks in the Artificial Analysis Intelligence Index (GDPval-AA v2, ³-Banking)
What happened? How can the scores change so much in a few seconds?
I believe it. It's extremely good at troubleshooting. I gave Qwen and Kimi K3 the same annoying, complicated, intermittent bug to track down. Kimi did a bit better in understanding the existing code, but Qwen built some diagnostic tools and did an excellent statistical analysis on the log data. Qwen got way closer to the truth.
I'm very much looking forward to their forthcoming smaller model Qwen 3.8 releases. A version that can easily run locally would be great.
Strange that the page https://artificialanalysis.ai/agents/coding-agents doesn't even mention "Qwen" once if it's now the "best" according to one of their one index?
Any benchmark showing Opus 5 as the best just loses credibility for me. Anyone who's actually used Opus 5 daily knows what I'm talking about.
Hopefully this boils down to the smaller versions they've teased. In my experience, Qwen models are the closest to the "less knowledge, more intelligence" (yes, the two are hugely correlated!) ideal some tool-dependent tasks need. Even the 3.5 2B can be easily prompted to always lean on tools and not jump to false conclusions (although its actual coding skills are abysmal, as you'd expect).
It is also the most expensive open source frontier model, per task. After seeing this table I see no reason to use it.
Why does an open weights model cost nearly the same as GPT5.6? $1.14 vs $1.23 on the cost index. Since you can't presumably run this on your own hardware given the model size and hence gain other things like privacy, I don't see any reason to move away from GPT at this rate.
I am so excited for Qwen 3.8 27B. It’s a shame how slow prefill (~3-400) is on a strix halo but it’s such a good model for agentic tasks.
Qwen is just crushing it overall. I regularly use 3.7-flash for everyday coding needs and it gets the job done.
I find that surprising.
I've been trying it on several projects and have found it's pretty sloppy. It leaves stuff broken, doesn't reliably write tests to check its own work unless explicitly prompted, misunderstands the assignment, etc.
It is smart and reasonably quick but not reliable.
This has me hopeful for Qwen3.8-27B!
A couple days ago they had published an overall score of 53 for this model, but that was removed and today it returned with a score of 56.
I wasn't able to find an explanation from them. Anyone knows what happened?
Is there a path to distill this model to do very specific things? Like a RAG strategy for a small (or even large) corpus?
That took what 2 years? I love how the chip ban made them more efficient
curious about methodology. ive seen them post results for claude/codex when they only ran over benchmarks 3 times per model...
Well, that's the bad index then. It is barely usable in my opinion compared to other Chinese frontier models.
Could someone like Apple be playing the long game - Good Enough(tm) intelligence will eventually fit in our pocket and homes?
Qwen is the way to go
It doesn’t even show up in the raw intelligence index, so how could it possibly be the best?
Go China!
Haven't tried this yet, but going to soon! I have to wonder what happened at Anthropic. We've cancelled our subscription in favor of OpenCode & Codex. Sol is just so good & OC goes so far for every $ spent. Claude's become a pain to work with - average output with an annoying personality. Who knew this would be an issue even a year ago? In any case, loving the stuff from the Chinese models!