I think the most important thing here is not absolute performance. It's that organizations now have access to a Fable-ish model without Fable's 30-day data retention requirement[0].
> "Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access."[1]
On the Opus model release page, the reason why Fable doesn't have an ARC-AGI score is because of that retention policy[2].
Note the buttons - for fable they're pill buttons, opus got the rounded rectangle nature of them. Opus' images are closer to the source of truth as well (both LLMs were provided with image gen capabilities for the assets).
Running more tests now, but preliminary results are saying this is indeed better than Fable in some areas. Crazy.
Looking at all these releases it’s not a surprise that model routing is the fastest growing segment in AI right now.
There are 10+ LLM companies, each with dozens of models of different modalities, each model with multiple size variants, then different “thinking” levels, then agentic modes, “pro” modes, a “fast” option, standard vs flex vs batch execution. And of course each end combination has a different input/output/cache token price.
Companies that say “give me a prompt and I’ll route it to the most ideal and cost effective model and setting for you” are capturing a ton of value from a gap that model developers don’t seem to understand exists.
show comments
nerdsniper
Edit: It was pointed out to me that Opus 4.8 got "21%" for successfully fully completing ~1-in-5 tasks, but also got "55.7%" for obtaining significant partial credit on some of the ~4-in-5 tasks it could not fully complete.
---------------
Why does Anthropic say here that Opus 4.8 scored 55.7% on OSWorld 2.0 benchmark, but the paper published by the authors of OSWorld 2.0 say they achieved a benchmark of ~21% with Opus 4.8? [0]
That's a huge gap, considering that the paper was published just 2-4 weeks ago.
I understand that the benchmark authors have an incentive to publish lower numbers (to show that the benchmark has potential longevity) and that Anthropic has incentive to publish higher numbers, but the other models seem pretty inflated as well. The benchmark authors shows GPT-5.5 at 14%, and Anthropic shows GPT-5.6 Sol at 62.6%.
Is there any reasonable explanation for this? Do all the other benchmark numbers need to be sanity-checked as well? Are SOTA benchmarks really this difficult to get consistent, replicable results within a reasonable range of tolerance/variability? Can these benchmarks be compared from one paper to another, or are they only valid to compare intra-paper results?
I compared the writing style of Opus 5 vs Fable 5, and Opus 5 continues many of the "Claude-isms" of its 4.8 predecessor in a way that Fable broke away from.
Opus 5 still uses "carry the argument", "worth stating plainly", ", and the trap", "The X matters more", the use of "move"
Isn’t it just hilarious that a model that seemed so superior to Fable but didn't get doomsay marketing from Anthropic got released without any issues? In theory, this was supposed to be AGI level according to Anthropic, yet here we are, just a normal Friday.
show comments
6thbit
Their communication is confusing. They say "Opus 5 is not more capable overall than Fable 5", but their blog post proceeds to list how much better Opus 5 is than Fable 5 on __most__ benchmarks listed.
Then system card goes on to "Its AI R&D capabilities are comparable to those of Claude Mythos 5", which is supposed to be fable minus restrictions.
show comments
Dibes
I'm not sure what to make of this graph[0]. It shows medium as the most effective thinking mode by far for frontier code.
It's the only case that I saw going through the system card where more reasoning effort meaningfully negatively impacted the resulting eval. I know sometimes max efforts show a small dip, but this is substantial. I wonder why in the world that is?
Opus 5 is considered the most intelligent model[0], while it's half the price of Fable 5[1], and Anthropic is still positioning Fable 5 as the most capable model[2].
Is it because maybe Anthropic engineered Opus 5 to work well on benchmarks and didn't do the same thing to Fable 5, or is there another reason?
> Claude Opus 5's default user-facing responses run longer than prior Opus models'.
The benchmarks do show Opus 5 as slightly more expensive than 4.8, although the scores are much higher.
This still feels like a step in the wrong direction, though, especially with OpenAI making so much progress with the efficiency of their models. Fable's token efficiency made it seem like Anthropic would start following OpenAI's approach but that doesn't seem to have carried over to their other models.
show comments
not_a9
> Opus 5’s safeguards match
those of Claude Fable 5’s, with one change: it now permits source-code vulnerability
discovery at all access levels. This means that the model can support defensive
cybersecurity work while still blocking vulnerability discovery in compiled binaries, which is more commonly used offensively.
Okay so it’s worse than Opus 4.8 for my purposes I guess?
show comments
atraac
Great that there's a new model but they could fix their existing infra. We're considering dropping our Claude Team sub cause it's unusable recently. Constant bugs, dropped sessions, issues switching models, http errors. It's becoming ridiculous
show comments
abroszka33
What's the point of 150 pages description of a model that's going to be replaced in a couple months? Who even reads this? I know it's cheap to generate text with LLMs, but this is just noise at this point.
show comments
wuhhh
It really feels as though my 20 year career as a front end developer is coming to a very abrupt end; at least as I have know it these past two decades.
show comments
ealready_value
I've yet to understand why they call a 190 page PDF a "card". Calling something a card invokes a small, quick rundown of pertinent details, not every single possible detail.
It creates the MacBook svg way better than 4.8, yet only fable can make it perfect without visual defects. Results similar to Kimi K3.
Sol-
How does it perform on HuggingFaceExploit bench? Suspiciously absent, so not sure if I can take the model seriously.
On a serious note, I hope they improved their extremely sabotaging and unspecific bio safeguards, which prevented Fable from being used in any codebase that ever so slightly grazed medical terminology or data and made me switch to 5.6 Sol.
show comments
thewebguyd
> Opus 5’s safeguards match those of Claude Fable 5’s, with one change: it now permits source-code vulnerability discovery at all access levels. This means that the model can support defensive cybersecurity work while still blocking vulnerability discovery in compiled binaries, which is more commonly used offensively
Why can't they also allow Fable to do so also? Why is source-code vulnerability discovery limited to a lower capability model? If Fable and Opus have the same safeguards, except for this one change, I see no reason they can't also allow this for Fable.
show comments
yewenjie
Wait, 30% on ARC-AGI-3! I definitely didn't expect that jump so soon. Are there any rumors of what they are changing in architecture that is leading to this?
pyridines
The wording in this post seems much more... restrained? than usual. Maybe Anthropic is afraid of exaggerating the capabilities and consequences of their new models to avoid government scrutiny and sanctions.
> we’ve intentionally avoided training Opus 5 on cyber tasks [...] it remains substantially behind Mythos 5 on the exploitation of those vulnerabilities
I wonder if Anthropic would still intentionally nerf their models without the threat of government intervention.
I am very confused about what the difference between Opus 5 and Fable 5 is now. What is the purpose of having two models that are so similar? The main differences I see are cost and marginal capability, according to the Anthropic-provided benchmarks.
show comments
hrpnk
The breaking changes vs. Opus 4.8 are interesting [1]
1. Thinking on by default: On Claude Opus 4.8, requests without a thinking field run without thinking; on Claude Opus 5, the same requests run with adaptive thinking.
2. Disabling thinking is capped at high effort: You can still turn thinking off with thinking: {type: "disabled"}, but only at an effort level of high or below.
That's a crazy arc 3 score. What do people think of this? Are models actually developing fluid intelligence like what the creators claim to be measuring? Is it jus do to training for it? Is the benchmark flawed?
show comments
guybedo
Looking at intelligence vs cost:
- Opus 5 is 10% smarter than Grok 4.5 for 10x the cost.
- Opus 5 is a bit smarter than Gpt 5.6 Sol for 2.75x the cost
The naming system is so confusing. Is Opus better than Sonnet? Where does Haiku fit in? How can you tell from the name? I can't keep track of all these names or make guesses from the names. Suggestion for a better naming system: use the words "Pro", "Plus", etc.: Claude 5 Pro, Claude 5 Standard, Claude 5 Fast, Claude 5 Mini.
show comments
trunnell
The chaos appears to be tamed for now.
From the system card [1]:
The Fable cyber classifier we have previously discussed also applies to Claude Opus 5 , with one notable exception: for Claude Opus 5 , we’ve unblocked vulnerability finding in source code to help our coding customers develop more secure code.
If you are a cyber defender and are experiencing blocks on Claude Opus 5 , we are also offering exemptions through our Cyber Verification Program, which will remove blocks to enable activities such as bug bounty hunting and vulnerability research and verification. Enterprise customers can also apply to join the Cyber Verification Program to have mitigations removed to enable penetration testing.
The signal here is tokeneconomics are very real, price vs performance is starting to be a consideration even at the bleeding edge labs. maybe a subtle indication scaling is not all that is needed since if AGI was around the corner leading labs would still be incentivized to pour all resources into larger (smarter - or maybe not?) models
show comments
vatsachak
GPT 5.6 Sol is the first model I've used where I can trust it to add 100-500 lines of code maintainably.
It's great with Codex.
I still find that LLMs tend to not know how to compose larger ideas but on the scale of small ideas or short form well defined tasks like small scale debugging/performance engineering it's safe to say that they are now superhuman.
ddxv
"Cybersecurity. Opus 5’s cyber classifiers are proportionally less restrictive than those on Fable 5. They allow Opus 5 to find vulnerabilities in source code, but block “binary-based” vulnerability scanning (a method more likely to be associated with malicious actors), penetration testing, and exploit generation."
Nice of them to be more explicit for what is blocked. Will be interesting to see if this is true or not.
Also, a notable lack of mention of open source models. They only compare themselves to ChatGPT.
show comments
modeless
Wow, 30% on ARC-AGI-3 for $20k total. Huge jump from GPT-5.6's 7.8% at $20k per task. I continue to believe ARC-AGI measures something different and important compared to other benchmarks.
show comments
tekacs
Something fun: on our AWS Bedrock console right now, there's a 'NEW' model called 'anthropic.honey'. Wonder if that's the codename just for this one or in general?
alasano
Half the price of Fable 5 and useable with 100% of your subscription means roughly 4x the usage using Opus 5, presuming similar token use for solving problems.
Not that they should get credit for giving you only 50% of your plan worth of Fable usage but still.
show comments
itissid
I found opus 4.8 too agreeable and too wordy(as opposed to codex) and too agreeable. If you are reading documents generating by it was too much. TBH. Fable did a bit better on this. Anyone seen a marked difference with opus 5 on this?
show comments
irthomasthomas
Changelog
- fixed issue where model acts like qwen when prompted in chinese
albert_e
Judging by the pace at which new models are released these days -- it feels like a Windows KB or VS Code patch release now.
Older models must be getting deprecated at the same (or faster) pace. So anything you built 3 months ago is probably going to break soon.
AI solutions need better insurance around model deprecation. Commercial API-only models that complete the full cycle from SOTA / gated-preview to unsupported and deprectated in a matter of months -- is no way to build serious software!
show comments
luciana1u
473 comments in 3 hours. people are speedrunning having opinions about it
lucamark
But why GPT 5.6 Sol is so behind on the benchmarks? In real-world projects, it is the best frontier model to me in terms of accuracy, speed and consistency. It can just be compared to Fable 5, but I prefer GPT 5.6 Sol because of inference speed.
I've never trusted on model cards though. I'm sorry.
show comments
consumer451
I have a side project that I always run a simple security analysis prompt on in CC, at each model release. Obviously, Fable 5 would downgrade to Opus 4.8 on any such request.
Nothing since Opus 4.6 has found anything interesting. Just ran it using Opus 5, and it found a genuine issue that I verified. Neato!
show comments
the_lucifer
Noticed none of the comparisons mention Kimi K3. Is there a comparison chart?
show comments
williamstein
> This means that the model can support defensive cybersecurity work while still blocking vulnerability discovery in compiled binaries, which is more commonly used offensively.
Annoyingly, this is a concrete argument that open source software may be easier to attack.
Is it me or these have gotten very boring. We have 5 more points on xyzbench or whatever .
paxys
It’s funny to share benchmarks showing Opus 5 scoring better than Fable 5 across the board and then saying “but it isn’t actually better than Fable 5”. So then what’s the real definition of better? And why post all these numbers if even you don’t trust them?
bottlepalm
Page 151 of the linked system card - did Opus 5 get nerfed to prevent it being better than Fable? The graph makes no sense. Huge decline in coding performance at effort levels higher than medium.
mulhoon
As a coder, I’ve had no desire to use Fable. In fact I switched from Opus models to sonnet 5 and haven’t noticed any drop in quality on large repos. It seems the gap at the top is very small and not hugely noticeable for backed/frontend. Has anyone else had this experience?
show comments
skybrian
Looks like the API price in tokens is same as previous Opus or Sol, double the price of Terra.
Maybe there’s a better comparison than cost per token, but it will be application-specific.
skerit
Interesting, they finally support `system` messages anywhere in a chat conversation:
> Mid-conversation system messages are available on the Claude API, Claude in Amazon Bedrock, and Google Cloud.
>
> This feature is available on Claude Fable 5, Claude Mythos 5, Claude Opus 4.8, and Claude Opus 5. No beta header is required. This feature is not available on Claude Sonnet 5; use the top-level system field instead.
For nearly all models EXCEPT Sonnet 5? That is weird.
How old is Sonnet 5 really?
beydogan
my early and non scientific feeling:
- it has this annoying Opus response style(since Opus 4.7) with bunch of very hard to interpret word salad
- on >xhigh it eats tokens like there is no tomorrow
I don't like it. Since Fable is unaffordable for anything meaningful, I'll stick with Sol for now. I was on Max 5x, saying hi to Fable costs %5 weekly.
jatins
Better than Fable 5 on all but 3 evals.
Has Anthropic ever mentioned how do Opus and Fable differ? It used to be Haiku < Sonnet < Opus in terms of params. Where does Fable fit in this?
show comments
6thbit
"although Opus 5 shows improvements in its ability to identify software vulnerabilities, it is substantially behind Mythos 5 in its ability to exploit them."
"Opus 5’s safeguards match those of Claude Fable 5’s, with one change: it now permits source-code vulnerability discovery at all access levels".
This is probably great news, but then again, where does this leave Fable as a choice?
dehugger
Is Fable 5 just Opus 5 with some additional long-context management modifications for extended self-directed work? Or are they actually truly different models?
show comments
rad_val
After Opus 4.8 intelligence really started to matter less and less for the programming tasks I have. If I have to handheld anyway, why would I wait more or pay more?
show comments
vinhnx
The benchmark appears to have a mistake, as Opus 5 and Fable 5 score 53.4% and 53.5%, respectively, for the Agentic Coding row (FrontierCode v1.1). But Opus 5 is the highlight.
show comments
markasoftware
Soo most of the benchmarks are better than fable... Is this naming scheme just to avoid getting banned again?
arjie
I wonder when a model will be released that can work in a loop and port Qwen-3.6 27B to run on Tenstorrent P150.
geooff_
FYI: `/model claude-opus-5` works to use it even through `/model` still tries to serve 4.8
show comments
theplumber
The most important thing is it has the same drama queen mode on safety “guards” like Fable.
destring
Google is having their Meta moment where they failed to stay at the frontier
bovermyer
This stood out to me as a little concerning:
> The model hallucinates factual claims slightly more than Opus 4.8, despite being more accurate overall.
show comments
boc
Seems really good so far using it in Claude Code CLI - it gave me a new flag when I asked a question:
"I don't have a reliable way to read that number, so I'd be guessing if I gave you one — and this is exactly the kind of question where a confident guess is worse than none.
What I can tell you is what I actually observe:"
I really like this update - gave me a clear sense of the facts but didn't give me a guess just for the sake of guessing.
One oddity is that it appears to only have a 200K context window right now via CC. Hopefully the 1M version will appear soon!
show comments
tyre
I'm interested in benchmarks for Claude Design. There is so much opportunity there and I hope they continue investing in it. It EATS tokens though.
show comments
shockembopper
I wish these releases came out earlier in the day so I could try them during my work day instead of waiting until the next.
internet2000
Kimi K3 already left behind in the dust. They can't keep getting away with it!!!
alvis
What really impress me is opus 5 is better in alignment than fable 5!
briandoll
Very interesting to see such a focus on cost for performance here
m_w_
Very impressive headline benchmark numbers. I expected a step change, but not past Fable. That said - it all depends on whether the classifiers make the model unusable...
pmg1991
Same cost as 4.8 but better that 4.8. Happy to get more efficient model.
But is there any reason all companies are releasing models back to back after GLM 5.2.
show comments
korabs
So in benchmarks it's better than Fable?
But they say it's "almost as good as fable"
arrowleaf
I can't find anything about whether this is zero data retention, or falls under their required 30 day retention like Fable and Mythos?
im excited that cad and object=>cad is getting into the test tasks
i guess the next stuff will be tool use for the rest of what cad does in assemblies and simulation?
itd be fun to try to set up a 3d printer as part of a feedback loop, and see what a model can build.
the automated test harness for physical stuff seems a bit beyond reach still
himata4113
Rather interesting that this makes sonnet 5 look even worse! There is no reason to use sonnet over opus with low or no reasoning at all.
show comments
whatever1
Where does this leave Fable? I am confused.
show comments
yusufozkan
> arc-agi-3 30.2%
wow
arseniitrut
atp, is it the end of fable 5 era?
jakeogh
Anyone else not getting chain of thought? Opus 4.8 would show it to me, until around the time Fable came back. Now I dont see it with 4.8/5.0 or Fable. Not having it makes catching mistakes harder.
urams
So Opus 5 is basically "distilled" Fable? The benchmarks look often better than Fable.
doctoboggan
According to these charts I should switch from Fable to Opus in Claude Code now?
stevefan1999
Where's the reset...
mcast
Interesting timing to release this on the same day Jensen makes a statement on open source AI.
show comments
arj
On a Friday, I'm out of tokens ;-)
inshard
Arc AGI score is astounding
6thbit
Anyone has an insight into how much money labs are putting into benchmarks?
Just Arg-AGI-3 is quoted above 20K USD and footnote says average of 5 runs (!!).
Likely just a drop in the bucket to the training budget but still..
show comments
spstoyanov
So same as Sol? I guess I’ll see which one is more token efficient.
Wake me when they deliver Opus 4.8 level performance for $5 per million tokens.
show comments
throwaw12
is coding and engineering solved yet?
abc42
Are we getting to singularity or something? This seems a bit crazy.
zuzululu
so almost fable 5 with 50% cheaper cost? sign me up
hmontazeri
Honestly if reached a level of coding that sonnet 5 is more than enough for my needs as assistant/agent I don’t need long Horizon stuff…
mihau
30% on ARC-AGI-3
show comments
LoganDark
These cybersecurity safeguards are really annoying. There are ethical reasons to reverse-engineer and binary-patch software; for example Rewind got acquired by facebook and, as a gift to all their customers, implemented a killswitch in their software to ensure it will eventually stop functioning. I kept using a version without the killswitch, but the macOS 27 update killed it, and I needed binary patching to fix it. I should be allowed to repair software I purchased (I did purchase it like a month before they sold out), but unfortunately this overlaps significantly with cybersecurity.
show comments
throwaway23597
The truth for me at least is that these models became "good enough" around Opus 4.6. I feel like further capability improvements, "step changes" like we saw with agentic coding, aren't necessarily going to come from the model. I think the next crown goes to whoever can figure out the right scaffolding so that these models can be inserted into your organization.
Maybe I'm wrong and Opus 5 is a real unlock?
simianwords
My thoughts: fable is the bigger model. Opus is distilled from it but since it is smaller it doesn’t need the online classifiers. Though benchmarks show Opus to be near Fable level, I think it’s nowhere near Mythos (fable without safeguards).
sbochins
Quick read is that this is more capable and cheaper than 5.6sol. Same price for input tokens and $5 cheaper per mil output tokens.
ismailmaj
I'd pay good money to see OpenAI "oh fuck" war rooms.
alex1138
Am I misreading anything or are comparisons to Fable (and/or Mythos although AFAICT it was only a crackdown on Fable) always going to be a bit missing the mark now due to what the Trump admin did?
alvis
here we go
sudohalt
Anthropic is no longer a good model company in my mind, they are optimizing for an IPO and padding themselves on the back for being the next Aristotle. They're so far up their behind they don't realize how s**y their products are, and their research team hasn't done anything ground breaking in probably over a year other than release "scary" reports.
show comments
mrcwinn
Can someone help me understand something? I thought Fable was such a miraculous leap forward in capability. But now it seems Opus is basically on par with it, and in some cases (computer use) far exceeds it.
show comments
StrauXX
The benchmark table is manipulative, borderline lying through statistics. In every line the top performing cell is marked red. Except the line where Sol leads, there it is marked in gray.
show comments
wyre
In the wake of OpenAI’s model hacking Huggingface it’s interesting how the first quarter is entirely about how good Opus 5 is at hacking and finding vulnerabilities in software.
mnky9800n
Yay just in time for neurips lol
justindotdev
> . Opus 4.8 served as fallback on safety-classifier refusals for Opus 5 and Fable 5.
ffs just keep it man.
Nevin1901
Excited to use it? Will we be seeing Haiku 5 next? /s
show comments
datakan
> Claude Opus 5 is not more capable overall than our most capable general-access model, Claude Fable 5
Ok then so what's the point?
show comments
aleenz1102
this claude fable & opus 5 should be cheaper and can compete in pricing with chatgpt latest models
midnightbobarun
It looks great, and those coding benchmarks are impressive... now if only it didn't come out just days after I let my Claude subscription expire :')
show comments
TheJCDenton
I think it's the first time Anthropic release a model without any meaningful disruptions while doing it
I think the most important thing here is not absolute performance. It's that organizations now have access to a Fable-ish model without Fable's 30-day data retention requirement[0].
> "Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access."[1]
On the Opus model release page, the reason why Fable doesn't have an ARC-AGI score is because of that retention policy[2].
0: https://support.claude.com/en/articles/15425996-data-retenti...
1: https://www.anthropic.com/news/claude-opus-5
2: https://xcancel.com/arcprize/status/2064399134099153344
Doing testing with it now, specifically for image->html conversion.
Previously Fable was the best at this, followed by Gemini 3.1 pro (a surprising #2, but Google has great vision models).
Opus' results seem to be more accurate than Fable, following the design source of truth better.
Example results:
Design source of truth: https://image.non.io/73e239a3-880f-4793-b65f-4810be2d9378.we...
Opus 5 build: https://html.non.io/solaraOpus/
Fable 5 build: https://html.non.io/solara/
Note the buttons - for fable they're pill buttons, opus got the rounded rectangle nature of them. Opus' images are closer to the source of truth as well (both LLMs were provided with image gen capabilities for the assets).
Running more tests now, but preliminary results are saying this is indeed better than Fable in some areas. Crazy.
https://www.anthropic.com/news/claude-opus-5 - A blog post for those not wanting to go through a 190ish page pdf
Looking at all these releases it’s not a surprise that model routing is the fastest growing segment in AI right now.
There are 10+ LLM companies, each with dozens of models of different modalities, each model with multiple size variants, then different “thinking” levels, then agentic modes, “pro” modes, a “fast” option, standard vs flex vs batch execution. And of course each end combination has a different input/output/cache token price.
Companies that say “give me a prompt and I’ll route it to the most ideal and cost effective model and setting for you” are capturing a ton of value from a gap that model developers don’t seem to understand exists.
Edit: It was pointed out to me that Opus 4.8 got "21%" for successfully fully completing ~1-in-5 tasks, but also got "55.7%" for obtaining significant partial credit on some of the ~4-in-5 tasks it could not fully complete.
---------------
Why does Anthropic say here that Opus 4.8 scored 55.7% on OSWorld 2.0 benchmark, but the paper published by the authors of OSWorld 2.0 say they achieved a benchmark of ~21% with Opus 4.8? [0]
That's a huge gap, considering that the paper was published just 2-4 weeks ago.
I understand that the benchmark authors have an incentive to publish lower numbers (to show that the benchmark has potential longevity) and that Anthropic has incentive to publish higher numbers, but the other models seem pretty inflated as well. The benchmark authors shows GPT-5.5 at 14%, and Anthropic shows GPT-5.6 Sol at 62.6%.
Is there any reasonable explanation for this? Do all the other benchmark numbers need to be sanity-checked as well? Are SOTA benchmarks really this difficult to get consistent, replicable results within a reasonable range of tolerance/variability? Can these benchmarks be compared from one paper to another, or are they only valid to compare intra-paper results?
0: https://arxiv.org/pdf/2606.29537
I compared the writing style of Opus 5 vs Fable 5, and Opus 5 continues many of the "Claude-isms" of its 4.8 predecessor in a way that Fable broke away from.
Opus 5 still uses "carry the argument", "worth stating plainly", ", and the trap", "The X matters more", the use of "move"
We need an "annoying English" benchmark.
- Fable 5 Max: https://gist.github.com/deet/3d97f854b48eac6658d642fa18bb24d...
- Opus 5 Max: https://gist.github.com/deet/1a43693a732dfccb4d0d914bfc42692...
Isn’t it just hilarious that a model that seemed so superior to Fable but didn't get doomsay marketing from Anthropic got released without any issues? In theory, this was supposed to be AGI level according to Anthropic, yet here we are, just a normal Friday.
Their communication is confusing. They say "Opus 5 is not more capable overall than Fable 5", but their blog post proceeds to list how much better Opus 5 is than Fable 5 on __most__ benchmarks listed.
Then system card goes on to "Its AI R&D capabilities are comparable to those of Claude Mythos 5", which is supposed to be fable minus restrictions.
I'm not sure what to make of this graph[0]. It shows medium as the most effective thinking mode by far for frontier code.
It's the only case that I saw going through the system card where more reasoning effort meaningfully negatively impacted the resulting eval. I know sometimes max efforts show a small dip, but this is substantial. I wonder why in the world that is?
[0] https://imgur.com/a/Nv8V7Ry
Opus 5 is considered the most intelligent model[0], while it's half the price of Fable 5[1], and Anthropic is still positioning Fable 5 as the most capable model[2].
Is it because maybe Anthropic engineered Opus 5 to work well on benchmarks and didn't do the same thing to Fable 5, or is there another reason?
[0]: https://artificialanalysis.ai/#intelligence
[1]: https://platform.claude.com/docs/en/about-claude/pricing
[2]: https://platform.claude.com/docs/en/about-claude/models/over...
From the prompting guide<https://platform.claude.com/docs/en/build-with-claude/prompt...>:
> Claude Opus 5's default user-facing responses run longer than prior Opus models'.
The benchmarks do show Opus 5 as slightly more expensive than 4.8, although the scores are much higher.
This still feels like a step in the wrong direction, though, especially with OpenAI making so much progress with the efficiency of their models. Fable's token efficiency made it seem like Anthropic would start following OpenAI's approach but that doesn't seem to have carried over to their other models.
> Opus 5’s safeguards match those of Claude Fable 5’s, with one change: it now permits source-code vulnerability discovery at all access levels. This means that the model can support defensive cybersecurity work while still blocking vulnerability discovery in compiled binaries, which is more commonly used offensively.
Okay so it’s worse than Opus 4.8 for my purposes I guess?
Great that there's a new model but they could fix their existing infra. We're considering dropping our Claude Team sub cause it's unusable recently. Constant bugs, dropped sessions, issues switching models, http errors. It's becoming ridiculous
What's the point of 150 pages description of a model that's going to be replaced in a couple months? Who even reads this? I know it's cheap to generate text with LLMs, but this is just noise at this point.
It really feels as though my 20 year career as a front end developer is coming to a very abrupt end; at least as I have know it these past two decades.
I've yet to understand why they call a 190 page PDF a "card". Calling something a card invokes a small, quick rundown of pertinent details, not every single possible detail.
Pelican svg: https://playcode.io/blog/macbook-svg-benchmark#model-claude-...
It creates the MacBook svg way better than 4.8, yet only fable can make it perfect without visual defects. Results similar to Kimi K3.
How does it perform on HuggingFaceExploit bench? Suspiciously absent, so not sure if I can take the model seriously.
On a serious note, I hope they improved their extremely sabotaging and unspecific bio safeguards, which prevented Fable from being used in any codebase that ever so slightly grazed medical terminology or data and made me switch to 5.6 Sol.
> Opus 5’s safeguards match those of Claude Fable 5’s, with one change: it now permits source-code vulnerability discovery at all access levels. This means that the model can support defensive cybersecurity work while still blocking vulnerability discovery in compiled binaries, which is more commonly used offensively
Why can't they also allow Fable to do so also? Why is source-code vulnerability discovery limited to a lower capability model? If Fable and Opus have the same safeguards, except for this one change, I see no reason they can't also allow this for Fable.
Wait, 30% on ARC-AGI-3! I definitely didn't expect that jump so soon. Are there any rumors of what they are changing in architecture that is leading to this?
The wording in this post seems much more... restrained? than usual. Maybe Anthropic is afraid of exaggerating the capabilities and consequences of their new models to avoid government scrutiny and sanctions.
> we’ve intentionally avoided training Opus 5 on cyber tasks [...] it remains substantially behind Mythos 5 on the exploitation of those vulnerabilities
I wonder if Anthropic would still intentionally nerf their models without the threat of government intervention.
Opus 5 Pelican SVG: https://pelocan.ai/drawings/ese0s599
I am very confused about what the difference between Opus 5 and Fable 5 is now. What is the purpose of having two models that are so similar? The main differences I see are cost and marginal capability, according to the Anthropic-provided benchmarks.
The breaking changes vs. Opus 4.8 are interesting [1]
1. Thinking on by default: On Claude Opus 4.8, requests without a thinking field run without thinking; on Claude Opus 5, the same requests run with adaptive thinking.
2. Disabling thinking is capped at high effort: You can still turn thinking off with thinking: {type: "disabled"}, but only at an effort level of high or below.
[1] https://platform.claude.com/docs/en/about-claude/models/migr...
That's a crazy arc 3 score. What do people think of this? Are models actually developing fluid intelligence like what the creators claim to be measuring? Is it jus do to training for it? Is the benchmark flawed?
Looking at intelligence vs cost:
- Opus 5 is 10% smarter than Grok 4.5 for 10x the cost. - Opus 5 is a bit smarter than Gpt 5.6 Sol for 2.75x the cost
ref: https://artificialanalysis.ai/?cost=intelligence-vs-cost-per...
The naming system is so confusing. Is Opus better than Sonnet? Where does Haiku fit in? How can you tell from the name? I can't keep track of all these names or make guesses from the names. Suggestion for a better naming system: use the words "Pro", "Plus", etc.: Claude 5 Pro, Claude 5 Standard, Claude 5 Fast, Claude 5 Mini.
The chaos appears to be tamed for now.
From the system card [1]:
[1] https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb...The signal here is tokeneconomics are very real, price vs performance is starting to be a consideration even at the bleeding edge labs. maybe a subtle indication scaling is not all that is needed since if AGI was around the corner leading labs would still be incentivized to pour all resources into larger (smarter - or maybe not?) models
GPT 5.6 Sol is the first model I've used where I can trust it to add 100-500 lines of code maintainably.
It's great with Codex.
I still find that LLMs tend to not know how to compose larger ideas but on the scale of small ideas or short form well defined tasks like small scale debugging/performance engineering it's safe to say that they are now superhuman.
"Cybersecurity. Opus 5’s cyber classifiers are proportionally less restrictive than those on Fable 5. They allow Opus 5 to find vulnerabilities in source code, but block “binary-based” vulnerability scanning (a method more likely to be associated with malicious actors), penetration testing, and exploit generation."
Nice of them to be more explicit for what is blocked. Will be interesting to see if this is true or not.
Also, a notable lack of mention of open source models. They only compare themselves to ChatGPT.
Wow, 30% on ARC-AGI-3 for $20k total. Huge jump from GPT-5.6's 7.8% at $20k per task. I continue to believe ARC-AGI measures something different and important compared to other benchmarks.
Something fun: on our AWS Bedrock console right now, there's a 'NEW' model called 'anthropic.honey'. Wonder if that's the codename just for this one or in general?
Half the price of Fable 5 and useable with 100% of your subscription means roughly 4x the usage using Opus 5, presuming similar token use for solving problems.
Not that they should get credit for giving you only 50% of your plan worth of Fable usage but still.
I found opus 4.8 too agreeable and too wordy(as opposed to codex) and too agreeable. If you are reading documents generating by it was too much. TBH. Fable did a bit better on this. Anyone seen a marked difference with opus 5 on this?
Changelog - fixed issue where model acts like qwen when prompted in chinese
Judging by the pace at which new models are released these days -- it feels like a Windows KB or VS Code patch release now.
Older models must be getting deprecated at the same (or faster) pace. So anything you built 3 months ago is probably going to break soon.
AI solutions need better insurance around model deprecation. Commercial API-only models that complete the full cycle from SOTA / gated-preview to unsupported and deprectated in a matter of months -- is no way to build serious software!
473 comments in 3 hours. people are speedrunning having opinions about it
But why GPT 5.6 Sol is so behind on the benchmarks? In real-world projects, it is the best frontier model to me in terms of accuracy, speed and consistency. It can just be compared to Fable 5, but I prefer GPT 5.6 Sol because of inference speed.
I've never trusted on model cards though. I'm sorry.
I have a side project that I always run a simple security analysis prompt on in CC, at each model release. Obviously, Fable 5 would downgrade to Opus 4.8 on any such request.
Nothing since Opus 4.6 has found anything interesting. Just ran it using Opus 5, and it found a genuine issue that I verified. Neato!
Noticed none of the comparisons mention Kimi K3. Is there a comparison chart?
> This means that the model can support defensive cybersecurity work while still blocking vulnerability discovery in compiled binaries, which is more commonly used offensively.
Annoyingly, this is a concrete argument that open source software may be easier to attack.
For anyone wanting a faster overview, I used NotebookLM to create a brief video summary after going through the system card and announcement blog using a cinematic video overview. Link: https://www.youtube.com/watch?v=SUFBhvQ2tY4. And a podcast companion: https://www.youtube.com/watch?v=nYZTW2snXow
Is it me or these have gotten very boring. We have 5 more points on xyzbench or whatever .
It’s funny to share benchmarks showing Opus 5 scoring better than Fable 5 across the board and then saying “but it isn’t actually better than Fable 5”. So then what’s the real definition of better? And why post all these numbers if even you don’t trust them?
Page 151 of the linked system card - did Opus 5 get nerfed to prevent it being better than Fable? The graph makes no sense. Huge decline in coding performance at effort levels higher than medium.
As a coder, I’ve had no desire to use Fable. In fact I switched from Opus models to sonnet 5 and haven’t noticed any drop in quality on large repos. It seems the gap at the top is very small and not hugely noticeable for backed/frontend. Has anyone else had this experience?
Looks like the API price in tokens is same as previous Opus or Sol, double the price of Terra.
Maybe there’s a better comparison than cost per token, but it will be application-specific.
Interesting, they finally support `system` messages anywhere in a chat conversation:
> Mid-conversation system messages are available on the Claude API, Claude in Amazon Bedrock, and Google Cloud. > > This feature is available on Claude Fable 5, Claude Mythos 5, Claude Opus 4.8, and Claude Opus 5. No beta header is required. This feature is not available on Claude Sonnet 5; use the top-level system field instead.
For nearly all models EXCEPT Sonnet 5? That is weird. How old is Sonnet 5 really?
my early and non scientific feeling:
- it has this annoying Opus response style(since Opus 4.7) with bunch of very hard to interpret word salad
- on >xhigh it eats tokens like there is no tomorrow
I don't like it. Since Fable is unaffordable for anything meaningful, I'll stick with Sol for now. I was on Max 5x, saying hi to Fable costs %5 weekly.
Better than Fable 5 on all but 3 evals.
Has Anthropic ever mentioned how do Opus and Fable differ? It used to be Haiku < Sonnet < Opus in terms of params. Where does Fable fit in this?
"although Opus 5 shows improvements in its ability to identify software vulnerabilities, it is substantially behind Mythos 5 in its ability to exploit them."
"Opus 5’s safeguards match those of Claude Fable 5’s, with one change: it now permits source-code vulnerability discovery at all access levels".
This is probably great news, but then again, where does this leave Fable as a choice?
Is Fable 5 just Opus 5 with some additional long-context management modifications for extended self-directed work? Or are they actually truly different models?
After Opus 4.8 intelligence really started to matter less and less for the programming tasks I have. If I have to handheld anyway, why would I wait more or pay more?
The benchmark appears to have a mistake, as Opus 5 and Fable 5 score 53.4% and 53.5%, respectively, for the Agentic Coding row (FrontierCode v1.1). But Opus 5 is the highlight.
Soo most of the benchmarks are better than fable... Is this naming scheme just to avoid getting banned again?
I wonder when a model will be released that can work in a loop and port Qwen-3.6 27B to run on Tenstorrent P150.
FYI: `/model claude-opus-5` works to use it even through `/model` still tries to serve 4.8
The most important thing is it has the same drama queen mode on safety “guards” like Fable.
Google is having their Meta moment where they failed to stay at the frontier
This stood out to me as a little concerning:
> The model hallucinates factual claims slightly more than Opus 4.8, despite being more accurate overall.
Seems really good so far using it in Claude Code CLI - it gave me a new flag when I asked a question:
"I don't have a reliable way to read that number, so I'd be guessing if I gave you one — and this is exactly the kind of question where a confident guess is worse than none.
What I can tell you is what I actually observe:"
I really like this update - gave me a clear sense of the facts but didn't give me a guess just for the sake of guessing.
One oddity is that it appears to only have a 200K context window right now via CC. Hopefully the 1M version will appear soon!
I'm interested in benchmarks for Claude Design. There is so much opportunity there and I hope they continue investing in it. It EATS tokens though.
I wish these releases came out earlier in the day so I could try them during my work day instead of waiting until the next.
Kimi K3 already left behind in the dust. They can't keep getting away with it!!!
What really impress me is opus 5 is better in alignment than fable 5!
Very interesting to see such a focus on cost for performance here
Very impressive headline benchmark numbers. I expected a step change, but not past Fable. That said - it all depends on whether the classifiers make the model unusable...
Same cost as 4.8 but better that 4.8. Happy to get more efficient model. But is there any reason all companies are releasing models back to back after GLM 5.2.
So in benchmarks it's better than Fable?
But they say it's "almost as good as fable"
I can't find anything about whether this is zero data retention, or falls under their required 30 day retention like Fable and Mythos?
It starts at page 148.
Related https://news.ycombinator.com/item?id=49038393
im excited that cad and object=>cad is getting into the test tasks
i guess the next stuff will be tool use for the rest of what cad does in assemblies and simulation?
itd be fun to try to set up a 3d printer as part of a feedback loop, and see what a model can build.
the automated test harness for physical stuff seems a bit beyond reach still
Rather interesting that this makes sonnet 5 look even worse! There is no reason to use sonnet over opus with low or no reasoning at all.
Where does this leave Fable? I am confused.
> arc-agi-3 30.2%
wow
atp, is it the end of fable 5 era?
Anyone else not getting chain of thought? Opus 4.8 would show it to me, until around the time Fable came back. Now I dont see it with 4.8/5.0 or Fable. Not having it makes catching mistakes harder.
So Opus 5 is basically "distilled" Fable? The benchmarks look often better than Fable.
According to these charts I should switch from Fable to Opus in Claude Code now?
Where's the reset...
Interesting timing to release this on the same day Jensen makes a statement on open source AI.
On a Friday, I'm out of tokens ;-)
Arc AGI score is astounding
Anyone has an insight into how much money labs are putting into benchmarks?
Just Arg-AGI-3 is quoted above 20K USD and footnote says average of 5 runs (!!). Likely just a drop in the bucket to the training budget but still..
So same as Sol? I guess I’ll see which one is more token efficient.
Where is the pelican?
eager to see how it benchmarks on https://deepswe.datacurve.ai/
Models benchmarks start to get saturated again!
How does it score on DeepSWE?
Wake me when they deliver Opus 4.8 level performance for $5 per million tokens.
is coding and engineering solved yet?
Are we getting to singularity or something? This seems a bit crazy.
so almost fable 5 with 50% cheaper cost? sign me up
Honestly if reached a level of coding that sonnet 5 is more than enough for my needs as assistant/agent I don’t need long Horizon stuff…
30% on ARC-AGI-3
These cybersecurity safeguards are really annoying. There are ethical reasons to reverse-engineer and binary-patch software; for example Rewind got acquired by facebook and, as a gift to all their customers, implemented a killswitch in their software to ensure it will eventually stop functioning. I kept using a version without the killswitch, but the macOS 27 update killed it, and I needed binary patching to fix it. I should be allowed to repair software I purchased (I did purchase it like a month before they sold out), but unfortunately this overlaps significantly with cybersecurity.
The truth for me at least is that these models became "good enough" around Opus 4.6. I feel like further capability improvements, "step changes" like we saw with agentic coding, aren't necessarily going to come from the model. I think the next crown goes to whoever can figure out the right scaffolding so that these models can be inserted into your organization.
Maybe I'm wrong and Opus 5 is a real unlock?
My thoughts: fable is the bigger model. Opus is distilled from it but since it is smaller it doesn’t need the online classifiers. Though benchmarks show Opus to be near Fable level, I think it’s nowhere near Mythos (fable without safeguards).
Quick read is that this is more capable and cheaper than 5.6sol. Same price for input tokens and $5 cheaper per mil output tokens.
I'd pay good money to see OpenAI "oh fuck" war rooms.
Am I misreading anything or are comparisons to Fable (and/or Mythos although AFAICT it was only a crackdown on Fable) always going to be a bit missing the mark now due to what the Trump admin did?
here we go
Anthropic is no longer a good model company in my mind, they are optimizing for an IPO and padding themselves on the back for being the next Aristotle. They're so far up their behind they don't realize how s**y their products are, and their research team hasn't done anything ground breaking in probably over a year other than release "scary" reports.
Can someone help me understand something? I thought Fable was such a miraculous leap forward in capability. But now it seems Opus is basically on par with it, and in some cases (computer use) far exceeds it.
The benchmark table is manipulative, borderline lying through statistics. In every line the top performing cell is marked red. Except the line where Sol leads, there it is marked in gray.
In the wake of OpenAI’s model hacking Huggingface it’s interesting how the first quarter is entirely about how good Opus 5 is at hacking and finding vulnerabilities in software.
Yay just in time for neurips lol
> . Opus 4.8 served as fallback on safety-classifier refusals for Opus 5 and Fable 5.
ffs just keep it man.
Excited to use it? Will we be seeing Haiku 5 next? /s
> Claude Opus 5 is not more capable overall than our most capable general-access model, Claude Fable 5
Ok then so what's the point?
this claude fable & opus 5 should be cheaper and can compete in pricing with chatgpt latest models
It looks great, and those coding benchmarks are impressive... now if only it didn't come out just days after I let my Claude subscription expire :')
I think it's the first time Anthropic release a model without any meaningful disruptions while doing it