I know consumers hate the situation with ram and storage prices right now, as do I. But at least on the bright side all this AI investment has unlocked a lot of progress in a space that didn't see huge advancements in a good while. All these 10-20% improvements gen-on-gen have resulted in upgrade cycles of well over 5 years for many use cases in order to really feel like it's worth it. RAM capacities especially have felt near stagnant for a decade.
The benchmark is very rudimentary. It does not test different levels/settings apart from its own -b 256/512 (does it affect decompression?), it doesn't measure compression time and memory usage. It does not specify parallel vs single-threaded (it mentions parallel on the one decoding number but what about the others?).
The lrzip test is interesting but it omits for example zstd and doesn't even have (de-)compression timings.
A lot more numbers are needed to present a fair and informative comparison.
I don't want this to be a swipe against bzip3, I only want to point out the presented benchmarks could be a lot better.
Debit cards can do chargebacks just as credit cards can do. That's a feature of Visa and Mastercard (and others). If your bank refuses, complain to the card network. Oh and consider changing banks, that's really unacceptable in 2026 and shows an anti customer stance.
This is completely normal. You can't just issue a chargeback because you feel like you want your money back after the fact. They want to know what went wrong. It could be fraud but it could be also misleading checkout experience and other reasons. The merchant is charged often a fee for a chargeback on the order of $15 and it might result in penalties.
So, if you feel like you've been charged unfairly (e.g. product not what was expected), in error, fraudulently etc. then chargeback is the route and you should provide a reason. It's a mechanism to protect consumers. If on the other hand you just want your money back without good reason then chargebacks are not the right tool. The consumer also has an obligation to pay if there was no fault on the merchant side.
Honestly, sounds like everything is as it should be?
In both US and EU, chargebacks are a general consumer protection tool, implemented by card networks to make you less afraid of using them with potentially-untrustworthy merchants, even when not legally required.
They can be used for any kind of legitimate grievance, outright fraud being just one of them. If you book a hotel and the room is not up to the standards on their website, take pictures, take screenshots and chargeback.
They are not, however, a way to cancel a subscription you decide you no longer want (incidentally, neither is cancelling your card). Ask your lawyer, but this is actually illegal and considered fraud in many places. If you make a dumb choice, you have to live with the consequences or follow the contractual route.
> If you book a hotel and the room is not up to the standards on their website, take pictures, take screenshots and chargeback.
This is unhinged. Then I can chargeback at McDonald’s because the burger doesn’t look the same as the picture. I still ate it and you probably still stayed in the room.
However if you try to cancel the subscription the normal way, and it doesn't work, or there is no normal way, then you escalate to a chargeback and that's valid grounds.
Well yes, that's how it works for credit cards too. Chargebacks aren't a thing you can just do on a whim, they are a mechanism for the bank to punish fraudulent merchants. Attempting to get one against a merchant who is not defrauding you is a criminal offence. The bar for fraud is lower than you might think though. Any unauthorized charge, incorrect charge, barrier to cancelling a recurring charge, all counts.
In the general Intelligence Index it scores exactly equal to Sol (61). In the Agentic Index it scores significantly lower than Sol (51 vs 58).
In both it scores lower than Fable 5.1, Opus 5 and even Muse Spark 1.3.
Am I missing something or is this not looking too... stellar?
And yet I(and many others publicly on x) chucked Claude for codex back in June. Codex is my workhorse. The main reason I use codex is because of its language. I just got really tired of reading long weirdly worded prose that I had to fix with some skill(though some people like matt pocock and dex horothy have good ideas on this, buts it's just wasteful). This is triply bad for learning newer stuff because it goes up and down the abstraction layer on any topic like mad. One moment it would be explaining a high level detail and then cite contrasts and then point indirectly to an implementation detail as an example. It's idea of explaining more abstractly was also weird in a different way.
Fable was better but I don't have 1000$/day to spend on it.
Cost and reliability are the two reasons why we don't use Claude in our product. Getting close to one nine, that's not something one can build a reliable product upon. We now use OpenAI with Gemini fallback (or vice versa depending on use case).
Personally I like Claude and have the 20x Max plan but even there I burned through the whole weekly quota with 3 prompts in less than a day using the new Fable 5.1 which is crazy. Now Opus 5 is down. These two issues are really testing my patience.
3.7 high and 3.8 medium are essentially the same on AA intelligence and cost. Output tokens on DeepSWE gives the same picture. So there might be something to it but they have done other things as well. At least the tokens are really fast.
According to Artificial Analysis, 5.1 cost 56% MORE than 5, $8523 vs $5455. Yes cache cost is lower but it was MUCH more verbose: 140M vs 83M output tokens.
This directly contradicts what Anthropic is presenting here. Yes it scores higher but that's to be expected from a new release. It's the opposite of what OpenAI has been doing which was reducing costs, increasing efficiency.
I am not sure if Fable is worth it, at least with version 5 vs Opus 5. Opus beats Fable in quite a few benchmarks and at twice the cost I just haven't seen it provide noticeably better results compared to Opus. Has anyone noticed big differences? I did notice Opus maybe making more mistakes repeatedly but I don't have hard numbers on this. I hope Fable 5.1 brings noticeable improvements. I am giving it a go now on my 20x Max plan on a problem that Opus 5 has struggled for more than week now and has made very slow progress with regular regressions on the way.
My impression is that Opus 5 can be very impressive if you don't care about maintenance, novel-length comments, and really having any input in general. But otherwise it's borderline-to-totally unusable. It seems tailor-made to not have a human in the loop.
Opus 5 is better than Fable 5 except for creative programming work (like graphics). Fable 5 might be slightly better but the token cost isn't worth it.
On the Fable 5.2 eval summary, Opus 5 only beats Fable on SWE-bench multilingual and multimodal.
I primarily use the models via interactive sessions enhanced with custom tools and skill. For that Opus 5's benchmark superiority has not materialized into greater productivity and frankly has been quite a let down.
The outputs are too often unreadable even after adding recommended prompts. There is an ongoing problem with the heron_brook system prompt affecting orchestration. [1]
I've used Opus 4.8 since the second week Opus 5 was released.
Over this time, Fable 5 has been reliably fantastic. Both in planning and direct execution on complex changes across code and infra.
I'm a bit surprised that there doesn't (seem) to be a section discussing ~performance across different modalities. This system card and blog post too-often default to an API-based use case when the gander primarily experience Anthropic's models via interactive sessions.
I understand waiting to comment until Opus 5.1 is available and handles these problems, though I am hopeful that Anthropic will confront the elephant in the room on Opus 5's failure to delivery great interactive sessions and the widespread negative feedback on the release.
It would show the org is paying attention, taking steps to balance model evals between interactive and API use. Also, some empathy for customers that wasted time trying to make opus 5 work for them.
If you're doing something cutting edge like math or formally verifying algorithms, Opus 5 is a steaming pile of shit compared to Fable 5 and Sol 4.6, it makes countless stupid mistakes and is essentially incapable of completing the task without extreme hand-holding.
reply