Hacker Newsnew | past | comments | ask | show | jobs | submit | eis's commentslogin

I know consumers hate the situation with ram and storage prices right now, as do I. But at least on the bright side all this AI investment has unlocked a lot of progress in a space that didn't see huge advancements in a good while. All these 10-20% improvements gen-on-gen have resulted in upgrade cycles of well over 5 years for many use cases in order to really feel like it's worth it. RAM capacities especially have felt near stagnant for a decade.

The progress is nice, but HBM appears to be impossible to reuse so... Kind of wasteful.

Chips at least had some aftermarket life in them.


The benchmark is very rudimentary. It does not test different levels/settings apart from its own -b 256/512 (does it affect decompression?), it doesn't measure compression time and memory usage. It does not specify parallel vs single-threaded (it mentions parallel on the one decoding number but what about the others?).

The lrzip test is interesting but it omits for example zstd and doesn't even have (de-)compression timings.

A lot more numbers are needed to present a fair and informative comparison.

I don't want this to be a swipe against bzip3, I only want to point out the presented benchmarks could be a lot better.


Making a wire transfer is easy, but can it be deducted from taxes? That's the tricky part.

No it can’t. Even across Europe it is not a given that the local tax authorities will easily accept the receipt. Charity law is national.

Debit cards can do chargebacks just as credit cards can do. That's a feature of Visa and Mastercard (and others). If your bank refuses, complain to the card network. Oh and consider changing banks, that's really unacceptable in 2026 and shows an anti customer stance.

They do, but require that an incident is raised claiming fraud, etc. I guess they want to avoid misuse of this feature.

This is completely normal. You can't just issue a chargeback because you feel like you want your money back after the fact. They want to know what went wrong. It could be fraud but it could be also misleading checkout experience and other reasons. The merchant is charged often a fee for a chargeback on the order of $15 and it might result in penalties.

So, if you feel like you've been charged unfairly (e.g. product not what was expected), in error, fraudulently etc. then chargeback is the route and you should provide a reason. It's a mechanism to protect consumers. If on the other hand you just want your money back without good reason then chargebacks are not the right tool. The consumer also has an obligation to pay if there was no fault on the merchant side.

Honestly, sounds like everything is as it should be?


Exactly this.

In both US and EU, chargebacks are a general consumer protection tool, implemented by card networks to make you less afraid of using them with potentially-untrustworthy merchants, even when not legally required.

They can be used for any kind of legitimate grievance, outright fraud being just one of them. If you book a hotel and the room is not up to the standards on their website, take pictures, take screenshots and chargeback.

They are not, however, a way to cancel a subscription you decide you no longer want (incidentally, neither is cancelling your card). Ask your lawyer, but this is actually illegal and considered fraud in many places. If you make a dumb choice, you have to live with the consequences or follow the contractual route.


> If you book a hotel and the room is not up to the standards on their website, take pictures, take screenshots and chargeback.

This is unhinged. Then I can chargeback at McDonald’s because the burger doesn’t look the same as the picture. I still ate it and you probably still stayed in the room.


However if you try to cancel the subscription the normal way, and it doesn't work, or there is no normal way, then you escalate to a chargeback and that's valid grounds.

Well yes, that's how it works for credit cards too. Chargebacks aren't a thing you can just do on a whim, they are a mechanism for the bank to punish fraudulent merchants. Attempting to get one against a merchant who is not defrauding you is a criminal offence. The bar for fraud is lower than you might think though. Any unauthorized charge, incorrect charge, barrier to cancelling a recurring charge, all counts.

It IS fraud. So report it and get your money back.

In the general Intelligence Index it scores exactly equal to Sol (61). In the Agentic Index it scores significantly lower than Sol (51 vs 58). In both it scores lower than Fable 5.1, Opus 5 and even Muse Spark 1.3.

Am I missing something or is this not looking too... stellar?


And yet I(and many others publicly on x) chucked Claude for codex back in June. Codex is my workhorse. The main reason I use codex is because of its language. I just got really tired of reading long weirdly worded prose that I had to fix with some skill(though some people like matt pocock and dex horothy have good ideas on this, buts it's just wasteful). This is triply bad for learning newer stuff because it goes up and down the abstraction layer on any topic like mad. One moment it would be explaining a high level detail and then cite contrasts and then point indirectly to an implementation detail as an example. It's idea of explaining more abstractly was also weird in a different way.

Fable was better but I don't have 1000$/day to spend on it.


Cost and reliability are the two reasons why we don't use Claude in our product. Getting close to one nine, that's not something one can build a reliable product upon. We now use OpenAI with Gemini fallback (or vice versa depending on use case). Personally I like Claude and have the 20x Max plan but even there I burned through the whole weekly quota with 3 prompts in less than a day using the new Fable 5.1 which is crazy. Now Opus 5 is down. These two issues are really testing my patience.

We just use bedrock in prod

That's a good point. Seems like Bedrock offers the same pricing while also providing an uptime SLA.

3.8 uses nearly twice as many tokens as 3.7. One might be inclined to think that they just increased the thinking budgets...

3.7 used 64M on high: https://artificialanalysis.ai/models/gemini-3-7-flash 3.8 used 120M on high: https://artificialanalysis.ai/models/gemini-3-8-flash

Even their own chart showed more than 2x higher cost compared to 3.7: https://storage.googleapis.com/gweb-uniblog-publish-prod/ima...


3.7 high and 3.8 medium are essentially the same on AA intelligence and cost. Output tokens on DeepSWE gives the same picture. So there might be something to it but they have done other things as well. At least the tokens are really fast.

i find deepswe not very reliable for instance it puts grok 4.6 xhigh over sol medium

3.8 uses nearly twice as many tokens as 3.7. One might be inclined to think that they just uppsed the thinking budgets...

3.7 used 64M on high: https://artificialanalysis.ai/models/gemini-3-7-flash 3.8 used 120M on high: https://artificialanalysis.ai/models/gemini-3-8-flash

Even their own chart showed more than 2x higher cost compared to 3.7: https://storage.googleapis.com/gweb-uniblog-publish-prod/ima...


According to Artificial Analysis, 5.1 cost 56% MORE than 5, $8523 vs $5455. Yes cache cost is lower but it was MUCH more verbose: 140M vs 83M output tokens.

This directly contradicts what Anthropic is presenting here. Yes it scores higher but that's to be expected from a new release. It's the opposite of what OpenAI has been doing which was reducing costs, increasing efficiency.

Fable 5: https://artificialanalysis.ai/models/claude-fable-5 Fable 5.1: https://artificialanalysis.ai/models/claude-fable-5-1


https://artificialanalysis.ai/models/claude-fable-5-1-high

On high it gets the same score as 5 with max effort while costing only half as much.


High, X-high, and Max are all on the $/intelligence Pareto

I am not sure if Fable is worth it, at least with version 5 vs Opus 5. Opus beats Fable in quite a few benchmarks and at twice the cost I just haven't seen it provide noticeably better results compared to Opus. Has anyone noticed big differences? I did notice Opus maybe making more mistakes repeatedly but I don't have hard numbers on this. I hope Fable 5.1 brings noticeable improvements. I am giving it a go now on my 20x Max plan on a problem that Opus 5 has struggled for more than week now and has made very slow progress with regular regressions on the way.

My impression is that Opus 5 can be very impressive if you don't care about maintenance, novel-length comments, and really having any input in general. But otherwise it's borderline-to-totally unusable. It seems tailor-made to not have a human in the loop.

Opus 5 is better than Fable 5 except for creative programming work (like graphics). Fable 5 might be slightly better but the token cost isn't worth it.

How are you evaluating the models?

On the Fable 5.2 eval summary, Opus 5 only beats Fable on SWE-bench multilingual and multimodal.

I primarily use the models via interactive sessions enhanced with custom tools and skill. For that Opus 5's benchmark superiority has not materialized into greater productivity and frankly has been quite a let down.

The outputs are too often unreadable even after adding recommended prompts. There is an ongoing problem with the heron_brook system prompt affecting orchestration. [1]

I've used Opus 4.8 since the second week Opus 5 was released.

Over this time, Fable 5 has been reliably fantastic. Both in planning and direct execution on complex changes across code and infra.

I'm a bit surprised that there doesn't (seem) to be a section discussing ~performance across different modalities. This system card and blog post too-often default to an API-based use case when the gander primarily experience Anthropic's models via interactive sessions.

I understand waiting to comment until Opus 5.1 is available and handles these problems, though I am hopeful that Anthropic will confront the elephant in the room on Opus 5's failure to delivery great interactive sessions and the widespread negative feedback on the release.

It would show the org is paying attention, taking steps to balance model evals between interactive and API use. Also, some empathy for customers that wasted time trying to make opus 5 work for them.

[1] https://github.com/anthropics/claude-code/issues/80988


> On the Fable 5.2 eval summary, Opus 5 only beats Fable on SWE-bench multilingual and multimodal.

Did you mean Fable 5.1, or do you have access to the next (unreleased) version of Fable?


I meant 5.1.

If you're doing something cutting edge like math or formally verifying algorithms, Opus 5 is a steaming pile of shit compared to Fable 5 and Sol 4.6, it makes countless stupid mistakes and is essentially incapable of completing the task without extreme hand-holding.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: