I’m not sure if the underlying data is counting subscription use for Fable, which is where a lot of people are using it because token pricing is very expensive. I wouldn’t be surprised if this was counting enterprise token usage only. As rich as enterprise customers are, they’re not exactly willing to double the cost of SWE salaries on tokens.
Either way, a model used to solve the top 10% of problems that people use AI to solve for, being used 10% of the time … seems like it’s in a decent place.
I find Fable indispensable, and measurably better than alternative models, for complex feature development in an existing codebase. It’s the closest thing I’ve seen to nearly one-shotting features. Even still, I only use it for the hardest features and Opus 5 does a good enough job on the rest.
They've put themselves in a corner. Fable was too good and they gave it away with the $20 plan. It had to be a big step from Opus 4.8 to show progress, and Opus 4.8 is GREAT at coding in many different domains.
But they're getting killed on token cost. They have to get people paying more for tokens. So then they put Fable in the $200 plan and release Opus 5. I'm suspicious of Opus 5. It is mostly worse than 4.8. It _seems_ like they nerfed it to create more distance between it and Fable.
So what have most of us done? Stayed on Opus 4.8. The statistics bear this out. 4.8 still dominates.
Now they're stuck. If they take 4.8 away, everyone will riot. If they make Opus 5.x better than 4.8, they disincentivize everyone from moving to Fable and most importantly, paying more.
Really, all they can do is take the L for now and just let 4.8 be the apex of the $20 pro plan for the foreseeable future while they work like hell to make Fable THAT much better that it earns the $200 to $infinity that they really want everyone to pay.
And they've hobble Fable and Opus so hard with their safety guardrails, I ask innocuous questions and tasks and they get flagged so often I gave up on it. I can get all the work I need done in GPT5.5 or 5.6 without the hassle.
Might as well not be - I routinely get rate limited in a single review session.
I've honestly stopped using CC and moved to codex. Sol has it's warts but I've never once hit limits on a 100€ and I get a similar level of performance for what I'm doing.
I wouldn't mind bumping Fable to 200$ plan if it was actually better but between the insane caps, reverting to opus/sonnet randomly and having similar perf as OAI - I'm done with it.
Next step is to put 100$ into open router and try some western hosted open models with Pi when OAI starts pulling up prices.
I did not know that - but what happens then I get denied and I plead with the model that my usecase doesn't violate its TOS ? Edit prompt till I get it to pass ?
Reverting is annoying but being flagged for security questions while I'm doing code review is insane.
I still havent used it much, and they crippled Opus for what feels like no reason. Some suspect that Opus knowing theres a better model degrades it some. Shame.
Fable is just way too expensive and limited compared to GPT 5.6 Sol, and the only task that requires that level of intelligence is frontier scientific research. I use GPT/Codex primarily for coding and usually keep Claude on Sonnet 5 most of the time as I use Claude primarily to debug/brainstorm/make frontends as a supplement to GPT.
I do not agree. Fable is the only model I can leave unattended and give me results for part of my work which is just devops related tasking.
I can hand hold opus but I would rather just ask fable to do it and give me the result that I review and works. Opus will waste tokens and still require me to help nudge it in the right directions.
I think the next gen models from china will put us in a spot that the cost can plummet and I won’t need the Sota from anthropic
Sol routinely catches stuff in review that fable misses for me. It's impossible to compare them meaningfully because it's a complete dice roll - but in practice using both in my projects I can get work done with both, and Opus 5 is far more tedious.
But Fable security false positives and pricing just make it not worth compared to Sol IMO.
Most of the benchmarks have exceeded their usefulness. Opus 5 beats fable 5 on many of them. Anyone who has used both models will notice immediately that this doesn't translate to the real world. Opus 5 is nothing short of a regression from Opus 4.8. Fable is genuinely a great model so long as you don't trigger a guard rail and it downgrades.
Sol in my experience isn't significantly different than fable ignoring that Sol burns usage 10x faster but the end result is hard to differentiate.
GLM 5.3 is a hair behind these two.
An anecdote but not an original one from the people I talk to.
Yeah, it's pretty much apples to oranges, and I don't consider GPT and Claude to be interchangeable at all. From my anecdotal experience, GPTs generally codes more creatively and verbosely but Claudes tend to code more carefully and precisely, so the result is that GPTs generally finds more creative solutions to problems but also writes buggier code, which is why I converged on the setup of GPT/Codex for implementation and Claude for debugging, which feels more like a force multiplier than using each model individually.
I wish Fable were as good as you make it sound. A plan created by Fable is good, but in my case, it always contain issues caught only when it's reviewed again (whether by itself, Opus, Sol etc.). That's (almost) not different from plans created by Sol, GLM 5.3 etc. The one thing where it's genuinely better is the front-end, but then again it's far from perfect, it just needs less iterations.
>> but in my case, it always contain issues caught only when it's reviewed again
Yes but those issues will be much less severe with Fable-written plans than those written by lesser models. I know this because my workflows at both my regular job and my startup involve multi-step agent reviews via codified adversarial review skills. Fable as a reviewer will frequently find blocker-level issues with plans written by GPT 5.6 Sol, and sometimes with Opus 5. The opposite almost never happens. In fact I cannot remember the last time it happened.
I have a sneaking suspicion that someone at Google may be making the same bet, looking at the faster and faster Flash models which provide acceptable results to a lot of people (outside of coding).
For Anthropic. But not for the AI industry at large.
Cheaper, more powerful AI will continue to expand the bubble. Projects will get more ambitious. Everyone will build out their own custom little software. Code diversity expands and requires even more AI.
The bubble is _not_ on models becoming more intelligent and solving arc-agi-999.
They are already good enough at what they mechanically are.
You have to use the right harness, right verifiers (automatic where possible, human where not), etc much much more specific than a generic one like claude code or codex, and it will also be able to work within constraints and be the "proposer" of an imaginary optimisation problem and an excellent one at that. But you have to frame the task at hand in that manner or maybe even reorganise the task you do itself so it is more amenable to being framed that way. If you use it this way, it is _already_ massively economically useful. But it will take many years for it to actually be usable in that way, since you need DC capacity to come up first which is few years away and also well, massive organisations that have to integrate these will usually take many years to do so.
It is also useful albeit less so in cases like general SWE, where you still need a human in a loop for non-verifiable requirements, and also in other general usecases where information retrieval is too intractable and you need to carefully use LLMs as a component of the overall system.
I am not saying Fable or whatever the biggest models are are useless - they will certainly be useful for tasks at the frontier of the day - which is today complex exploits and open math problems, and well, tomorrow it could be something in biotech. But this is not what the entire bet is on at all - just automating day to day drudge at the tens of thousands of massive companies and governments we all know and love is more than enough. With the right training data (which _also_ is a bottleneck and takes time) you could even automate certain processes entirely. Sure, if we get a crazy medical innovation and end up saving trillions in healthcare great, but that's just a bonus.
None of this is to say that I think there is zero sketchy financial engineering going on
The bubble is based on the promise that these LLMs will cure cancer and find the solution to global warming. The pragmatic users of these tools (like you seem to be) are enjoying the subsidized use of the tools right now, but it's not a sustainable business model
Transformer models are quickly becoming a commodity, and I suspect in time we'll all be running them locally. Even now, you can run something pretty useful on a 16 GB graphics card, and I suspect a decade into the future, entry level hardware will be running better models than high-end graphics cards can run now, as entry-level hardware gets better and models get more efficient.
It doesn't mean hosted frontier models wont exist, they'll just be rare. It's no different than any other commodity market, for example most cars are cheap commodity models, with rare individuals buying expensive luxury cars and businesses buying expensive trucks and specialized equipment.
In my opinion, the big issue with Fable is that Claude Code cannot use it properly. I know, that sounds weird, but I've had Fable run down the wrong lane (and never stop) or give up and claim that something was impossible so many times (until I pointed at a GitHub repo that solves the "impossible" issue).
But a while ago, I had access to a Fable harness that just never gives up. And that verifies itself. It burned $100 in API tokens in 15 minutes ... but it succeeded for all the prompts where Fable + Claude had failed.
And I believe that's a real issue for Anthropic. Fable+Claude is not too expensive thanks to the subscription, but Claude severely nerfs Fable. To save money, I guess. Fable API + Custom Harness is a different class, it's so much better. But API tokens are so expensive, you're cheaper off hiring a freelancer.
As a small background, I have a local server and I've been trying out different models with different inference engines, quants, configurations etc... I'm also using Opus and Sol at work consistently. I've used AI since the first wave, first as a toy, then as a highly specific tool, last 6+ months as the primary LoC generator.
This is the first time I've felt, and I use the word *felt* since I don't have a suite of benchmarks or any sort of material approach towards comparing models, that Opus has declined in quality compared to before. Primarily I think its powers of deduction and understanding, even on xhigh, have become much worse. Before, being vague and providing a simple prompt would be enough, it could deduce and expand the details it needed, plus ask you clarifying questions, now this is no longer the case. A concrete, personal example, for a personal project, I've asked it to setup ssl over local IP. I didn't go into too much detail in the prompt as there are many approaches it could take and I didn't care too much to choose. It did horrible. The first thing it did was say the best lightweight approach is to add a reverse proxy. I'm like ok, makes sense. Then after asking it to proceed, it went and added a bunch of config to my golang service and didn't even setup a reverse proxy even when it said that is the way to go. It even said it didn't set it up lol. Then after I told it to do so it failed building the config in a way it was asked of it (support LAN IP and tailscale IP). Etc etc...
When Fable came out it was huge, the benchmarks told the story, and the story mostly matched the experience. It felt, again, intentionally saying felt, like it was miles ahead. Now benchmarks say that there are many models that are close, but in actual use Fable still *feels* much better. I think benchmaxxing the new open weights models is ruining the value of benchmarks, if they ever had any. When you actually put them to the test you see 500k tokens of reasoning with "Actually..." and "Wait..." in every third paragraph of their reasoning trace.
The price for Fable is definitely too much for any personal use now that it's no longer included in the subscription, and GLM 5.2, Deepseek Flash and Qwen 3.8 served locally or via cloud provide a lot, requiring a bit more babysitting though. Considering the price of Fable, my 5k USD Epyc server would pay itself off in less than a year if I used Fable or Opus in the same manner so at least for me the decision seems easy. And considering the point I'm poorly trying to make, that Opus doesn't feel like frontier anymore, this is probably the last month of my Claude subscription.
Anthropic’s issue is churn because of the peak verbosity vomit coming out of Opus 5/Mythos/Fable. What the hell did they train it on. The sane model is still Opus 4.6.
- Opus 4.6 was the last Opus generation that got a lot of use by Anthropic's own employees
- After that they primarily used Mythos internally
- 4.7, 4.8 and 5 were RLAIFd by Mythos "teachers"
- Hence why 4.6 is the last Opus gen who doesn't report back like a robot wanting to cover every potential hole another AI system would've spotted and criticized
- Hence why coding style in Opus 5 also gets criticized, not only behavior in CC
Hasn't it always been the premise that intelligence would get cheaper? To me, on the enterprise side, it seems like firms are finally getting the memo that, whether you are locked into the Ant/OAI ecosystem or not, you don't need the smartest, most expensive model to do every single task. This is a good thing for overall adoption. Whether that trickles down into regular user behavior, especially with subscription pricing, remains to be seen; even though I intellectually know I don't need Sol for a simple refactor, I am sometimes hesitant to choose Luna/Terra, as it's hard to accept using something positioned, even implicitly, as 'worse'. Remembering that the smaller models tend to be faster is what usually pushes me over the edge.
Anthropic in particular is much more compute-constrained than OpenAI and SpaceXAI and has relied on partnerships to provide inference. This reality factors into their pricing and usage limits (they started 'adjusting' the 5-hour limits during peak hours, and it certainly wasn't an upward adjustment). Accordingly, this is presumably what Anthropic wants, given they develop and release the lower-end models, suggest users use them in various nudges within their product, position the bigger/more expensive models as "For the most complex tasks" in their UIs, and so on.
> whether you are locked into the Ant/OAI ecosystem or not
I think the problem (for Ant/OAI) is that there is no sensible lockin or moat. LLMs are essentially interchangeable and stuff like a harness doesn't offer enough value on its own for someone to be locked into using one of them.
Now with the onslaught of the Chinese models that offer almost the same quality for much less money they have a very serious problem on how to proceed. Investors now might be looking through rose tinted glasses but their patience has its limits.
Yeah I mean if I run out of tokens every couple of hours and have to pause my work or shell out more money I’ll switch to other tools that don’t have this problem. Though they turned this down a bit it seems, I can work with Fable reasonably now and I enjoy it actually. I think they were just testing out how much they can raise the cost without users leaving when having the best model. I guess not much after all!
I've always wondered why everyone flocks to SV's latest darling company. Have we not learned from our history of glorifying these SV darlings that turn hostile?
Most of the people pushing this are just hoping that they can cash out before the hype pops and financial gravity crashes the party. Sam Altman recently claiming that the singularity is here is so stupid on its face he should just be treated as what he is, a huckster.
None of this stuff ever made any sense on what it was being sold initially. It was always insulting that the media and business leaders tried to argue that the tech could replace entire call centers or vast swaths of entire industries.
People keep arguing, but it will or it has based on extrapolating certain, reasonable use cases. Klarna has shut up about replacing call centers with bots because Markov chains with memory only can do so much.
Reads like: brilliant Carnegie Mellon University computer science grad struggles to find job where he is not replaced by cheap, inferior Indian labor that still gets the job done, even if it takes marginally longer.
No shit we're all paying 清冲 Flash to do the grunt work. Turns out though, paying 清冲 Flash a few more cents does exactly what Ivy Wasp Pro Mythical does. Crazy how that works.
I am because Claude has figured out I am a biochemist and therefore even asking Fable what the weather is gets me bumped back to Opus, sometimes even Opus 4.8 instead of 5. Kimi? GLM 5.3? DeepSeek? No such problem.
I can literally open a new chat with just "Hello" and it gets bumped.
Just cancelled my Claude subscription and used Ox Alpha Free through OpenCode (so that's one).
In my view, from the testing I did since Thursday, it's better than Fable. I had just finished a rather large task that Fable completed, including a /review and an /ultrareview.
Ox Alpha found bugs that Fable and Opus missed, and it continued to build things like a pro.
It does have issues with availability - but it's on a free promo right now. That also means that I don't know how much it would have cost if I had to pay API prices for it, which might not be cheaper than the subsidized small company / consumer usage, but for large companies paying API prices for either offering, the difference will be considerable.
I almost feel like there's some sort of paid campaign going on in Hacker News promoting the Chinese/open models. I feel like every day I'm hearing about how the frontier labs are dead but your experience is the same as mine. My company pays for Claude AND Codex but we never really use the open models for anything critical.
Doesn’t matter if your company pays for Claude, Anthropic and OpenAI valuations and expenditures commitment requires them to win the vast majority of the market to make economical sense. And the Chinese competition makes that very unlikely, to say the least. They likely won’t disappear fully but their “free” lunch as the AI darlings is done, on paper. Will be interesting to see how they adapt
It's absolutely a campaign. Every time anyone says anything about Claude, immediately and inorganically there's a bunch of people claiming to be biochemists who are constantly shut down by their work, people saying that they run out of tokens instantly even on Premium, people saying they get even better results on their GTX 4050, and people saying that Grok is better, or OpenAI is better. It's probably several different campaigns each run by different tranches of the competition. Reminds me of the good old days in which every criticism of bitcoin was immediately jumped on by nine or ten pretend Venezuelans who asserted that it was the only thing permitting their family to evade government currency controls.
http://archive.today/ZLojz
I’m not sure if the underlying data is counting subscription use for Fable, which is where a lot of people are using it because token pricing is very expensive. I wouldn’t be surprised if this was counting enterprise token usage only. As rich as enterprise customers are, they’re not exactly willing to double the cost of SWE salaries on tokens.
Either way, a model used to solve the top 10% of problems that people use AI to solve for, being used 10% of the time … seems like it’s in a decent place.
I find Fable indispensable, and measurably better than alternative models, for complex feature development in an existing codebase. It’s the closest thing I’ve seen to nearly one-shotting features. Even still, I only use it for the hardest features and Opus 5 does a good enough job on the rest.
They've put themselves in a corner. Fable was too good and they gave it away with the $20 plan. It had to be a big step from Opus 4.8 to show progress, and Opus 4.8 is GREAT at coding in many different domains.
But they're getting killed on token cost. They have to get people paying more for tokens. So then they put Fable in the $200 plan and release Opus 5. I'm suspicious of Opus 5. It is mostly worse than 4.8. It _seems_ like they nerfed it to create more distance between it and Fable.
So what have most of us done? Stayed on Opus 4.8. The statistics bear this out. 4.8 still dominates.
Now they're stuck. If they take 4.8 away, everyone will riot. If they make Opus 5.x better than 4.8, they disincentivize everyone from moving to Fable and most importantly, paying more.
Really, all they can do is take the L for now and just let 4.8 be the apex of the $20 pro plan for the foreseeable future while they work like hell to make Fable THAT much better that it earns the $200 to $infinity that they really want everyone to pay.
And they've hobble Fable and Opus so hard with their safety guardrails, I ask innocuous questions and tasks and they get flagged so often I gave up on it. I can get all the work I need done in GPT5.5 or 5.6 without the hassle.
I just kept a $20 plan going for use on my phone.
I had Fable bail on me because I used the word autopilot on a project. Changed to something unrelated, and off we went.
Fable is still on the $100 plan for me. (Maybe this is A/B testing or something.)
Might as well not be - I routinely get rate limited in a single review session.
I've honestly stopped using CC and moved to codex. Sol has it's warts but I've never once hit limits on a 100€ and I get a similar level of performance for what I'm doing.
I wouldn't mind bumping Fable to 200$ plan if it was actually better but between the insane caps, reverting to opus/sonnet randomly and having similar perf as OAI - I'm done with it.
Next step is to put 100$ into open router and try some western hosted open models with Pi when OAI starts pulling up prices.
You can turn off the reverting to opus with an option.
I did not know that - but what happens then I get denied and I plead with the model that my usecase doesn't violate its TOS ? Edit prompt till I get it to pass ?
Reverting is annoying but being flagged for security questions while I'm doing code review is insane.
I still havent used it much, and they crippled Opus for what feels like no reason. Some suspect that Opus knowing theres a better model degrades it some. Shame.
The problem is that they are not solving the problem they ought to be solving.
I don't want Shakespeare, I want Bob the builder.
Half of my work is telling claude how to behave. I'm pretty certain they have enough _data_ to realize people do the same thing time and time again.
Check this comment of mine for a better explanation of this: https://news.ycombinator.com/item?id=49413353
Fable is just way too expensive and limited compared to GPT 5.6 Sol, and the only task that requires that level of intelligence is frontier scientific research. I use GPT/Codex primarily for coding and usually keep Claude on Sonnet 5 most of the time as I use Claude primarily to debug/brainstorm/make frontends as a supplement to GPT.
I do not agree. Fable is the only model I can leave unattended and give me results for part of my work which is just devops related tasking.
I can hand hold opus but I would rather just ask fable to do it and give me the result that I review and works. Opus will waste tokens and still require me to help nudge it in the right directions.
I think the next gen models from china will put us in a spot that the cost can plummet and I won’t need the Sota from anthropic
I thought Sol was on par with Opus, so comparing it to Fable is apples and (very expensive) oranges?
Sol routinely catches stuff in review that fable misses for me. It's impossible to compare them meaningfully because it's a complete dice roll - but in practice using both in my projects I can get work done with both, and Opus 5 is far more tedious.
But Fable security false positives and pricing just make it not worth compared to Sol IMO.
I've used all three extensively.
Most of the benchmarks have exceeded their usefulness. Opus 5 beats fable 5 on many of them. Anyone who has used both models will notice immediately that this doesn't translate to the real world. Opus 5 is nothing short of a regression from Opus 4.8. Fable is genuinely a great model so long as you don't trigger a guard rail and it downgrades.
Sol in my experience isn't significantly different than fable ignoring that Sol burns usage 10x faster but the end result is hard to differentiate.
GLM 5.3 is a hair behind these two.
An anecdote but not an original one from the people I talk to.
Yeah, it's pretty much apples to oranges, and I don't consider GPT and Claude to be interchangeable at all. From my anecdotal experience, GPTs generally codes more creatively and verbosely but Claudes tend to code more carefully and precisely, so the result is that GPTs generally finds more creative solutions to problems but also writes buggier code, which is why I converged on the setup of GPT/Codex for implementation and Claude for debugging, which feels more like a force multiplier than using each model individually.
Fable is not a tool for the average user. It’s a professional tool for highly complex work.
I would compare it to a extremely high end $15k PC, or an expensive pro-grade video camera, or a freight train, or a …
I would say at least 95% of the global population will not encounter a situation once in their life where it would be actually useful/warranted.
I wish Fable were as good as you make it sound. A plan created by Fable is good, but in my case, it always contain issues caught only when it's reviewed again (whether by itself, Opus, Sol etc.). That's (almost) not different from plans created by Sol, GLM 5.3 etc. The one thing where it's genuinely better is the front-end, but then again it's far from perfect, it just needs less iterations.
>> but in my case, it always contain issues caught only when it's reviewed again
Yes but those issues will be much less severe with Fable-written plans than those written by lesser models. I know this because my workflows at both my regular job and my startup involve multi-step agent reviews via codified adversarial review skills. Fable as a reviewer will frequently find blocker-level issues with plans written by GPT 5.6 Sol, and sometimes with Opus 5. The opposite almost never happens. In fact I cannot remember the last time it happened.
Yeah this is a fair point. I only go to it when I have some big architectural problem I want its help in working out. Or a super nasty bug.
It works much better on regular software development e.g. for complex refactoring where cheaper models would produce a lot of garbage results.
I hate to be brusque but this is cope. GPT 5.6 Sol is just as good and cheaper
You have a $3 trillion bubble riding on this not being true.
Right? If we’ve already reached ‘good enough’ then there’s rough waters ahead.
I have a sneaking suspicion that someone at Google may be making the same bet, looking at the faster and faster Flash models which provide acceptable results to a lot of people (outside of coding).
For Anthropic. But not for the AI industry at large.
Cheaper, more powerful AI will continue to expand the bubble. Projects will get more ambitious. Everyone will build out their own custom little software. Code diversity expands and requires even more AI.
The bubble is _not_ on models becoming more intelligent and solving arc-agi-999.
They are already good enough at what they mechanically are.
You have to use the right harness, right verifiers (automatic where possible, human where not), etc much much more specific than a generic one like claude code or codex, and it will also be able to work within constraints and be the "proposer" of an imaginary optimisation problem and an excellent one at that. But you have to frame the task at hand in that manner or maybe even reorganise the task you do itself so it is more amenable to being framed that way. If you use it this way, it is _already_ massively economically useful. But it will take many years for it to actually be usable in that way, since you need DC capacity to come up first which is few years away and also well, massive organisations that have to integrate these will usually take many years to do so.
It is also useful albeit less so in cases like general SWE, where you still need a human in a loop for non-verifiable requirements, and also in other general usecases where information retrieval is too intractable and you need to carefully use LLMs as a component of the overall system.
I am not saying Fable or whatever the biggest models are are useless - they will certainly be useful for tasks at the frontier of the day - which is today complex exploits and open math problems, and well, tomorrow it could be something in biotech. But this is not what the entire bet is on at all - just automating day to day drudge at the tens of thousands of massive companies and governments we all know and love is more than enough. With the right training data (which _also_ is a bottleneck and takes time) you could even automate certain processes entirely. Sure, if we get a crazy medical innovation and end up saving trillions in healthcare great, but that's just a bonus.
None of this is to say that I think there is zero sketchy financial engineering going on
The bubble is based on the promise that these LLMs will cure cancer and find the solution to global warming. The pragmatic users of these tools (like you seem to be) are enjoying the subsidized use of the tools right now, but it's not a sustainable business model
Transformer models are quickly becoming a commodity, and I suspect in time we'll all be running them locally. Even now, you can run something pretty useful on a 16 GB graphics card, and I suspect a decade into the future, entry level hardware will be running better models than high-end graphics cards can run now, as entry-level hardware gets better and models get more efficient.
It doesn't mean hosted frontier models wont exist, they'll just be rare. It's no different than any other commodity market, for example most cars are cheap commodity models, with rare individuals buying expensive luxury cars and businesses buying expensive trucks and specialized equipment.
In my opinion, the big issue with Fable is that Claude Code cannot use it properly. I know, that sounds weird, but I've had Fable run down the wrong lane (and never stop) or give up and claim that something was impossible so many times (until I pointed at a GitHub repo that solves the "impossible" issue).
But a while ago, I had access to a Fable harness that just never gives up. And that verifies itself. It burned $100 in API tokens in 15 minutes ... but it succeeded for all the prompts where Fable + Claude had failed.
And I believe that's a real issue for Anthropic. Fable+Claude is not too expensive thanks to the subscription, but Claude severely nerfs Fable. To save money, I guess. Fable API + Custom Harness is a different class, it's so much better. But API tokens are so expensive, you're cheaper off hiring a freelancer.
As a small background, I have a local server and I've been trying out different models with different inference engines, quants, configurations etc... I'm also using Opus and Sol at work consistently. I've used AI since the first wave, first as a toy, then as a highly specific tool, last 6+ months as the primary LoC generator.
This is the first time I've felt, and I use the word *felt* since I don't have a suite of benchmarks or any sort of material approach towards comparing models, that Opus has declined in quality compared to before. Primarily I think its powers of deduction and understanding, even on xhigh, have become much worse. Before, being vague and providing a simple prompt would be enough, it could deduce and expand the details it needed, plus ask you clarifying questions, now this is no longer the case. A concrete, personal example, for a personal project, I've asked it to setup ssl over local IP. I didn't go into too much detail in the prompt as there are many approaches it could take and I didn't care too much to choose. It did horrible. The first thing it did was say the best lightweight approach is to add a reverse proxy. I'm like ok, makes sense. Then after asking it to proceed, it went and added a bunch of config to my golang service and didn't even setup a reverse proxy even when it said that is the way to go. It even said it didn't set it up lol. Then after I told it to do so it failed building the config in a way it was asked of it (support LAN IP and tailscale IP). Etc etc...
When Fable came out it was huge, the benchmarks told the story, and the story mostly matched the experience. It felt, again, intentionally saying felt, like it was miles ahead. Now benchmarks say that there are many models that are close, but in actual use Fable still *feels* much better. I think benchmaxxing the new open weights models is ruining the value of benchmarks, if they ever had any. When you actually put them to the test you see 500k tokens of reasoning with "Actually..." and "Wait..." in every third paragraph of their reasoning trace.
The price for Fable is definitely too much for any personal use now that it's no longer included in the subscription, and GLM 5.2, Deepseek Flash and Qwen 3.8 served locally or via cloud provide a lot, requiring a bit more babysitting though. Considering the price of Fable, my 5k USD Epyc server would pay itself off in less than a year if I used Fable or Opus in the same manner so at least for me the decision seems easy. And considering the point I'm poorly trying to make, that Opus doesn't feel like frontier anymore, this is probably the last month of my Claude subscription.
Anthropic’s issue is churn because of the peak verbosity vomit coming out of Opus 5/Mythos/Fable. What the hell did they train it on. The sane model is still Opus 4.6.
There's actually a tool called vomit [1] of all names to fix exactly what what is being discussed here.
[1] https://github.com/zachahn/vomit
https://x.com/wolframs91/status/2090159644849353058
What if:
- Opus 4.6 was the last Opus generation that got a lot of use by Anthropic's own employees
- After that they primarily used Mythos internally
- 4.7, 4.8 and 5 were RLAIFd by Mythos "teachers"
- Hence why 4.6 is the last Opus gen who doesn't report back like a robot wanting to cover every potential hole another AI system would've spotted and criticized
- Hence why coding style in Opus 5 also gets criticized, not only behavior in CC
Perhaps there is a sophisticated subtle poisoning attack that makes models behave like that?
Hasn't it always been the premise that intelligence would get cheaper? To me, on the enterprise side, it seems like firms are finally getting the memo that, whether you are locked into the Ant/OAI ecosystem or not, you don't need the smartest, most expensive model to do every single task. This is a good thing for overall adoption. Whether that trickles down into regular user behavior, especially with subscription pricing, remains to be seen; even though I intellectually know I don't need Sol for a simple refactor, I am sometimes hesitant to choose Luna/Terra, as it's hard to accept using something positioned, even implicitly, as 'worse'. Remembering that the smaller models tend to be faster is what usually pushes me over the edge.
Anthropic in particular is much more compute-constrained than OpenAI and SpaceXAI and has relied on partnerships to provide inference. This reality factors into their pricing and usage limits (they started 'adjusting' the 5-hour limits during peak hours, and it certainly wasn't an upward adjustment). Accordingly, this is presumably what Anthropic wants, given they develop and release the lower-end models, suggest users use them in various nudges within their product, position the bigger/more expensive models as "For the most complex tasks" in their UIs, and so on.
> whether you are locked into the Ant/OAI ecosystem or not
I think the problem (for Ant/OAI) is that there is no sensible lockin or moat. LLMs are essentially interchangeable and stuff like a harness doesn't offer enough value on its own for someone to be locked into using one of them.
Now with the onslaught of the Chinese models that offer almost the same quality for much less money they have a very serious problem on how to proceed. Investors now might be looking through rose tinted glasses but their patience has its limits.
Every software engineer in my company uses Claude code heavily. However we’ve never enabled Fable and only use Opus, Sonnet and Haiku.
Nobody has complained and seems like for every use case we have Opus is more than powerful enough, especially with Opus 5
Yeah I mean if I run out of tokens every couple of hours and have to pause my work or shell out more money I’ll switch to other tools that don’t have this problem. Though they turned this down a bit it seems, I can work with Fable reasonably now and I enjoy it actually. I think they were just testing out how much they can raise the cost without users leaving when having the best model. I guess not much after all!
Karma backed over their dogma.
I switched from Opus 4.7/4.8 to test Kimi K3 a few weeks back and the test hasn't finished; it's my daily driver now.
Given their general behavior and preference toward social engineering to scare the shit out of normal people...this seems fitting.
"Almost Nobody Is Using Anthropic’s Fable 5" - https://analyticsindiamag.com/ai-features/almost-nobody-is-u...
I've always wondered why everyone flocks to SV's latest darling company. Have we not learned from our history of glorifying these SV darlings that turn hostile?
I think the glib answer is, “Greed blinds all”.
Most of the people pushing this are just hoping that they can cash out before the hype pops and financial gravity crashes the party. Sam Altman recently claiming that the singularity is here is so stupid on its face he should just be treated as what he is, a huckster.
None of this stuff ever made any sense on what it was being sold initially. It was always insulting that the media and business leaders tried to argue that the tech could replace entire call centers or vast swaths of entire industries.
People keep arguing, but it will or it has based on extrapolating certain, reasonable use cases. Klarna has shut up about replacing call centers with bots because Markov chains with memory only can do so much.
Are people flocking to them? I see people buy the products but if you ask I think they are just about as hated in the big techs.
Reads like: brilliant Carnegie Mellon University computer science grad struggles to find job where he is not replaced by cheap, inferior Indian labor that still gets the job done, even if it takes marginally longer.
No shit we're all paying 清冲 Flash to do the grunt work. Turns out though, paying 清冲 Flash a few more cents does exactly what Ivy Wasp Pro Mythical does. Crazy how that works.
Absolutely zero people, rounding up generously, are replacing Fable with local or Chinese models.
I am because Claude has figured out I am a biochemist and therefore even asking Fable what the weather is gets me bumped back to Opus, sometimes even Opus 4.8 instead of 5. Kimi? GLM 5.3? DeepSeek? No such problem.
I can literally open a new chat with just "Hello" and it gets bumped.
Just cancelled my Claude subscription and used Ox Alpha Free through OpenCode (so that's one).
In my view, from the testing I did since Thursday, it's better than Fable. I had just finished a rather large task that Fable completed, including a /review and an /ultrareview.
Ox Alpha found bugs that Fable and Opus missed, and it continued to build things like a pro.
It does have issues with availability - but it's on a free promo right now. That also means that I don't know how much it would have cost if I had to pay API prices for it, which might not be cheaper than the subsidized small company / consumer usage, but for large companies paying API prices for either offering, the difference will be considerable.
Intriguing. Who is it? What is it? any ideas?
https://openrouter.ai/stealth/ox-alpha
Some people say it is a new version of Gemini Pro - this is based on some tweets from their employees.
Rumored to be mimo
I almost feel like there's some sort of paid campaign going on in Hacker News promoting the Chinese/open models. I feel like every day I'm hearing about how the frontier labs are dead but your experience is the same as mine. My company pays for Claude AND Codex but we never really use the open models for anything critical.
Doesn’t matter if your company pays for Claude, Anthropic and OpenAI valuations and expenditures commitment requires them to win the vast majority of the market to make economical sense. And the Chinese competition makes that very unlikely, to say the least. They likely won’t disappear fully but their “free” lunch as the AI darlings is done, on paper. Will be interesting to see how they adapt
It's absolutely a campaign. Every time anyone says anything about Claude, immediately and inorganically there's a bunch of people claiming to be biochemists who are constantly shut down by their work, people saying that they run out of tokens instantly even on Premium, people saying they get even better results on their GTX 4050, and people saying that Grok is better, or OpenAI is better. It's probably several different campaigns each run by different tranches of the competition. Reminds me of the good old days in which every criticism of bitcoin was immediately jumped on by nine or ten pretend Venezuelans who asserted that it was the only thing permitting their family to evade government currency controls.
"Everyone that disagrees with me is astroturfing" is not going to lead you to being in touch with reality.
It’s not wrong if it’s true