Opus 5 is beating Fable because it never gives up
Anthropic's smaller model was never supposed to outscore the flagship it sits under. It does, and the reason tells you a lot about how these models actually work. Here is how I use the two together.
Claude Opus 5 is out, and on most of the benchmarks that matter it is beating Fable. Ten days ago this model was a rumor called Honeycomb flashing in Cursor's model picker. Now it is live, it costs a fraction of the price, and it is outscoring the flagship it sits under. Nobody saw that coming, including me. So I want to talk about why it happened, when you should use Fable versus Opus, and whether there is still room for GPT-5.6 Sol or open weight models like Kimi K3.
The smaller model was not supposed to win
Fable 5 has been out for about a month now, and it is truly amazing right up until you run out of usage on your subscription and have to start paying API prices. I wrote a whole column about that markup. Opus 5, on the other hand, is a smaller model. There was no reason to think a smaller model would beat the big one across most of the benchmarks. It absolutely does. Anthropic's launch numbers put it at or near the top of the frontier at half of Fable's price, and what I have seen since Friday mostly backs that up.
It also beats GPT-5.6 Sol across almost every benchmark, on both price and performance. Before this release, OpenAI could at least claim Sol was the cheaper version of Fable. Not so much anymore. I have been using Opus nonstop since it launched, and right now it is hard for me to see why anyone would pick Sol.
Fable is still the smarter model
And now you are thinking, AJ, you just said Opus beats Fable on most of the benchmarks. It does. But not because it is smarter. It wins because it never gives up. Fable is incredibly smart, and it knows it is smart, which is why it almost never doubts itself. It plows forward with whatever clever solution it has come up with, and to be fair, the solutions are usually clever. Opus doubts itself every step of the way, which is why it double checks and triple checks all of its work. Benchmarks do not hand out style points for brilliance. They score the finished answer, and the model that keeps checking finishes right more often.
How to use them together
If you are doing anything simple or straightforward, just use Opus. It ends up cheaper and it does a great job. If you are doing something complex and important, where you need the model to have good taste and judgment, where it has to reason through a lot of competing factors and priorities, use Fable on high effort. I have written before about letting Fable make its own calls, and that advice stands. But add one extra instruction: tell Fable it is in charge of design and review, and that all implementation should be handled by Opus subagents. Fable comes up with the plan. Fable gives the final sign off. Opus does the typing, because Opus is actually better at the implementation.
There is a practical reason to work this way too. On Claude subscriptions you can only put half your usage toward Fable. Working this split, only about 30 percent of mine goes there. Fable will typically spin up three to six subagents, all running Opus, and the heavy token burn lands on the cheaper model.
But what about Kimi K3?
But AJ, what about all these open weight models coming out of China that are supposedly just as smart and way cheaper? Kimi K3 is genuinely impressive. It is the largest open weight model ever released, it beats Claude on a few benchmarks, and it loses the vast majority to both Opus and Fable. It is also not that cheap, at least not yet. The hosted version runs $3 per million input tokens and $15 out, which is well over half of Opus pricing, and the weights themselves are not even public yet.
And that is all per token math. If you are on a Claude subscription and you actually use your allowance, you are paying something like 4 percent of the metered price. Against Kimi K3 that works out to roughly 8 percent of its cost. Fable is only available on the Max plan, but Opus is on Pro. Unless you are an enterprise that is forced to pay by the token, Claude is where the value is right now.
Opinion columns reflect the personal views of the author. Our reporting on the stories referenced here lives on the linked pages.