Kimi K3 vs Fable 5. Scored on your agent, not a leaderboard.
Create your the platform account
Spin up the platform this agent runs on, then clone the agent and point it at your account.
Walkthrough coming soon
A full video walkthrough of Kimi K3 vs Fable 5 Benchmark, start to finish, is on the way.
The agent recipe on GitHub
The Claude Code agent, wired and documented. Read the recipe, fork it, and run it in your own session today. No black box, every gate is in the open.
One command runs both models on the Claude Code skills you actually use, blind judges every output, and tells you when the cheap model is good enough.
What changed
The breakdown, gate by gate
What goes in, what comes out
The honest verdict. On the work that has a right answer, the cheap model is as good, at about a third of the price. Fable holds a small edge on customer facing polish. And the run's own routing policy routed nothing to Kimi, because the evidence was not strong enough to justify a switch.
The real lesson is about benchmarks. Judge once, with one model, in one order, and half your verdicts are noise. This one judges twice, with two families, across both orders, and admits a tie when it cannot call a winner.