← writing / article

The leaderboard moved again: K3, Fable 5, Sol, and the K3.1 nobody has announced

Four frontier models shipped in a month and a fifth is already rumored. The teams that win are not the ones who pick right — they're the ones who stop picking.

27 Jul 20265 min readai · leadership · strategy

Four frontier models, one month

I keep a note of which model my teams treat as the default, and this month I edited it three times. Anthropic shipped Claude Fable 5, OpenAI put GPT-5.6 Sol into a restricted preview, Google rebuilt Gemini from scratch and aimed 3.5 Pro at a mid-July release, and then Moonshot dropped a 2.8-trillion-parameter open-weight model that beat Fable 5 on the Frontend Code Arena on its first day. Before I had finished reading K3’s benchmark table, TechCrunch was already reporting the next one. The interesting question this month is not which of these is best. It is what you do when “best” has a shelf life measured in weeks.

To answer that honestly you have to look at the four models clearly first, because the differences are real and they are not the differences the leaderboard advertises.

They are good at different jobs, not different amounts

Kimi K3 is the open-weight shock: frontier-adjacent quality, a one-million-token window and always-on reasoning, at $3 in / $15 out and, within days, weights you can host yourself. It is the value frontier and the sovereignty option in one release.

Claude Fable 5 is the opposite trade. It costs $10 in / $50 out — the premium tier by a wide margin — and earns it on exactly the work K3 leaves unproven: Anthropic reports it topping Cognition’s FrontierBench and breaking 90% on a long-running analytics benchmark, and independent analysts put it first on SWE-Bench Pro, the agentic-coding eval where a model has to hold a plan across many turns. Stripe told Anthropic it ran a codebase-wide migration on 50 million lines of Ruby in a day. That is the capability K3’s launch conspicuously did not measure.

GPT-5.6 Sol is the controlled frontier, in two senses. OpenAI reports it leading Terminal-Bench 2.1 — 88.8%, and 91.9% in its subagent-spawning “ultra” mode — and unlike the others it ships with a reasoning dial you can turn down to cheaper Terra and Luna tiers for volume work. It is also the only one of the four released under a government-coordinated access restriction, available to selected partners rather than anyone with a credit card — a reminder that at the frontier, availability is now a policy variable, not just a pricing one.

And then the rising edge, which is where I would point anyone who thinks this is a two-horse race between America and one Chinese lab. Google rebuilt Gemini rather than ship what it had, betting a multi-tier lineup covers more of the cost-quality plane than any single flagship. DeepSeek’s V4, graduating from preview under an MIT license at twenty-eight cents per million output tokens, makes K3 look expensive and proves the open-weight frontier is a crowd, not an exception. The floor and the ceiling are both moving, in opposite directions, at the same time.

The one you cannot benchmark yet

Which brings us to K3.1, a model no one has announced and everyone is planning around. The pattern is unmistakable: K2.6 to K3 was a 732-point Elo jump shipped in a single version, and the reporting already frames the next release as the one expected to close the remaining gap with Opus 4.8. Moonshot has compressed the release cadence to the point where the market prices in the follow-up before the current model’s weights are even public. You will not benchmark K3.1 before you have to make decisions in a world where it exists. That is the actual condition now: the model you would architect around is frequently the one that has not shipped.

The reflex this produces in a lot of organizations is a standing evaluation committee, a quarterly bake-off, a slide that ranks six models by a composite score. I understand the instinct and I think it solves the wrong problem. By the time the committee reports, the ranking is stale, and worse, it trains the organization to believe the model choice is the decision that matters. It almost never is.

Stop picking the model; build the thing that picks

Here is the leadership move, and it is unglamorous, which is usually the sign it is right. Do not architect your product around a model. Architect it around a seam — a thin interface every model call passes through — and make the model a runtime choice behind it. Then the flood of releases stops being a threat to your roadmap and becomes a menu you order from.

The teams that will look composed a year from now are not the ones who called the July leaderboard correctly. Nobody is calling a leaderboard that changes four times a month. They are the ones who built an organization that treats a new frontier model as an input to swap, not a foundation to pour — so that K3, and Fable 5, and Sol, and the K3.1 nobody has announced, all arrive as the same easy question instead of a fresh emergency. The frontier will keep moving. The advantage was never in guessing where it lands. It was in not needing to.

If this maps to problems you're working on, my inbox is open — the conversation continues on LinkedIn.