Frontier AI Just Got Better and Cheaper. This Is Capitalism at Its Best.

In early June, exactly two labs fielded a model scoring above 50 on the Artificial Analysis Intelligence Index. As of July 17, six do. Four frontier launches — Grok 4.5, GPT-5.6, Muse Spark 1.1, and Kimi K3 — landed inside a single eight-day window. The scores at the very top barely moved. The prices underneath them fell off a cliff.

That is the story, and it is not another leaderboard coronation. Claude Fable 5 is still number one in that snapshot. Its lead is now one point. What actually changed in eight days is not who wears the crown — it is that near-frontier intelligence got 2–3x cheaper, from more suppliers, all at once. This is competition doing the one thing competition is supposed to do: deliver the same or better capability for less money and hand the leverage to whoever is holding the invoice.

The honest scoreboard

The instinct that kicked this post off was “these new models are all Opus-or-Fable level now.” That instinct is broadly right, and I want to defend it — but not by overstating it, because the precise version is more interesting than the hype version.

Here is the July 17 composite, top to bottom: Claude Fable 5 at 60, GPT-5.6 Sol at 59, Kimi K3 at 57, Claude Opus 4.8 at 56, Grok 4.5 at 54. Five models inside six points. A year ago that spread was a chasm; now it is a rounding cluster.

But “all Opus-or-Fable level” is not literally true, and pretending otherwise just gets you corrected by the first person holding the table. GPT-5.6 Sol is genuinely Fable-class: 59 against 60, one point, and it leads the Coding Agent Index at 80 against Fable’s roughly 77. Kimi K3 clears Opus 4.8 on the composite — 57 to 56 — while sitting below Sol and Fable, and Moonshot itself says K3 still trails Fable 5 and Sol overall with a noticeable user-experience gap. Grok 4.5 lands at 54, below both Opus and Fable on this measure, yet performs at or near them on several agentic and coding evals while costing a fraction of the price.

So: not three new universal champions. What you actually have is frontier or near-frontier capability available from six labs instead of two — and that is the more durable fact. One caveat before the numbers, the same one I put on every one of these posts: this is a single third-party snapshot on a single dated day. Prefer the scores to the rank labels, which go stale within a week, and treat all of it as signal rather than scripture.

The economics, which is where this actually lands

Capability got cheaper; that sentence is easy to type and hard to feel until you put the money next to the score.

ModelIntelligence IndexList price ($/1M in → out)AA cost per task
Claude Fable 56010.00 → 50.002.75
GPT-5.6 Sol595.00 → 30.001.04
Kimi K3573.00 → 15.000.94
Claude Opus 4.8565.00 → 25.001.80
Grok 4.5542.00 → 6.000.31
GPT-5.6 Luna511.00 → 6.000.21

Those scores and per-task costs are Artificial Analysis figures from the July 17 snapshot, not universal truth. The list prices carry footnotes the table can’t hold: Fable’s input drops 90% on cached tokens, Kimi K3’s input falls to $0.30 per million on a cache hit versus the $3 cache-miss rate shown, and GPT-5.6’s cache reads keep a 90% discount while cache writes cost 1.25x uncached input. Read the column as a starting point, not a quote.

The reason I lead with the per-task figure rather than price per token is that the two can point in opposite directions. A model with a low sticker price that rambles through twice the tokens or reasons verbosely is not as cheap as its rate card suggests. Artificial Analysis’s cost per task captures exactly that: it is the weighted spend to run a task across the benchmark workload, priced on each model’s actual token use at its actual rates, so a chatty model’s verbosity lands in the number where a per-token quote hides it. What that figure does not do is divide by successful completions — it is spend per benchmark task, not cost per successful task. That second metric is the one you have to measure yourself, because only you know your acceptance criteria: a model that fails a third of your tasks and retries costs far more per accepted result than any benchmark table can show, and the only honest version of that number comes from running your own workload against your own definition of done.

With that distinction kept straight, the AA figures still tell a remarkable story about the eight days: Fable’s cost per task is $2.75, and directly below it sit four models within five points of its score at $1.04, $0.94, $0.31, and — for GPT-5.6 Luna at 51 — twenty-one cents. The leader held the top of the scoreboard and lost the entire floor beneath it.

GPT-5.6: one point off the top, leading where it counts

OpenAI’s GPT-5.6 shipped July 9 as a family — Sol, Terra, and Luna — priced at $5/$30, $2.50/$15, and $1/$6 per million input/output tokens respectively. Sol is the one that matters for this argument. In the independent analysis it tops out at 59 on the Intelligence Index, one point below Fable’s 60, at roughly a third of Fable’s cost per task. On the Coding Agent Index it posts 80 and leads the field, with Terra at 77 and Luna at 75.

The honest asterisk is worth keeping, because it cuts against the easy narrative: in AA-Briefcase, Sol posted the highest Presentation Elo yet still came second to Fable overall, and Fable held the stronger Analytical Quality Elo. Fable still leads on some real knowledge-work measurements. But a model that ties the leader on general intelligence, beats it on the coding agent index, and does it at a third of the per-task cost is not a runner-up. It is a genuine alternative for the exact workloads I care about, at a price that changes what you can afford to run.

Kimi K3: above Opus, cheaper, and not yet the model its own maker sells hardest

Kimi K3 landed July 16, and the spec sheet alone is arresting: 2.8 trillion parameters, Kimi Delta Attention, Attention Residuals, native vision, and a 1M-token context window. On the independent measurement it scores 57 — above Opus 4.8’s 56 — at roughly half Opus’s cost per task. API pricing is $0.30 per million cache-hit input tokens, $3 cache-miss, and $15 output. For a model this large clearing a Claude flagship on the composite, that is a real result.

Now the parts the launch thread skips. Moonshot is refreshingly blunt that K3’s overall performance still trails Fable 5 and GPT-5.6 Sol, with a noticeable user-experience gap. It measured about 62 output tokens per second in the tested API and ran relatively verbose — which, as always, eats into that per-task math the sticker price won’t show you. Only the maximum reasoning effort is available right now; it is sensitive to preserved thinking history and can be excessively proactive.

And one thing I want to be precise about, because I have watched the open-weights story get told wrong before. Moonshot calls K3 an open model, but as of today, July 19, the full weights are not out. The promise is release by July 27, which is why Artificial Analysis still labels it proprietary — the weights are not public yet, so you cannot download or self-host K3 today no matter what a headline implies. That distinction is the whole difference between a rental and an asset, and it is exactly the line GLM-5.2 actually crossed last month with a real MIT release. If Moonshot ships the weights on schedule, the story changes. It has not shipped them yet.

Grok 4.5: lower on the composite, hard to argue with on price

I already wrote a full launch review of Grok 4.5, so I will keep this short. On the Artificial Analysis launch snapshot it posts an Intelligence Index of 54 and a Coding Agent Index of 76 in Grok Build — below Opus and Fable on the composite, near them on the agentic and coding work. What makes it relevant here is the denominator. Its cost per Intelligence Index task is $0.31, and its Coding Agent Index task ran about $2.49–$2.59 against $11.80 for Fable 5 in Claude Code in that same evaluation. The official docs list $2/$6 per million tokens and a 500K context.

xAI’s own numbers — roughly 80 tokens per second and about 4.2x fewer output tokens than Opus 4.8 at max effort on SWE-Bench Pro — are vendor-reported, and I label them that way until someone independent reproduces them. The independent read also flagged a higher hallucination rate on AA-Omniscience, so this is not an undisputed overall winner. It is a near-frontier model that is dramatically cheaper and more token-frugal than the leaders, which for a lot of agent loops is the trade that actually matters.

Why this is capitalism at its best

Strip the logos off and look at what happened structurally. Six suppliers now field a credible frontier or near-frontier model where there were two. Each new entrant did not just add a name to a menu — it forced every other vendor to justify its price. Fable can hold the top of the scoreboard and still watch four models line up underneath it offering 90-plus percent of its capability at a third, a half, or a tenth of its per-task cost. That is not a marketing problem for the leader. That is a market.

This is the good version of the thing, and it is worth naming plainly because the word gets thrown around loosely. Competition here is not an abstraction about incentives; it is a concrete, measurable outcome. In eight days the price below the leader fell 2–3x while the quality gap narrowed to a point. Vendors now have exactly four options: improve quality, cut price, get faster, or lose routed traffic to whoever did. No committee decided that near-frontier intelligence should cost a dollar a task instead of three. A crowded field of suppliers did, by competing for the same buyers — and the same dynamic is why Chinese models winning usage through price was never really a benchmark story either. Developers route to whatever finishes their work cheapest, and enough suppliers make that a rout.

Who wins: the person using AI as a tool

Not a lab. Not a logo. The winner is whoever is building with this stuff, and the win is unusually concrete.

You get more viable fallbacks, so a single vendor’s outage, price hike, or rate limit stops being an existential event. You get real routing by workload — send the coding agent to whichever model leads the coding index this month, send the cheap high-volume classification to the twenty-cent tier, and reserve the flagship for the hard fraction that genuinely needs it. You get cheaper experimentation, because trying a new model is now a config change and a small bill instead of a migration. And you get the big structural discount where it compounds hardest: agent loops that chew through enormous token volumes are precisely where a per-task cost of $0.31 instead of $2.75 stops being a rounding error and starts being the difference between a feature that ships and one that gets killed in the budget review.

That is the whole game. The same dollar buys meaningfully more completed work than it did in early June, and it buys it from more than one place.

The necessary caveat

Competition delivers this outcome under specific conditions, and it is worth being honest that they can fail. It works right now because switching costs are relatively low and the interfaces are portable — most of these models speak a near-identical API, so moving traffic is a routing decision rather than a rewrite. Take that away and the whole mechanism inverts. Lock-in, closed platforms, benchmark gaming, and vendor concentration can each quietly turn “cheaper and faster” into “harder to leave,” and the history of this industry is not short on companies that gave away the razor to sell the blades later.

The defense is boring and it works: keep model selection configurable, keep a fallback you have actually run, and evaluate cost per successful task on your own real workloads rather than on someone’s launch-day table. The portability is what turns six suppliers into leverage. Protect it, and the competition keeps working for you. Surrender it, and you have simply picked a new monopoly to depend on.

Where I land

I do not care which logo wins the model war, and neither should you. The scoreboard winner can change weekly — it nearly did this month, and Fable’s one-point lead is not a fortress. Care about the thing that does not reverse: the same dollar now buys much more finished work, from more than one supplier, and no single vendor gets to name the price of frontier capability anymore. Fable can keep the crown. The rest of us just stopped paying monopoly rent for it.