Every public number attached to Kimi K3 describes a rented API, not the weight file that is supposed to land on July 27. Artificial Analysis scores the model 57 on its Intelligence Index, ranks it fourth of 186, measures 35.2 output tokens per second, and prices the endpoint it measured at $3.00 per million input tokens and $15.00 per million output tokens 1. The same page lists Kimi K3 as proprietary, weights not publicly available 1. Simon Willison’s hands-on carries that price sheet and the promised open weight release by July 27, 2026 2. I take the promise at face value. The problem survives it.
The leaderboard scored a configuration
A rank is a property of a served configuration: specific hardware, a specific precision, a batching and routing policy, a price sheet. It is not a property of a weight file. For most models that distinction stays academic, because the two objects track each other closely enough that the leaderboard row works as a proxy. Kimi K3 is where the proxy breaks, and the cleanest evidence is the measuring institution itself: Artificial Analysis ranks the model fourth of 186 while filing it in the proprietary column, weights not publicly available 1. Nathan Lambert describes the arriving artifact as a 2.8 trillion parameter mixture-of-experts model, weights due July 27, the closest open models will have been to the frontier since DeepSeek R1 3. At that size the distance between the vendor’s serving stack and anything a normal team can afford is not a tuning delta. It is a different quantization and a different memory hierarchy, so it is a different quality and cost profile, and no leaderboard row covers it.
The strongest case against this
Put the objection at its best. A ranking published before a weight release measured the only artifact that existed, which is true of every staged launch, and nobody writes a memo about those. Moonshot has committed to a date 2. If the weights ship on July 27, the complaint expires in four days and reads as cynicism about a vendor that did exactly what it said.
That objection would land if my claim were about sincerity. It is not. Grant the release, on schedule, byte for byte. What does not arrive on July 27 is a measurement of the object being released: leaderboards are not re-run at your precision on your cards, and the figure that keeps circulating as the model’s identity will go on being the endpoint’s.
There is also a tempting version of my own argument that is simply wrong, so I will kill it here. 35.2 output tokens per second is not a ceiling for local reproduction. A latency-optimized single-stream deployment can beat an endpoint tuned for concurrency, and a memory-constrained one offloading experts can land an order of magnitude below it. The honest statement is weaker and more useful: the number does not transfer in either direction, and neither does the quality-derived rank once the precision changes.
What I would do on Monday
Price the migration against the API you can benchmark today, not against weights you cannot yet run. Run your own evaluation through the endpoint priced at $3 and $15 per million tokens 2, keep the result, and label it an API number. When the weights land, measure output tokens per second and cost per million on your own harness, at the quantization you would actually buy, single-stream first and then under your real concurrency, before anyone signs anything. Until that measurement exists, the honest line in the migration doc is one sentence: open weights is an option whose price nobody has published, the vendor included.
My prediction: the gap between the ranked object and the runnable one widens with every frontier-scale open weight release, and the indexes will keep scoring the endpoint, because the endpoint is the only thing they can call.