The summer continues in hyperdrive at the frontier. In the twenty days since my last post, "Everyone's Watching the Wrong Benchmark," the market saw many new models drop: OpenAI's GPT-5.6 line (Sol, Terra, Luna) went generally available the day of my post (July 9), Thinking Machines' Inkling (July 15), Moonshot's Kimi K3 (July 16), Google's Gemini 3.6 Flash, and Anthropic's Claude Opus 5 (July 24), all on the heels of Anthropic's Fable 5 and Sonnet 5 and Z.ai's GLM-5.2, which landed (to varying degrees of access) in the weeks just before it.
Both the open source and closed source frontiers moved meaningfully, so: a retrospective is due some ~480 hours later.
The most interesting development, to me, is that we now have a natural experiment. Kimi K3 is the new open-source frontier: 2.8 trillion parameters, the largest open model ever released. Claude Opus 5 is the new closed frontier. For the first time, both were measured on the same private, fixed-harness benchmark suites within eight days of each other. This means that for the first time, we can actually see the thing I claimed was invisible.
Three weeks ago I made a set of predictions. Here's the point-by-point accounting, with receipts. One prediction was validated in detail. One was wrong, due to a market shift in generally accepted industry practices and not a technological development. The way it was wrong is more interesting than if it had been right.