In June we sat eleven AI models down for the Dutch HAVO final exams and published their report card. Barely six weeks later, four new models were released that we simply had to test: Claude Opus 5 from Anthropic, GPT-5.6 Sol from OpenAI, and two open models, Kimi K3 from Moonshot AI and GLM 5.2 from Z.ai. We got the exam booklets back out and ran everything again, exactly the same way. This is a short follow-up, not a rewrite of the original. Four new students joined the class, and their grades are in.
Same exam, same rules
Identical setup to the June benchmark: seven HAVO subjects (Dutch, English, mathematics, physics, economics, history and geography), graded question by question against the official answer keys on the Dutch 1-to-10 scale, where 5.5 is a pass. No web search, no tools: what a model knows on its own.
The new report card
The full grade list is below, fifteen models now. The four new ones sit up front with a badge and their own colour; everything to the right of the divider is June's class, unchanged.
The report card: 11 AI models, 7 subjects
Dutch HAVO final exams, graded question by question against the official answer keys.
The front row has changed
First, the standings. Kimi K3 averages a 9.1, the highest grade we have measured so far, with two perfect 10s: economics and history. Claude Opus 5 gets an 8.9, exactly the same average as Claude Fable 5, June's best. GPT-5.6 Sol scores an 8.4, a small but real step up from GPT-5.5's 8.2, including a flawless economics exam, 56 out of 56 points, the second perfect score on that subject in this round. GLM 5.2 lands on an 8.3, which would have put it second in June's class. All four newcomers score higher than almost everything we tested six weeks ago. The top of the class got crowded fast.
A lot still looks familiar. Economics is still the subject where AI shines: clear questions, a small answer space, and now four candidates at or near a perfect score. Physics and mathematics still separate the careful reasoners from the rest, and there Claude Opus 5 is the strongest of the four: a 9.5 in physics that none of the others gets close to, and a tie with Kimi K3 for the best mathematics grade. The old lesson from June still holds: paying more does not simply buy better grades. Opus 5 matches Fable 5's average at half the price per token, and the model at the top of the table is not the most expensive of the four.
Two open models, two very different report cards
The headline is hard to miss: for the first time, the best report card in the class belongs to an open model, one you can download and run yourself. Kimi did not win it on the easy subjects. It scored an 8.5 in Dutch, the subject the whole June class struggled with, well above anything we had measured there before. It added an 8.4 in English and perfect 10s in economics and history. Its lowest grade, an 8.4, would have been a perfectly good average in June. The only place it gives ground is physics, where Opus 5's 9.5 is clearly better than its 8.7.
The second open model tells you not to treat "open" as a single category. GLM 5.2 is excellent where the answer is well defined: a 9.8 in economics, a 9.7 in history. Then it drops to a 6.7 in English and a 6.9 in Dutch, its two weakest subjects by a wide margin, and exactly the pattern the whole June class showed. Within the same six weeks, one open model solved the language problem that had held every model back, and another one did not. Whether a model is open or closed tells you far less about its report card than what it was trained to be good at.
The bill, revisited
Round one ended with a chart showing that price and grades barely have anything to do with each other. Round two makes that even clearer. Claude Opus 5 delivers Fable 5's average at half Fable's price. Kimi K3 tops the entire class at three dollars per million tokens in, less than a third of what Fable 5 charges for a lower average. That does not make the expensive tier pointless: Opus 5's physics grade, and the fact that it scores at least a 7.2 in every single subject, is exactly what you pay for. But the best report card in the class no longer costs flagship money.
Paying more barely buys a higher grade
Average grade set against what using each model costs.
The conclusion from June still stands, it just got stronger: the question is not whether AI passes the exam, but where each model slips, and what that costs you. Six weeks ago our advice was a strong, affordable base model, with the expensive tier where it really counts. Today the best report card in the class costs a third of the most expensive one. Keep testing in your own context, keep asking to see the grade list, and count on this list changing again. Four new names in six weeks says enough.



