Localign

Report card, round two: four new AI models sit the Dutch finals

Doga ArasDA
Bink HolBH
Doga Aras & Bink Hol
August 3, 2026
4 min read
Final Exams
Education
AI Models

Six weeks after we graded eleven AI models on the Dutch final exams, four new candidates showed up: Claude Opus 5, GPT-5.6, Kimi K3 and GLM 5.2. Same exams, same official answer keys. All four land in the top half of the class, and for the first time the highest grade belongs to an open model.

A classroom of friendly AI robots sitting the Dutch final exams

In June we sat eleven AI models down for the Dutch HAVO final exams and published their report card. Barely six weeks later, four new models were released that we simply had to test: Claude Opus 5 from Anthropic, GPT-5.6 Sol from OpenAI, and two open models, Kimi K3 from Moonshot AI and GLM 5.2 from Z.ai. We got the exam booklets back out and ran everything again, exactly the same way. This is a short follow-up, not a rewrite of the original. Four new students joined the class, and their grades are in.

Same exam, same rules

Identical setup to the June benchmark: seven HAVO subjects (Dutch, English, mathematics, physics, economics, history and geography), graded question by question against the official answer keys on the Dutch 1-to-10 scale, where 5.5 is a pass. No web search, no tools: what a model knows on its own.

The new report card

The full grade list is below, fifteen models now. The four new ones sit up front with a badge and their own colour; everything to the right of the divider is June's class, unchanged.

The report card: 11 AI models, 7 subjects

Dutch HAVO final exams, graded question by question against the official answer keys.

School year 2025/2026
Loading…
8.0 and up6.5 to 8.05.5 to 6.5failing grade
Pass mark is 5.5. Below each model name: its size in parameters (~ marks an estimate).

The front row has changed

First, the standings. Kimi K3 averages a 9.1, the highest grade we have measured so far, with two perfect 10s: economics and history. Claude Opus 5 gets an 8.9, exactly the same average as Claude Fable 5, June's best. GPT-5.6 Sol scores an 8.4, a small but real step up from GPT-5.5's 8.2, including a flawless economics exam, 56 out of 56 points, the second perfect score on that subject in this round. GLM 5.2 lands on an 8.3, which would have put it second in June's class. All four newcomers score higher than almost everything we tested six weeks ago. The top of the class got crowded fast.

A lot still looks familiar. Economics is still the subject where AI shines: clear questions, a small answer space, and now four candidates at or near a perfect score. Physics and mathematics still separate the careful reasoners from the rest, and there Claude Opus 5 is the strongest of the four: a 9.5 in physics that none of the others gets close to, and a tie with Kimi K3 for the best mathematics grade. The old lesson from June still holds: paying more does not simply buy better grades. Opus 5 matches Fable 5's average at half the price per token, and the model at the top of the table is not the most expensive of the four.

Two open models, two very different report cards

The headline is hard to miss: for the first time, the best report card in the class belongs to an open model, one you can download and run yourself. Kimi did not win it on the easy subjects. It scored an 8.5 in Dutch, the subject the whole June class struggled with, well above anything we had measured there before. It added an 8.4 in English and perfect 10s in economics and history. Its lowest grade, an 8.4, would have been a perfectly good average in June. The only place it gives ground is physics, where Opus 5's 9.5 is clearly better than its 8.7.

The second open model tells you not to treat "open" as a single category. GLM 5.2 is excellent where the answer is well defined: a 9.8 in economics, a 9.7 in history. Then it drops to a 6.7 in English and a 6.9 in Dutch, its two weakest subjects by a wide margin, and exactly the pattern the whole June class showed. Within the same six weeks, one open model solved the language problem that had held every model back, and another one did not. Whether a model is open or closed tells you far less about its report card than what it was trained to be good at.

The bill, revisited

Round one ended with a chart showing that price and grades barely have anything to do with each other. Round two makes that even clearer. Claude Opus 5 delivers Fable 5's average at half Fable's price. Kimi K3 tops the entire class at three dollars per million tokens in, less than a third of what Fable 5 charges for a lower average. That does not make the expensive tier pointless: Opus 5's physics grade, and the fact that it scores at least a 7.2 in every single subject, is exactly what you pay for. But the best report card in the class no longer costs flagship money.

Interactive chart

Paying more barely buys a higher grade

Average grade set against what using each model costs.

Best valueTop of the class
Cost
Loading…
Cost in dollars per million tokens, the snippets of text an AI computes in. Blended averages reading (what the model takes in) and writing (what it returns). Hover over a dot for the details.

The conclusion from June still stands, it just got stronger: the question is not whether AI passes the exam, but where each model slips, and what that costs you. Six weeks ago our advice was a strong, affordable base model, with the expensive tier where it really counts. Today the best report card in the class costs a third of the most expensive one. Keep testing in your own context, keep asking to see the grade list, and count on this list changing again. Four new names in six weeks says enough.

About the authors

Doga ArasDA
Doga Aras
AI Engineer

AI engineer at Localign and builder of the exam pipeline: from ingesting the exam booklets to automated grading against the official answer keys.

An average tells me very little. Show me where a model slips and I'll tell you what it's worth.

Bink HolBH
Bink Hol
AI work student

AI work student at Localign. Ran the exams, combed through the answers of fifteen models and knows the grade list by heart by now.

The biggest AI is not always the smartest one.

Questions about this research?

Reach out. We're happy to explain the methodology or discuss what this means for your use case.