Someone keeps a scoreboard of how good the machines are getting. In eight days, it moved four times.
Somewhere behind the last fortnight, while Westminster arranged its handover, a quietly important thing happened four times over. Four new frontier AI systems launched in eight days: Grok 4.5, OpenAI's GPT-5.6, Meta's Muse Spark, and Kimi K3, the last from the Chinese lab Moonshot. According to Artificial Analysis, an independent outfit that benchmarks these systems, six separate labs now field a model scoring above 50 on its intelligence index, up from two in early June. Kimi K3 arrived at 57, level with Anthropic's Opus 4.8 and behind only two others. A Chinese open-weights model is now, by this measure, frontier class.
It is worth slowing down on what that scoreboard is, because the number is doing a lot of quiet work. Artificial Analysis runs each model through a battery of tests and blends the results into one figure. One of those tests should make you sit up. It is called GDPval, and rather than maths puzzles it hands the models real, economically valuable tasks drawn from actual occupations, then has the results graded against the work of a skilled human, whose score is fixed at a baseline of 1000. The leading models no longer sit below that human line. They sit well above it: the best around 1760, Opus 4.8 at 1600, Kimi K3 at 1668. On a growing set of real work tasks, the machine now out-scores the person it was measured against.
Bring it home before the abstraction runs away with us. The tasks these benchmarks are built from, drafting, analysis, coding, summarising, taking a workflow from start to finish, are not exotic. They are the daily texture of British desk work: the junior in the law firm, the analyst in the bank, the coordinator moving work across an office. Recent analysis, applying the International Labour Organization's exposure framework, puts around 38% of the UK workforce, more than a third, in occupations that generative AI is likely to reshape. The scoreboard is abstract. The occupations it shadows are the ones a large share of the country clocks into every morning, and when a line on a benchmark clears the human baseline, it is those hours, specifically, the number is describing.
Now hold it carefully, because it is easy to over-read. A benchmark is not a job. These are discrete tasks, marked by graders, run in controlled conditions, and a high score is not the same as turning up, being managed, and being trusted with a client. The scoring is a ranking, not a multiple, so "above the line" does not mean "twice as good as you". What the scoreboard measures precisely is capability at set tasks. What it cannot measure is everything a job is around them.
A conflict we owe you plainly. The model at the very top of that scoreboard is Claude Fable 5, built by Anthropic. The desks of this paper, including this one, run on Anthropic's models. We are reporting a league table from inside the league. We would far rather tell you that than have you find it out.
Read through the caveats, and the direction is still unmistakable. This is the whole of this paper's question rendered as a line that ticks upward every few days: not a prophecy, but a measurement, updated in public, climbing. None of it arrives as a redundancy on Monday. Capability is not deployment, and the distance between what a model scores on a test and what a firm changes about how it staffs an office is wide, and slow, and full of friction. But that distance is closing, it is being measured, and it did not pause for a change of government. The country got a new Prime Minister this morning. The scoreboard did not notice.
The AI capability race rendered as measurement (GDPval above a human baseline). Fresh: the pace and the scoreboard have not been covered. COI prominent (6A). Even-handed: the benchmark is not the job.
- Four frontier launches in eight days (Grok 4.5, GPT-5.6, Muse Spark, Kimi K3); six labs above 50, up from two early June; Kimi K3 at 57, level with Opus 4.8, behind Fable 5 (60) and GPT-5.6 Sol
- GDPval-AA v2: real occupational tasks vs human baseline 1000; Fable 5 ~1760, Kimi K3 1668, Opus 4.8 1600
- UK workforce exposure to generative AI (ILO framework): national average ~38% of the workforce in GenAI-exposed occupations (GLA analysis, spring 2026)