A job has appeared in British working weeks that no one designed, no one budgeted for and no one is paid to do
Start with the arithmetic, because it is unusually clean.
Workers using AI say automation saves them about eleven hours a week. Of the total time they spend interacting with AI, 37 per cent goes on supervising and repairing it, 36 per cent on actually producing work with it, and 27 per cent on learning the tools and building workflows. The supervision comes to roughly six and a half hours a week. More time goes into managing the machine than into using it.
Break the six and a half hours down and the shape of the work becomes visible. Loading context the system should already have: 2.3 hours. Reviewing output for errors that the polish conceals: 2.2 hours. Debugging, meaning re-prompting and swapping models until something usable returns: 1.7 hours. Clearing up after handoffs: the remainder.
Any engineer will recognise the pattern. The load did not vanish. It moved, and it moved from a place where it was measured to a place where it is not.
Consider what happens to a task that used to take a person ninety minutes. The machine now does it in ten. Sixty of the saved minutes come straight back as the work of making those ten minutes trustworthy. The task is faster. The person is not obviously freer. And the eighty minutes that used to sit in a workload model, a headcount plan and a delivery estimate now sit in a gap between tasks that no system records.
That is the load-bearing point, and it is a design failure rather than a moral one. Organisations measure what they deploy. They do not measure what deployment costs the person operating it.
Three properties make this worse than an ordinary measurement gap.
The first is that the work is invisible by construction. Time spent using an AI tool is logged by the tool. Time spent deciding an output is wrong, reconstructing why, and starting again is logged nowhere. The metric captures the productive half and misses the half that pays for it. Any system measured that way will look like it is succeeding while the people inside it are absorbing an unbudgeted cost.
The second is that the errors are disguised. Bad human work has traditionally announced itself. The clumsy sentence, the typo in the first line, the graph with the wrong axis: these are friction, and friction makes a reader slow down. Machine output arrives polished regardless of whether it is right. The cheap warning signs that professionals have relied on for a century have been removed, and almost nothing has replaced them. The people most exposed are those without enough experience to know what a correct answer looks like, which is to say the junior staff this paper has spent two months watching disappear from hiring plans.
The third is the feedback loop, and it is the one that should worry anyone running a team. Supervision is tiring. The researchers found that the more time workers spend feeding context, the more likely they are to report being worn out by it, and debugging is the most exhausting component of the lot. Exhausted people stop checking. They ship output they have not verified and could not defend if asked, which lands on a colleague downstream who did not produce it, does not have the context to repair it, and now has to repair it anyway. That colleague's supervision burden rises. The cycle tightens.
Roughly two thirds of AI users in the survey admit to shipping work in that condition. Among the heaviest users it approaches four in five. Workers doing the most supervision are markedly more likely to be looking for another job.
Now the British part, and it cuts against the obvious remedy.
The United Kingdom, of the three countries measured, has the most developed governance. More British workers have read their employer's AI policy in full than American or Australian ones. More are confident in their organisation's AI strategy. More say there are clear rules about who may build or deploy an agent. On every institutional measure in the study, Britain is ahead.
British workers spend 38 per cent of their AI time on supervision. Americans spend 36. The study states that differences below three percentage points are not statistically significant, so the honest reading is not that Britain is worse. It is that Britain is the same, having done considerably more of what is usually recommended.
The reason is structural. Governance describes who may use a tool, under what conditions, and who answers when it fails. It is a permissions layer. Supervision is what the work costs once permission has been granted. A policy can say a human must review the output. It cannot say how many hours that review takes, whose day it comes out of, or what else stops happening as a result. Those are questions of work design, and almost nobody is asking them.
The falsifiable version, so this can be tested rather than believed: if supervision were primarily a governance problem, the country with the strongest governance should show a lower burden. It does not. If it were a tooling problem, the most capable tools should reduce it. On the study's own figures, the tools whose users report the largest productivity gains also report the highest rates of shipping unverified work.
Two caveats the reader is owed. Every figure here is self-reported, which is a genuinely weak instrument for measuring your own time. And the research group behind the study sells enterprise software designed to solve precisely the problem it has identified, which does not make the numbers wrong but does explain which conclusion the report reaches.
What would settle it is an employer, anywhere in Britain, publishing the hours. Not adoption rates, not licence counts, not a strategy. The hours the machine gave back, and the hours it took to make them safe. Nobody has published that. Until somebody does, every productivity claim on this technology is missing its denominator.
If the freed time is spent supervising the machine, the founding question has an answer already, and it is not the one anyone was promised.