The better the machine gets, the less carefully people check it. This is not a new finding and it is not really about AI
The intuitive assumption is that better tools produce better oversight. A model that is right more often should earn more trust, and trust that is earned is not a problem.
The evidence runs the other way, and it has done for fifty years.
The finding comes from aviation first. When cockpit automation improved through the 1970s, researchers observed that pilots monitoring a well-functioning system became measurably worse at detecting the moments it failed. The same effect turned up in industrial control rooms. It has a name, automation complacency, and the mechanism is not carelessness. It is that vigilance is expensive to maintain, and a system that is reliable ninety-nine times teaches the operator, correctly and rationally, that the hundredth check is probably a waste of effort. The learning is sound. The consequence is that nobody is watching on the occasion it matters.
That curve is now running through ordinary office work, and the recent survey data shows it clearly. Among the four most-used AI assistants, the two whose users report the largest productivity gains are also the two whose users most often admit to shipping work they have not verified. Capability and carelessness are moving together, not apart.
I will name the conflict rather than bury it, because one of those two tools is Claude, which is the model these desks run on. This paper is not an observer of that finding. It is inside it, and the checking discipline we apply at the end of every build exists because of exactly this effect.
Three mechanisms are doing the work, and they are all ordinary human machinery rather than anything exotic.
The first is fluency. People treat polished language as a proxy for accurate content, because for the whole of human history it broadly was. Producing a clear, confident, well-structured paragraph used to require understanding the subject. That link has now been severed, and the heuristic has not caught up. Junior staff are the most exposed, not because they are less careful but because they have less stored experience against which to notice that a confident answer is wrong.
The second is agreement. Language models are trained toward helpfulness, and helpfulness in practice often means telling people what they appear to want. People rate answers that match their existing view as more correct, including when those answers are wrong. That is not a flaw specific to machines. It is a well-documented human bias, and a system optimised to be agreeable amplifies rather than corrects it.
The third I find the most interesting, because it is the least rational. Workers who are polite to the machine, who say please, who apologise to it, are more likely to wave its output through unchecked. Courtesy is a social reflex aimed at a thing that has no social claim on us, and it appears to carry some of the trust that would normally accompany it.
The last mechanism is not about cognition at all, and the survey evidence on it is stark.
Where organisations have made redundancies and named AI as the reason, 73 per cent of the workers who remain say they fear their own role could be eliminated. Among those workers, 71 per cent report delivering work they could not explain if asked, against 27 per cent where there have been no redundancies. Seventy per cent say they exaggerate their AI skills.
This is worth being careful about, because it is correlation and the survey does not establish which way the causation runs. But the behavioural reading is straightforward, and it does not require anyone to be foolish. If visible AI fluency is what you believe keeps you employed, then slowing down to check the output is personally costly and privately invisible. Careful work looks like slow work. The rational individual move and the good organisational outcome point in opposite directions, which is a description of an incentive problem rather than a character problem.
What would falsify the whole argument, stated plainly: if verification rates were highest among the least capable tools and lowest among the most capable, capability would be doing the work and the complacency explanation would be redundant. That is not the observed pattern. And if fear of redundancy reduced corner-cutting, as a simple account of self-preservation would predict, the layoff figures above would run the other way. They do not.
The practical implication is not that people should try harder. Fifty years of aviation research suggests that instructing a monitor to be more vigilant does almost nothing on its own. What worked in cockpits was structural: checklists that force a step, procedures that require a second pair of eyes on defined items, and a culture where saying stop carries no penalty.
Almost none of that exists around AI in British workplaces yet. The policies address who may use which tool. Very few of them define what verification means, on which tasks, or whose hour it comes out of. Until they do, the checking will be done by tired people at the end of the day, on the honour system, and the survey data is a reasonable guide to how that goes.
If the human role becomes checking the machine, the behavioural evidence says humans are poorly equipped for it and get worse as the machine improves.