Before anyone could measure these tools, somebody had to write down what one is
A measurement is only as good as the thing it is measuring, and for two years nobody could say precisely what an AI scribe was. Suppliers described their products differently. Some transcribed. Some summarised. Some suggested clinical codes. Some claimed to do things that would make them medical devices.
Three pieces of plumbing were built before anyone could compare one product with another, and they are worth understanding because no other sector this paper covers has them.
A definition
The evaluation team built a taxonomy: ten elements a product must have to count as ambient voice technology at all, and 40 more covering transcription, summarising, accessibility and other functions. It was assembled from product reviews, interviews and a workshop with NHS England.
That sounds like paperwork. It is the difference between a market where every supplier grades its own homework and one where a buyer can ask a common question.
A list
NHS England published a national registry of suppliers in January 2026. Suppliers self-certify, which is a real limit, but the register exists and can be checked.
When the evaluation team mapped the market against its taxonomy, it found 32 products, 14 claiming some use in the NHS or social care, 11 registered with the medicines regulator and 11 on the registry. It also recorded that claimed use was often difficult to verify independently. The plumbing does not make every claim true. It makes the gaps visible.
A line in law
On 29 July the medicines regulator, working with NHS England, published guidance on where these products sit in law. Products used only to transcribe, summarise, draft letters or suggest codes for a clinician to review are not regulated as medical devices. Products intended to support diagnosis or treatment, or that act without a clinician reviewing the output, are.
NHS England had previously treated all of them as medical devices and revised its guidance to match.
Read that as a boundary being drawn around the human. The tool may listen and draft. The moment it decides, it becomes something the state inspects.
What does that make possible?
Because those three things exist, the February 2027 evaluation can do something unusual. It can compare products that meet a shared definition, in organisations that appear on a public list, against a regulatory line that says which functions are in scope.
Now set that against the rest of the economy. In legal services this paper found three widely quoted time-saving figures, all from suppliers, no shared definition of the task being measured, no register of who is using what, and no regulator drawing a line around the point where software stops assisting and starts deciding.
The NHS did not get this by accident. It got it because a public body bought the technology at scale and therefore had to answer for it. The question that leaves is what happens in every sector where the buyer is a thousand private firms and nobody has to answer for anything.
What has to exist before a claim about AI at work can be tested?
- NIHR Rapid Service Evaluation Team, Evaluation of ambient voice technology in the NHS, project page, read 17 September 2026
- NIHR Rapid Service Evaluation Team, Mixed-method evaluation of ambient voice technology: phase 1, 4 February 2026
- MHRA, regulatory status of ambient voice technologies, 29 July 2026
- The Quernal, Issue 67, 14 September 2026
- Matt Brazil This paper has spent a month saying nobody measures what AI saves. In the NHS, somebody is
- Dr. Leah Sandoval The NHS has what legal services does not: an independent evaluation of an AI tool, with the method published before the answer
- Tomás Reyes The measurement my colleagues are celebrating reports in February 2027. The rollout finished long before that
- From the Editor If a machine takes a task off your desk, three questions about the hour it hands back