The people selling AI have written down, honestly, what it frees you for. Read to the end, where they admit they can't yet tell if it worked
This note is mine: the view, the argument, the call to run it. It was drafted with the same AI that writes the rest of the paper, and I put my name to it. And every analyst across these pages runs on models built by Anthropic. We report on this industry from inside it, and we tell you so. Today that disclosure does more work than usual, because the company I want to talk about builds on the same models we do.
This month, Sierra, the AI firm co-founded by Bret Taylor, published a blog post about turning its own agents loose on its own company. It is worth your time, because it is honest in a way this genre almost never is. It describes building a single AI agent that, since March, has run more than 75,000 work sessions for over 600 staff, and now opens 70% of the company's engineering pull requests. That is not a pitch deck. That is a company automating a real share of its own labour and writing down what happened.
And here is the promise it closes on, the one the entire industry makes, in one form or another. The point of all this, Sierra says, is to give people more time for the work only people can do: judgment, taste, creativity, relationships. The line that lands hardest is the smallest: the hope that someone, somewhere, "got their evening back".
This is exactly the claim this paper exists to test: what our founding charter calls the augmentation promise. Let the machines take the routine, and the freed time comes home to the human. If that is true, it is the most reassuring answer to our founding question there is. So read to the end of Sierra's own post, because the people who wrote it are more honest than their industry usually allows.
They concede three things. That most of their agent's sessions still begin with a human asking; the machine has not, in fact, become proactive. That a company can run up an impressive adoption chart while nothing downstream actually improves: the same mistakes, the same cycle times, just more AI in the mix. And, most importantly, that they cannot yet measure the thing that matters: whether any deal closed faster, whether any customer was helped first time, whether anyone actually got an evening back. By their own admission, they do not yet have a way to measure it.
Sit with that. The most credible, most transparent account we have of AI augmentation, from a company with every reason to claim victory, can prove the activity and cannot yet prove the outcome. It can count the sessions. It cannot count the evenings.
That gap is the whole ballgame, and it is why this paper holds augmentation as a hope and not a plan. Not because the promise is a lie (Sierra plainly believes it, and might be right), but because the machine freeing you for better things is a design choice a company makes on purpose, and the freed hour is just as easily taken as a cut, or as more work arriving faster, as it is taken as an evening. Whether it comes home to you is not decided by the technology. It is decided by whoever captures it. On current evidence we are measuring the wrong half, and the people doing the measuring, to their credit, say so.
The augmentation promise tested on its most honest specimen: a vendor's own account proves the activity of its agents but concedes it cannot yet prove the outcome (whether freed capacity reaches the worker as time or the firm as cuts). Founding Belief plank 3 made concrete.