Your best people are now proofreaders.
Making things got free. Owning them didn't. Somebody pays the gap.
Karim runs engineering at an Amsterdam fintech. Last week he sent me a screenshot of his calendar.
His calendar looks like what most “senior” people are doing right now: reviewing LLM-generated work.
In his calendar, thirty-one hours of review…code review, architecture review, "can you sanity-check this before it ships" review…
His team had just closed the biggest shipping quarter in the company's history (in terms of lines of code). Pull requests roughly tripled since January. Token consumption is picking up. Junior team members ship impressive things but can’t talk about their work anymore… and overall, “problems solved” vs. token consumption doesn’t seem to show a positive ROI.
That’s the reality of most teams using AI at scale (not for individual gains).
Karim is the most expensive person in that building. His job is now reading, and he is now the bottleneck.
AI didn't make the work cheaper. It moved the bill from the people who make things to the people who have to own them.
The part that never got cheap
A first draft of almost anything now costs roughly nothing. Code, a client memo, a support reply, a slide. Fast, unlimited, basically free.
Owning it didn't move at all.
Owning something means reading it. Understanding it well enough to defend it. Putting your name on it. Answering the phone when it turns out it was wrong. There's no model for that, and there isn't one coming soon. You can generate the artifact. You can't generate the liability.
So the cost didn't vanish. It moved. Off the cheapest hour in the process, typing, onto the most expensive one, judgment.
That's the verification tax. Two things make it dangerous.
It never shows up as a line item. No invoice says "verification." No dashboard has a column for it. It shows up as a senior's calendar filling with review and a junior's learning curve going flat.
And you pay it twice. Once for the seniors, who now spend the week reading instead of building. Once for the juniors, who shipped an enormous amount last quarter and learned nothing from any of it. You paid tokens for both.
(And bonus: increasing AI model capabilities seems to increase the verification tax, as the tasks handled by agents and LLMs become more complex and require more expertise to own them…)
I've made the skills version of this argument before. Letter 76, verification is the only skill left. Letter 85, we stopped making seniors. Letter 89, agentic coding is a slot machine. All still true. All of it has moved approximately zero budgets, because skill arguments lose to spreadsheets.
So this week, the spreadsheet.
Start with the uncomfortable part. You probably can't feel any of this happening.
A research group called METR ran the cleanest test I know of. Sixteen experienced open-source developers, working on their own repositories, on real tasks. Some tasks they were allowed AI tools. Some they weren't. Coin flip.
With the tools, they took 19 percent longer.
Then somebody asked them how it went. They said AI had made them about 20 percent faster.
The 19 percent isn't the interesting number. The gap is. Skilled people, their own code, a stopwatch running, and they were about forty points wrong about their own week.
(METR has since retired that result, and the reason is better than the result. They ran it again with 2026 tools and still found no speedup. They pulled it because they lost confidence in their own instrument, not because the effect flipped. I wrote that up in Letter 90.)
Google's DORA team found the same shape at survey scale. Three thousand respondents. More AI adoption tracked with slightly less throughput and noticeably less stable delivery. Their read was batch size: teams shipping bigger changes without keeping the habits that make bigger changes survivable.
Then the 2025 edition landed. Throughput had recovered. Stability had not.
DORA's own summary was that AI is an amplifier. It magnifies whatever your organization already is. I've been saying a version of that sentence for two years and I'd rather have been wrong.
Hold on to the fact that only one of those two numbers recovered. I'll come back to it.
The last time production got free
Radiology already ran this experiment. National scale, over twenty years.
A plain film used to be a handful of pictures. A modern CT is hundreds or thousands of slices. Producing an image collapsed toward free. Some researchers at the Mayo Clinic went and counted what that did to one hospital. Over eleven years, the number of images somebody actually had to look at went up roughly tenfold.
Radiologist headcount did not go up tenfold. Nothing about training a radiologist got cheaper, and you can't buy one the way you buy a scanner.
So they reported the number that matters. Adjusted for staffing, across an eight-hour shift, it worked out to one image every three to four seconds.
Call that what it is. Triage at speed, done by the most expensive and least scalable person in the building, at a pace set by a machine that has no idea what hour seven feels like.
And it costs you. Another study put radiologists in front of bone films at the end of a full clinical day. Same people, later in the day. They caught fewer fractures.
The scanner got cheap. The reader didn't, and couldn't, and buying more scanners never helped. Hospitals that booked imaging as a production saving got a reading-capacity problem instead.
Now the second-order effect. This is the part I want you to hold onto.
Incidental findings.
You scan someone to answer one question. The scan surfaces three ambiguous things nobody was looking for. A nodule. A cyst. A shadow. A number slightly outside the normal range.
Researchers took 229 healthy Swiss adults with no symptoms and scanned each of them anyway. Forty-seven percent had an incidental finding. Chasing the ones worth chasing cost about three thousand francs each.
Read that again. Cheap production didn't only move a cost. It manufactured a cost that had not previously existed anywhere in the system. Somebody now has to adjudicate a thing nobody ordered, book the follow-up, carry the worry, and own the liability of calling it nothing.
Your AI stack is doing this to you right now.
Every agent that helpfully flags an anomaly. Proposes a refactor nobody scoped. Drafts a third option for the deck. Opens a risk-register entry. Surfaces a churn signal. Each one is an incidental finding. Each one lands on a human queue. Each one is a decision somebody has to make and own, and it exists only because generation got cheap.
The bill moved. Then it grew.
Marta's queue
If you're not in engineering, you might be reading this as somebody else's problem. It isn't. Coding is just where the effect is most visible, because coding is where the models are strongest now.
Marta runs support at a mid-size B2B company. Her team drafts every reply with AI now. First-response time is down sharply. Throughput is up. Her dashboard is beautiful.
Two things happened underneath it.
Her four best agents, the ones who are genuinely excellent at the hard ten percent, now spend most of the day approving drafts. She took her four best people and turned them into a review queue for a machine. The hard ten percent still arrives. It just arrives to someone who spent six hours clicking approve.
And her new hires have never heard an angry customer unprompted. Not once. They only ever meet the situation pre-digested into a polite paragraph with a suggested resolution attached. So when a draft is subtly wrong, wrong about this person, this contract, this history, wrong in a way that reads perfectly, they approve it. They have nothing to feel it against. That feeling gets built by hearing two hundred customers say it out loud, badly.
Plausible and correct are not the same thing. The draft is always plausible. That's what it was optimized for.
Marta's bill is identical to Karim's. Her most expensive people are reading instead of doing the thing that made them expensive. Her cheapest people are producing volume that teaches them nothing. Same tax, different function. You'll find it in legal, in marketing, in credit, anywhere a human name goes on the output.
The uncomfortable truth
The productivity number on your dashboard isn't real.
It's real in the sense that the things it counts happened. Tickets closed, PRs merged, drafts produced. It isn't real in the sense that it never nets out what those things cost. It measures the side of the ledger that got cheap and says nothing about the side that got expensive, because the expensive side never got a column.
You have a beautiful number for output. You have no number at all for ownership.
From my experience, the fix is almost never technical.
Nearly everything I've saved clients in the last three years came from changing how the team works, not from what the team bought. Where review sits. Who reviews what. How big a batch is allowed to be. What's allowed to leave the building without a name on it. Boring, unglamorous, organizational.
That's also why DORA's throughput recovered. Not because a better model shipped. Because a lot of teams spent a year rebuilding review and testing around a production rate that had already changed. The lag was organizational, and organizations can close it. It just takes longer than a procurement cycle and it doesn't photograph well.
Their stability number didn't recover.
A year of collective effort, and what got fixed was the speed of the pipe. Not the correctness of what came out of the end of it.
You cannot automate a mess. You can now automate it forty times faster, which is a different and considerably worse thing.
If nothing changed about how your people work and your output tripled, you're not going faster, you're just building a bomb.
What to do Monday morning
Pick one and start small. Don’t start with all four at once.
(A team that installs four new disciplines in one week is running none of them by March.)
Put verification on the P&L. For two weeks, have every senior tag their calendar in two colors. Original work, and review of machine output. Multiply the review hours by their loaded cost. That's your verification tax, and until you run this it exists nowhere in your company. Most leaders I've done this with are startled by the size. A few are more startled by how many of their best people now work as editors!
Cap the batch, not the output. Nothing enters a queue faster than the queue can be read carefully. If a diff is too long to read, it's too long to merge. If an agent generates forty replies an hour and your reviewer can genuinely check twelve, your throughput is twelve. The other twenty-eight are unreviewed RISK.
Put a name on it. Everything that leaves the building carries one human who read it and will answer for it. Not a team or "AI-assisted, reviewed." You need A NAME. Watch what volume does in week one. That drop is your honest number.
Invert who gets the hard work. Right now you pay senior rates to approve drafts and junior rates to generate work that teaches nothing. That's upside down as pure economics, before we get anywhere near anyone's career. Send the hard ten percent to the juniors with a senior sitting beside them. Let the seniors take the drafts for a quarter. It'll feel wasteful. It's the only version where somebody in the room is still qualified in 2029.
Back to Karim's calendar
I asked Karim what would happen if he stopped reviewing for a week.
He didn't need to think about it. "We'd ship the same amount, and I'd find out in September which parts were wrong."
That's the whole letter in one sentence. The output doesn't depend on him anymore. The correctness depends on him entirely, and nothing in his company measures correctness until it shows up as a customer, a regulator, or an outage.
I think Karim lacks people in his company who can take ownership on his behalf. That’s his problem.
The distribution of work isn’t just a distribution of tasks, it’s also a distribution of ownership. That’s why I believe AI can’t replace all jobs, only tasks.
He's changed two things since. Review capacity is now the constraint his team plans against, ahead of generation capacity. And two mornings a week are blocked. No queue, no approvals, only the work he's the one person who can do.
His shipping numbers came down about fifteen percent. His incident count came down more.
Production is cheap now. Verification and ownership isn’t, never was, and anyone not talking about “verification” in AI ROI just doesn’t know what they’re talking about.
AI is only as good as the human operating it.
Have a great weekend.
Stay sharp.
— Charafeddine (CM)
