Every founder I coach has run the same private experiment by now. They gave the whole team ChatGPT, or Claude, or a Copilot seat, watched a few people get visibly faster at writing and drafting, and then waited for the number that never came: the one on the P&L. The revenue line did not move. The margin did not widen. And nobody could quite explain where the promised productivity went.

I spent most of 2026 assuming that was a failure of implementation on my clients' part, or on mine. Then I read the study that reframed the whole thing for me, and I want to walk through it here, because it is the single most clarifying piece of evidence I have seen on what AI actually does to your performance and your income. The short version: AI genuinely makes you faster at specific tasks, the gain across a real job shrinks to about 3% of your hours, and almost none of that reaches your pay unless you deliberately go and collect it. That last clause is the entire game, and it is the part every vendor demo skips.

The study that linked AI use to actual paychecks

Most AI productivity research measures a task. Someone writes a press release 40% faster in a lab; a support agent closes 14% more tickets in an hour. Those numbers are real, and I will defend them below. But they measure a twenty-minute slice of work, not a career, and certainly not a bank account.

Two economists, Anders Humlum at the University of Chicago and Emilie Vestergaard at the University of Copenhagen, did the thing almost nobody else did. In their NBER working paper "Large Language Models, Small Labor Market Effects" (w33777), they linked AI-adoption surveys covering roughly 25,000 workers across about 7,000 Danish workplaces to actual government payroll records. Not self-reported vibes. Tax data. And their finding, written in a sentence that should be printed on the wall of every company that just bought enterprise AI seats, was that "AI chatbots have had no significant impact on earnings or recorded hours in any occupation." The measured time savings were real but small, around 2.8% of work hours, call it an hour a week, and only 3 to 7% of that productivity gain showed up in anyone's pay.

Hold that against the demo reel. In the lab, AI makes people 15%, 40%, sometimes 55% faster. Both things are true, because they are measuring different worlds. In a controlled short task, AI is a rocket. Across a real month, on a real payroll, it is a small, leaky gain that mostly evaporates before it reaches a paycheck or a profit line. The right question was never "does AI make knowledge work faster," because on the right task it plainly does. The right question is "does the time you save turn into money," and the honest answer, so far, is: only if you make it.

The gains are real — on a narrow, specific set of tasks

I want to be precise here, because the anti-hype position gets lazily flattened into "AI does nothing," and that is not what the evidence says. The task-level wins are genuine and, on the right work, large. In a randomized experiment with 453 professionals published in Science, giving people ChatGPT for mid-level writing, press releases, short reports, sensitive emails, cut the time spent by 40% and raised graded quality by 18%. In a field study of 5,179 customer-support agents, an AI assistant lifted resolved issues per hour by 14% on average, and by roughly 34% for the newest, least experienced agents.

Notice what every one of those studies actually measures: a task. A writing assignment done faster. A ticket closed quicker. None of them measured whether the freed-up time became income. That gap between "task faster" and "richer" is not a rounding error. It is the whole story, and it maps almost perfectly onto a book I keep handing to clients.

Why the hours leak: Newport's pseudo-productivity

Cal Newport, in Slow Productivity (2024), names the disease directly: pseudo-productivity, the habit of using visible activity as a proxy for useful output. For a century the knowledge economy has had no good way to measure real value creation, so it measured motion instead, emails sent, hours logged, meetings attended. AI is the most powerful pseudo-productivity engine ever built. It lets you generate more motion, faster, than any tool in history: more drafts, more replies, more decks, more summaries. If your organization rewards motion, AI will hand you an infinite supply of it, and none of it will touch the number that matters.

This is the mechanism underneath the payroll data. The hour you save with AI does not vanish, and it does not automatically become profit. It flows to the path of least resistance, which in almost every company is more low-value activity: answering more email, attending one more sync, producing a report nobody requested. Newport's builder-coach point, and mine, is that saved time is not a benefit. It is a raw material. Left alone, it converts back into busywork at roughly a one-to-one rate. The Danish payroll records are just what that conversion looks like when you measure it at national scale.

There is a sharper version of this that the 2026 data has started to name: "workslop." Stanford and BetterUp researchers documented AI-generated output that looks like finished work but is actually low-effort and low-quality, so a colleague downstream has to spend their reclaimed hour cleaning it up. A Workday analysis found that nearly 40% of the value from AI-driven speed is lost to exactly this, correcting low-quality output and resolving the confusion it creates. The time did not disappear into laziness. It disappeared into rework, one person's saved hour becoming another person's lost one.

The jagged frontier: where the saved hours turn negative

There is a second, more dangerous way the gain leaks, and it is the one I watch for most closely with clients. Ethan Mollick calls it the "jagged frontier" in Co-Intelligence: AI's competence has a strange, invisible shape, and stepping off the edge is expensive. A Harvard and BCG field experiment with 758 consultants captured it in two numbers. On tasks inside AI's range, the consultants using GPT-4 finished 12.2% more work, moved 25.1% faster, and produced output rated more than 40% higher in quality. Then the researchers handed everyone a task deliberately chosen to sit just outside AI's range. On that one, the people using AI were 19 percentage points less likely to reach the correct answer than the ones working without it.

That is the trap in a single statistic. AI does not announce when a task has crossed its edge. It answers just as fluently, just as confidently, and wrong. A confident wrong answer costs far more to catch than a right one saved, and if you do not catch it, it costs more still. So the real distribution of AI's effect on your work is not "small positive." It is large positive on a narrow band of tasks, roughly zero on most, and sharply negative on the tasks where it fails silently. Averaged across a real job, you get the Humlum and Vestergaard result: a small, leaky 3%.

A framework: the Capture Test

So what do you actually do with this? For the past six months I have run every proposed AI use, my own and my clients', through a four-part filter before deciding whether it is worth the license fee and the oversight. I call it the Capture Test, and it is built directly from the evidence above. A use of AI only pays if it clears all four.

TestQuestionWhy it matters
FitIs this task inside AI's range, or near the jagged frontier?Off the frontier, AI is 19 points worse than nothing (Harvard/BCG).
VolumeIs the task high-frequency and repeatable?A 40% win on a rare task rounds to nothing across your week.
CaptureWhere does the saved hour go, by name?Unbanked time leaks back into busywork (Humlum/Vestergaard).
QualityWho checks the output, and is that cheaper than the win?Rework destroys ~40% of the speed gain (Workday).

The one that changes behavior is Capture. Most people, including me at first, cannot answer it. "I'll have more time" is not an answer. "I will use the two hours I save each week on support drafts to take on one more client, or to ship the feature that unblocks the enterprise deal, or to end my day at six instead of eight" is an answer. If you cannot name where the hour goes before you adopt the tool, the payroll data predicts exactly where it will go: nowhere you can measure.

What I actually changed

Three things, concretely. First, I stopped counting AI adoption as a win and started counting captured hours. My clients now write down, per AI workflow, the specific higher-value activity the reclaimed time is being redeployed into. If they cannot name it, we do not roll the tool out, because we now have national payroll evidence for what happens next.

Second, I map the jagged frontier explicitly for each role before we deploy. Which tasks are inside AI's range (structured drafting, summarizing, first-pass code, boilerplate replies) and which sit at the edge (novel strategy, anything requiring tacit judgment, high-stakes external communication)? The inside-range tasks get AI hard. The edge tasks get a human first and AI only as a check, never as the author.

Third, and this is the coaching half of the builder-coach lens, I treat the saved hour as a decision, not a gift. Newport's discipline of doing fewer things at a higher quality is not compatible with letting AI quietly refill your day with more things. The founders who actually get richer from AI in my client base, and there are a few, are not the ones with the most seats. They are the ones who used the tool to kill a task entirely and then guarded the empty space it left, instead of letting it fill back up.

AI will make you faster. The evidence on that is settled and I use it every day. It will not make you richer on its own; the payroll records are blunt about that. The gain goes to whoever deliberately banks it, and mostly, still, no one does. The edge is not in having the tool. It is in being the rare person who decides, in advance and by name, what the time it saves is for.

For the deeper practitioner mechanics behind this, see my related field notes: Does AI actually make knowledge workers more productive?, Healthy versus unhealthy productivity, How executives use AI to reclaim ten hours a week, Can AI improve your judgment, or just your confidence?, How to use AI without losing your gut feel, and How to stay sharp when AI does the work. See also my post How founders actually learn new skills with AI in 2026.

Sources

Anders Humlum & Emilie Vestergaard, "Large Language Models, Small Labor Market Effects," NBER Working Paper 33777 (2025), nber.org/papers/w33777; Fortune coverage, "Study looking at AI chatbots in 7,000 workplaces finds 'no significant impact on earnings or recorded hours in any occupation'," May 2025; Shakked Noy & Whitney Zhang, "Experimental evidence on the productivity effects of generative AI," Science, 2023 (453 professionals); Brynjolfsson, Li & Raymond, "Generative AI at Work," NBER w31161 (5,179 support agents); Dell'Acqua et al., "Navigating the Jagged Technological Frontier," Harvard Business School / BCG working paper 24-013 (758 consultants); Cal Newport, Slow Productivity (2024); Ethan Mollick, Co-Intelligence (2024); Kate Niederhoffer et al. (Stanford / BetterUp), "workslop" research, Harvard Business Review, 2025; Workday, "The AI time-savings paradox," 2026.

Share this post