AI is saving only about 3% of most workers' hours because the individual gains — 5-11 hours per week per user — are being burned back on "botsitting," rework, and lower-quality output that Kahneman's System 1 can't catch. Meta CEO Mark Zuckerberg admitted on July 2, 2026 that agent progress has stalled. The fix is not more AI; it's designing System 2 checkpoints around it.
I have been running AI in my own work for three years and coaching founders through their AI stacks for two of those years, and the July 2026 productivity numbers finally match what I see in every founder session: the gap between AI's promise and AI's arrival at the bottom line is enormous, and it is not closing.
Here is the current picture, dated. A Boston Consulting Group study (June 2026) found that over 40% of regular AI users in white-collar non-managerial roles save a full workday or more per week — but organisations struggle to convert that into measurable value. Workday's global research (January 2026) reported 85% of employees saving one to seven hours a week, while nearly 40% of the saved time is lost to rework because of low-quality AI output. The Work AI Institute survey (June 2026) is the sharpest number: individuals save ~11 hours a week and then spend over six hours "botsitting" — checking, correcting and rerunning AI output. Net at the organisation level? The oft-cited figure floating around HN this month is ~3% of hours saved, and almost none of it reaches the money. On July 2, 2026, Mark Zuckerberg told Meta staff at an internal town hall that agent progress has been slower than he expected over the last four months, and that the restructure was not "clean" — this is the CEO of the company spending up to $145B on AI infrastructure in 2026.
What is actually happening (the System 1 trap)
Daniel Kahneman's Thinking, Fast and Slow maps this cleanly. AI writes with cognitive ease — clean fonts, confident cadence, coherent stories — which is precisely the condition under which System 1 takes over and System 2 stops checking. WYSIATI, "what you see is all there is": if the AI's answer looks complete, we don't notice what's missing. So the founder saves 45 minutes drafting an investor update, then spends 40 minutes correcting a hallucinated metric plus 15 minutes rewriting the tone. The experiencing self logged a 45-minute save; the remembering self, by peak-end rule, remembers "I used AI, it was fast," and reports that to the McKinsey survey. Duration neglect does the rest. This is how you get an 11-hour individual save that shows up as 3% at the org level: the botsitting is real work but it is invisible to the person doing it.
Where the leaks actually are
Across ~40 founder sessions this year, three leaks account for most of it. First, the review tax: AI output is 80% right, so 100% needs reading, and reading a plausible-looking draft is slower than reading a rough human draft you know to distrust. Second, quality regression: the LA Times reported (June 12, 2026) that AI-generated content is often subtly worse, and the recipient (an LP, a customer, a hiring committee) picks up the tell within a paragraph. Third, displacement, not automation: people fill saved hours with more low-leverage work, not with the strategic thinking that would compound. Judgment does not scale by adding more drafts.
The System 2 checkpoint protocol I actually run
This is the four-step workflow I have kept for the last six months — the only version that has stopped my own botsitting drift. It is deliberately slow at the check-points because that is where the value is.
- Name the decision, not the task. Before opening ChatGPT or Claude, write one sentence: "What decision does this document change?" If there is no decision, do it faster (short email, bullet, one-liner). AI is a decision accelerator; on non-decisions it is a cost.
- Do the outline in your own head first. Five minutes, no screen. This is Kahneman's slow-System-2 anchor — you now have your own frame to compare the AI draft to, which is the only way to notice what it silently omitted.
- Ask for the counter-argument, always. After the first draft, prompt: "What is wrong with this? What would the sharpest critic say?" This borrows Kahneman's premortem: imagine this document already failed — why? Roughly half the time the AI surfaces a real hole its own first draft ignored.
- Log the botsitting time honestly for a week. Use a note app or a stopwatch. Most founders I've done this with are shocked: the ratio of AI-drafting time to rework time is often 1:1.5 for anything customer-facing. The number is the intervention.
The honest limit: this workflow only converts AI hours into value when the underlying decision was going to be made anyway. If you use AI to generate decisions (more content, more emails, more analyses that nobody asked for), the 3% number is generous — you are on the wrong side of leverage. Zuckerberg's July 2 admission is worth sitting with: the model does not yet do agentic strategy. What it does is compress specific, well-scoped tasks. Treat it that way.
What this changes about how I work
I have stopped measuring AI ROI in hours. I measure it in decisions changed — did the output alter what I did, hire, ship or refuse? If the answer is no, that hour was botsitting theatre. This is the boring, honest reframe the July 2026 studies have finally forced on the industry, and it is much closer to how coaching frames performance: the number that matters is not effort, it is what shifted.
Sources
Straits Times, "Meta's Zuckerberg says AI agent tech progressing slower than expected" (July 3, 2026). TechCrunch / PYMNTS coverage of Zuckerberg's July 2, 2026 Meta town hall. Boston Consulting Group, "AI at Work" (June 2026). Workday, "Companies Are Leaving AI Gains on the Table" (January 2026). Work AI Institute survey (June 2026). LA Times, "AI saves office workers hours but then demands hours of babysitting" (June 12, 2026). Daniel Kahneman, Thinking, Fast and Slow (2011).
Related: How to Find Your Passion · Best Self-Improvement Books · How to Make Better Decisions · AI Coach App — Building It in 8 Hours
