In some workflows, yes. Microsoft and Uber both reported in May 2026 that token-heavy agents cost more per completed task than a salaried employee doing the same work. The lesson, echoing Davenport and Mittal in All-In On AI, is that the cheap demo and the production unit-economics are two different conversations.
The Fortune piece that crossed my desk in late May made a quiet but important admission: at Microsoft, internal benchmarks now show that running an autonomous agent end-to-end on a real business task — research, drafting, multiple tool calls, retries — often costs more than paying a person to do the same work. Uber reportedly burned through a full year of AI budget in four months, not because the models were broken but because nobody had modeled the tokens an agent actually spends when it loops, second-guesses, and re-reads its own context. I spent a week with my own logs after that came out, and the result was uncomfortable: two of the four agentic workflows I had been quietly proud of were, on a per-task basis, more expensive than hiring a part-time contractor in the Baltics.
The mechanism is easy to miss because every individual call looks cheap. A single Claude or GPT call is a few cents. The trap is that a useful agent is not one call. It is a planning step, a tool-use step, a re-plan when the tool returns nothing, a long context window that grows with every retrieval, and often three or four iterations before the output is acceptable. Anthropic's published Claude cost-per-million-token figures from the second half of 2025 told you what one mouthful costs, not what a meal costs. When you instrument the meal, a "five-cent" agent task quietly becomes a dollar and change. Multiply by ten thousand runs a month and you are now paying more than the human you replaced, who at least came with judgment and accountability.
Daniel Kahneman would call this a base-rate failure. We anchor on the dazzling first demo and never update on the boring tail of production runs. I see founders do this constantly. They watch an agent finish one beautiful customer-research report in four minutes, extrapolate that experience to a workforce of digital employees, and skip the unit-economics math entirely. The math, when you do it, is unglamorous. You need to compare the fully-loaded human cost — not just salary, but tools, management overhead, and lost output during retraining — against the fully-loaded agent cost, including failed runs, retries, evals, and the engineer who maintains the prompt. In most knowledge work, the human still wins on cost per acceptable outcome, especially once you weight by quality.
Where agents do win, and the Davenport and Mittal All-In On AI thesis still holds, is when the work has three properties: it is repetitive enough that quality variance is low, the input is structured, and the cost of one bad output is small. Classifying inbound support tickets, summarizing meeting notes, drafting first-pass cold emails — these are the agent's natural turf because retries are cheap, judgment is bounded, and a human can rubber-stamp at the end. The mistake is dragging agents into the opposite kind of work: open-ended research, judgment-heavy decisions, anything where a wrong answer compounds. There, the agent's token spend climbs and the value it produces falls, because the cheap part of the work — the typing — was never the bottleneck. The bottleneck was thinking, which is still where humans are dramatically cheaper per useful insight.
The practical move, the one I now run with every founder I coach, is the unit-economics question before the build. For a given workflow, what is the cost per acceptable output of the human-only baseline? What is the realistic cost per acceptable output of the agent version, with retries and evals included? And what is the cost of being wrong? If the agent does not beat the human on at least two of those three numbers, you do not have an agent project. You have an enthusiasm project. Dorie Clark's strategic-patience framing in The Long Game is useful here: short-term, you can ship anything with an agent; long-term, the workflows that survive are the ones whose unit economics quietly improve as model prices fall, which they will, but unevenly and not on your fundraising timeline.
None of this means stop building. It means build with eyes open. The Microsoft and Uber numbers are not a referendum on AI; they are a referendum on cost discipline. The companies that will pull ahead in 2026 are not the ones with the most agents — they are the ones who killed the agents that lost money and reinvested the savings in the agents that compounded. That is harder than it sounds, because killing your own experiment hurts. But the alternative is the founder I spoke to last month who realized, eight months in, that her autonomous research stack cost three times what a graduate student would have charged, and produced output her advisors trusted less. The agent was not the villain. The missing spreadsheet was.
Related: How to Find Your Passion · Best Self-Improvement Books · How to Make Better Decisions · AI Coach App — Building It in 8 Hours
