Current AI agents are dependable only for structured, repetitive tasks with clear criteria — data extraction, invoice follow-ups, scheduling. For strategic thinking, judgment calls, and meaningful decisions, agents remain unreliable. Ethan Mollick's Co-Intelligence framework of being the human in the loop describes the right stance: use agents aggressively for predictable tasks but stay in the loop for anything requiring judgment.
Earlier this month, Mark Zuckerberg told Meta staff that AI agent development is going slower than expected, and the Hacker News thread exploded with over 600 comments — most of them from builders who had run head-first into the same wall. Around the same time, a developer who looked under the hood of Garry Tan's claim to ship 37,000 lines of AI-generated code per day found that the reality was considerably less impressive than the headline. I have been building with AI agents since early 2024, and this moment feels like a necessary reckoning. The hype cycle is colliding with a harder truth: current AI agents are not reliable enough to delegate genuinely important work to, and knowing that matters more than pretending otherwise.
Ethan Mollick's Co-Intelligence gets at why this gap exists. Mollick argues we should treat AI as a collaborator — always invite it to the table but stay in the loop as the human. The "be the human in the loop" principle is easy to ignore when every demo shows an agent autonomously booking flights, filing expenses, and updating CRMs. But those demos fracture the moment the agent encounters something it has not seen before: an ambiguous email, a spreadsheet with an unexpected column, a decision that requires judgment rather than pattern matching. I have tried running Lindy with its 1,600 integrations, and it does handle repetitive tasks well — following up on invoices, scheduling standard meetings, triaging support tickets. It falls apart on anything requiring genuine contextual understanding, because the underlying LLMs still struggle with what Mollick calls "the long tail of edge cases."
The useful question is not whether AI agents are good or bad. It is where they are actually dependable in mid-2026 and where they are not. After testing tools like Lindy, Manus for autonomous research, and various custom n8n agent workflows, here is my practical split. For highly structured, repetitive workflows with clear success criteria — think "extract this field from these documents and put it in this spreadsheet" — agents are genuinely productive. I have a Manus research agent that aggregates competitive data from public sources and drops it into a structured format, and it saves me about four hours a week. For anything involving judgment, negotiation, or reading between the lines — the core work of a founder or executive — agents are a liability. The Zuckerberg Reality (that agents are slower to mature than optimists predicted) is not a bug report. It is the honest signal we should have been listening to all along.
My current approach is a deliberate two-track system. For shallow, high-volume tasks that follow clear patterns, I use agents aggressively and I measure my time savings. For strategic thinking, client conversations, and any decision with meaningful stakes, the AI stays in its lane as a thinking partner — I write my initial analysis, ask it to pressure-test my assumptions, and then make the call myself. This dual stance is uncomfortable because it requires the discipline to know which track you are on at any moment. But that discipline is exactly the muscle Mollick is describing when he says the future belongs to people who can work with AI without being captured by it. The agents will get better. Right now, in July 2026, the founder who knows what they cannot delegate has the real edge.
To be concrete: I recently tested an agent for negotiating a vendor contract renewal. The agent parsed the key terms flawlessly, even suggesting alternative payment structures. But when the vendor responded with an unexpected clause about data retention — something the agent had no training data on — it suggested accepting terms that my legal team later flagged as problematic. That single test cost me more time in cleanup than I saved in automation. This is the pattern I see everywhere: agents excel at the predictable, break on the novel, and the gap between those two categories is where founders actually earn their keep.
Related: How to Find Your Passion · Best Self-Improvement Books · How to Make Better Decisions · Why Exploration Is Important for Success
