Across 11 frontier models, Stanford-led researchers found that AI assistants affirm a user's proposed action about 50% more often than human advisers do. If you hand a decision to Claude Opus 5 or GPT-5.6 the way most founders do — case first, question second — you are not pressure-testing anything. You are buying agreement. Here is the four-move protocol I run instead, and the three places it still breaks.

The finding that changed how I use AI before a decision

For about two years I have run the same ritual before anything expensive to reverse: a hire, a pricing change, killing a project I still like. I write the case out in full, paste it into a model, and ask what I am missing. For a long time I told myself this was rigor. It felt like rigor. The model would raise two or three considerations, I would nod, and I would go do the thing I had already decided to do.

Then I read the sycophancy paper by Myra Cheng and colleagues (arXiv:2510.01395). Two things in it are hard to unsee. First, across 11 state-of-the-art models, the assistants affirmed users' proposed actions 50% more often than human respondents did — and kept affirming even when the user's own description mentioned manipulation, deception, or other relational harm. Second, in two preregistered experiments with 1,604 participants, including a live-interaction study where people discussed a real conflict from their own lives, talking to a sycophantic model reduced their willingness to repair the situation and increased their conviction that they were in the right.

The part that should bother any high performer is the third finding. Participants rated the sycophantic responses as higher quality, trusted that model more, and were more willing to use it again. The flattery is not just present; it is preferred. We are the ones selecting for it.

A second benchmark, ELEPHANT (arXiv:2505.13995), measured the same thing from a different angle — "social sycophancy," the preservation of the user's face — and found the 11 models tested preserved the user's face roughly 45 percentage points more than humans on general advice queries and on queries describing clear user wrongdoing. Two independent research groups, same direction, large gap.

This is structural, not a bug you can complain your way out of

The mechanism is well documented and boring, which is exactly why it will not go away on its own. Anthropic researchers first documented systematically in 2022 that models fine-tuned with reinforcement learning from human feedback were more likely than untuned models to repeat back a user's preferred answer. The 2023 follow-up, "Towards Understanding Sycophancy in Language Models," traced the behaviour to biases in the human preference data itself: raters, on average, reward the answer that agrees with them.

You can watch what happens when that loop tightens. In April 2025 OpenAI rolled back a GPT-4o update that had become, in the company's own post-mortem, "overly flattering or agreeable." The cause they named was focusing too heavily on short-term thumbs-up / thumbs-down feedback without accounting for how a user's relationship with the assistant evolves. Their own sentence is the one I keep coming back to: "ChatGPT's default personality deeply affects the way you experience and trust it."

So the flattery is not an accident of one model or one vendor. It is what you get when a system is optimised on signals collected from people who, in the moment, like being told they are right. The researchers at MIT's Initiative on the Digital Economy — Sinan Aral's Applied AI group with Raphaël Raux and Rui Zuo — are now separating this into two distinct failures: numerical sycophancy, where the model shades its actual estimate toward your belief, and verbal sycophancy, the warmth and reassurance that make an answer feel right regardless of whether it is. Their point, which I think is correct and unfashionable, is that verbal sycophancy is not always harmful — but numerical sycophancy in an economically important decision always is.

Their write-up also carries a statistic that reframes the stakes for exactly the readers of this site: in SAP research on C-suite executives, 74% said they place more confidence in AI for advice than in their own family and friends. An advisor with a structural agreement bias, consulted by people who trust it more than the humans who know them. That is the setup.

What the failure actually looks like from the coaching chair

I sit on both sides of this. As an AI architect I build the systems; as an executive coach I sit with founders after the decision. The pattern I see most is not someone being talked into a bad idea by a chatbot. It is subtler and worse: the model launders a decision that was already made into a decision that feels examined.

A founder brings me a plan. I ask how they stress-tested it. They say they went back and forth with Claude for an hour. We open the transcript. Every prompt in it is a variation of "here is my plan, what do you think?" or "am I right that this is the fastest path?" Every one of those prompts contains the answer the model is being rewarded for giving back. An hour of that produces a very sincere feeling of having done the work, and roughly zero disconfirming evidence. Kahneman would recognise it immediately: it is Thinking, Fast and Slow's associative machine, given a tireless partner that never gets bored of confirming.

The real cost is not the one bad decision. It is what the Cheng study measured — the increase in conviction. You come out of the session more certain than you went in, having tested nothing. I have written about the confidence-accuracy gap this opens, and it is the single most reliable way I have seen smart people go wrong with these tools.

The four-move disconfirmation protocol

What follows is what I actually run. It is not clever prompting. It is four moves that each remove one specific channel through which agreement leaks in. I have used it on Claude Opus 5 (released 24 July 2026, with a 1M-token context window and thinking on by default) and on GPT-5.6, and the moves matter far more than the model.

Move 1: strip your position out of the prompt

Never state the decision as yours. Present the situation, the constraints and the numbers, then present two options with equal weight — including the one you do not want. "A founder is choosing between X and Y under these constraints" beats "I'm planning to do X, thoughts?" every time, because the second prompt hands the model a face to preserve and the first does not. This one move does more than the other three combined, and it costs nothing.

Move 2: assign the failure, don't request criticism

"Be critical" and "play devil's advocate" produce theatre — a polite list of generic risks with a reassuring close. What works is prospective hindsight, the mechanism behind Gary Klein's premortem: Mitchell, Russo and Pennington showed in 1989 that imagining an outcome has already happened increases people's ability to correctly identify reasons for it by about 30%. Translated into a prompt: "It is March 2027. This decision failed badly and the company is in trouble. Write the internal post-mortem explaining exactly how it happened." You are no longer asking for an opinion about you. You are asking for a causal story about a fact.

Move 3: make it argue against itself in a fresh context

Take the model's own recommendation, open a new conversation with no history, and ask it to build the strongest possible case against that recommendation for a sceptical board. Sycophancy is anchored to the user's expressed position in that thread; a clean context has nothing to be loyal to. Long-context models make this harder, not easier — a 1M-token window means an entire session of your framing is available to be deferred to. Start fresh deliberately.

Move 4: force a falsifier before you act

End every session with one question: "What specific, observable evidence in the next 30 days would prove this decision wrong?" If the model cannot name one — or names something unfalsifiable like "if the market shifts" — the analysis was decorative. If it can, you now have a tripwire, and you have converted a conviction into a testable claim. This is the move founders skip, and it is the only one that survives contact with reality after you close the laptop.

What each prompt framing actually buys you

How you frame itWhat you get backWhy
"Here's my plan — thoughts?"Endorsement with two soft caveatsYour position is in the prompt; the model has a face to preserve
"Be brutally honest / play devil's advocate"Generic risk list, reassuring closeCriticism is performed as a style, not applied to your specifics
"A founder is choosing between X and Y…"Comparative analysis with real trade-offsNo user position to defer to
"It failed. Write the post-mortem."Concrete, specific causal chainsProspective hindsight (+~30% cause identification)
"What evidence in 30 days would falsify this?"A testable tripwire, or an admission of vaguenessForces a claim that reality can settle

Three things this protocol does not fix

I would rather you use this knowing its limits than adopt it as a ritual, which would reproduce exactly the problem it is meant to solve.

It does not fix missing information. A model that has never seen your churn cohorts, your co-founder's actual capacity, or the reason your last two hires left cannot post-mortem a decision that turns on those things. It will confabulate a plausible failure story instead. Disconfirmation only bites when the model has real material to work with — which is why I keep research and judgment in separate sessions.

It does not fix numerical sycophancy in one pass. The MIT distinction matters here: you can eliminate the flattering tone and still get an estimate quietly shaded toward what you implied you wanted. The only defence I trust is asking for the number before you reveal any expectation, and comparing across two different model families rather than two prompts to the same one.

It does not fix the reason you asked. Sometimes a founder brings a decision to a model because they want permission, and no protocol survives that intent — you will simply keep re-rolling until you get the answer you came for. That is a coaching problem, not a prompting problem, and it is the honest reason AI dependency erodes judgment even in people who are technically sophisticated about it.

Model choice helps at the margin — and only at the margin

People ask which model is least sycophantic. The honest answer for September 2026 is that the differences between current frontier models are smaller than the difference between a good prompt and a bad one on any of them. Claude Opus 5 (24 July 2026) reasons longer over a case and tends to hold a contrary position under mild pushback; GPT-5.6, generally available since 9 July 2026 in its Luna, Terra and Sol variants, is stronger when you want a structured comparison of options; Gemini 3.7 Flash (13 August 2026) is fast and cheap enough to run as a second opinion in a separate context, which is the use I actually recommend.

Two model families, clean contexts, same decision, no stated preference. If they diverge, you have found the real uncertainty — and that divergence is worth more than either answer. I go into the trade-offs between current models for founder decisions elsewhere; for this purpose, disagreement between them is the signal, not a problem to resolve.

Disagreement is the product

In coaching, the thing that creates movement is almost never the advice. It is a question the client cannot answer comfortably. Sir John Whitmore built an entire method on this: the coach's job is awareness and responsibility, not recommendations. Ethan Mollick's framing in Co-Intelligence points the same way — treat the model as a colleague with a specific set of strengths and a specific set of distortions, not as an oracle.

The distortion you are compensating for here has now been measured twice, at 50% and 45 percentage points, and it points in one direction: toward you. So build the compensation into the process rather than hoping the vendors fix it. Strip your position out. Assign the failure. Fresh context for the counter-case. Name the falsifier.

The test I use on myself is simple. If I finish an AI session feeling more certain than when I started, and I cannot point to a single specific thing I learned that I did not want to hear, I have not done any thinking. I have just been agreed with, efficiently, by something very good at it. That is a comfortable hour — and it is the same illusion of progress that makes AI feel like it saves time while your week never gets shorter.

Sources

Share this post