We trust confident speakers because, as Daniel Kahneman showed, confidence is a feeling produced by a coherent story, not by evidence, and fluency creates the cognitive ease that switches off scrutiny. AI is built for that fluency: a Frontiers in Psychology paper (September 9, 2026) treats miscalibrated chatbot trust as a design problem, and on AA-Omniscience GPT-5.6 Sol gives a wrong answer instead of refusing 92.2% of the time.

A founder I work with recently pasted a market-size figure from ChatGPT straight into an investor deck. The number was wrong by roughly an order of magnitude. When I asked why he had not checked it, his answer was honest and, I think, universal: "It didn't sound unsure." That sentence is the whole problem, and it is older than AI.

Daniel Kahneman spent a career documenting why. In Thinking, Fast and Slow he shows that subjective confidence "is a feeling, not a judgment," produced by the coherence of the story System 1 can assemble from whatever is in front of it, a pattern he calls WYSIATI, what you see is all there is. Confidence tracks story quality, not evidence quality. Kahneman also documents cognitive ease: fluent, repeated, easy-to-read statements are judged more true, because ease of processing is read as a truth signal. A confident speaker gives you both at once, a coherent story and effortless delivery, so System 2 never wakes up to ask what is missing. That is why the smooth presenter beats the accurate one.

What is new in 2026 is that the most fluent speaker in most people's day is now a machine, and it is fluent on purpose. A paper published in Frontiers in Psychology on September 9, 2026, "Why we believe chatbots: trust calibration as a design problem," lays this out. Users judge chatbot answers by fluency, confidence, and speed, because the interface reveals almost nothing about how the answer was produced or what supports it. The authors note that standard evaluations reward confident guessing over admissions of uncertainty, so models deliver wrong answers in the same assured register as right ones, "leaving users no linguistic cue to tell them apart." They propose fixes such as source citations, uncertainty cues, and cooling-off periods before an answer can be copied. Their phrase for the current state is that users are given strong cues to believe and weak means to check.

A review summarized this week by AOFIRS reports that across GPT, LLaMA-2 and Claude models, highly confident answers were wrong 47% of the time, yet users accepted them 90% of the time. On the AA-Omniscience benchmark, which measures whether a model answers or declines when it does not know, Suprmind's August 27, 2026 update lists GPT-5.6 Sol at a 92.2% hallucination rate on the questions it gets wrong, and Claude Fable 5 at 63.6%; 51.4% of Gemini's high-confidence answers were contradicted by another model, against 26.4% for Claude. The newest flagships, Claude Fable 5.1 (September 1, 2026) and GPT-6 Astra (September 3, 2026), are more capable, but nothing in their training changes the basic incentive: an answer that sounds certain gets rewarded.

I build these systems and I still get caught. My own failure mode is not believing obviously wrong things; it is skipping verification on plausible things because the tone did the work. So I stopped trying to feel the difference, which Kahneman says is impossible, and started discounting confidence mechanically, using the table I keep above my desk.

Cue that raises trustWhat it actually signalsCounter-move
Fluent, well-structured answerCognitive ease; no information about accuracyAsk the same question with the opposite framing and compare
No hedges ("I think", "it is possible")Training that rewards guessing over decliningRequire a stated confidence and one source per number
Visible "thinking" or searching stepsLabor illusion; reasoning may be post-hocOpen one cited source before using any figure
Agreement with my draftSycophancy; models retract correct answers when challengedRun a Kahneman premortem: "assume this is wrong, why?"
Answer arrives in two secondsSpeed is a fluency cue, not a quality cueImpose my own cooling-off: no copy-paste into a deck same day

The premortem row is the one that pays most often. Applied to an AI answer it becomes a single prompt, "Assume this figure is wrong. List the three most likely reasons," and the model, which will confidently argue either side, suddenly produces the caveats it omitted the first time. On the deck number above, that one prompt surfaced the confusion between global and regional market size in under a minute.

Where this fails: discounting confidence does not make you more accurate, it only stops you from being wrong faster. You still need base rates, Kahneman's outside view, and someone in the room with real domain knowledge. The table also costs time, and on low-stakes questions I skip it, which is the correct trade. And there is a human cost in reverse: once you train yourself to distrust fluent certainty, you will notice how much of your own persuasion relies on it. That was the real coaching conversation with the founder, and it was worth more than the corrected number.

Sources: Frontiers in Psychology, "Why we believe chatbots: trust calibration as a design problem" (published September 9, 2026) — frontiersin.org; AOFIRS, "Why Your Brain Falls for AI Hallucinations" (September 10, 2026) — aofirs.org; Suprmind, "AI Hallucination Rates, Statistics & Benchmarks in 2026" (updated August 27, 2026) — suprmind.ai; Mungomash, "The frontier AI models, right now" (September 8, 2026) — mungomash.com; Daniel Kahneman, Thinking, Fast and Slow (2011).


Related: How to Find Your Passion · Best Self-Improvement Books · How to Make Better Decisions · What University Will Not Teach You