Two weeks ago a founder I coach was choosing a customer-support platform for a twelve-person team. She did what most operators now do first: she asked an AI assistant. Then, because she is careful, she asked two more. She got three different shortlists. Two of them contained a vendor she had never heard of, and under one of those answers sat a small labelled card for that same vendor. She sent me the screenshots with one question: "So which of these is the actual answer?"
None of them, and the reason none of them is the answer is the useful part of this post. I build AI systems for a living and coach the people who buy software with them, and the vendor decision is where I now see the most confident, least examined use of AI in a founder's week. It was already a weak spot before the assistant started carrying ads. It is a weaker one now.
What changed, with dates
ChatGPT has carried ads for logged-in adults on the Free tier and the $8-a-month Go tier since the US pilot in February 2026; self-serve buying opened on May 5, and placements reached 31 European markets on August 24. On September 16, 2026, OpenAI announced Sponsored Agents: after an ad appears, a user can click into a clearly labelled conversation with an agent sponsored by the advertiser, ask follow-up questions, and follow a link to the business's site. OpenAI says the sponsored conversation is "distinct from ChatGPT's independent answers and separate from the original conversation," and that the test is limited to select advertisers in the United States. The same announcement wired ChatGPT Ads into HubSpot and Shopify as the first CRM and ecommerce partners. A week later, on September 23, OpenAI added seven Asian markets, said the product now runs in over 60 countries, and repeated that ChatGPT Ads had reached a $1 billion annualised revenue run rate in under 200 days, with tens of thousands of advertisers.
How often you actually meet one: Similarweb's AI Ads dataset, released August 17, 2026, found that 26% of ChatGPT responses shown to Free and Go users already carried a sponsored placement, and that nearly 30% of ad-eligible Google AI Mode queries showed ads too. OpenAI's stated principles are unambiguous and I take them at face value: "Ads do not influence the answers ChatGPT gives you," ads are labelled and separated from the organic answer, and Plus, Pro, Business and Enterprise subscriptions carry none. Anthropic has gone further and committed that Claude will remain ad-free: no sponsored links beside conversations, no third-party product placements in responses.
So the honest summary is: if you pay for a business tier, or use Claude, the ad problem as narrowly defined does not reach you. I wrote up the trust question itself in can you still trust ChatGPT's recommendations now that it has Sponsored Agents. This post is about the part that reaches everyone, paid or not, and that the ads only made visible.
The ads are the visible part of an older problem
Here is what I know as someone who builds these systems. An assistant's "organic" recommendation of a payroll tool or a support desk is not drawn from some neutral registry of software. It is drawn from the web, and the web's coverage of business software is overwhelmingly written by three parties: the vendors, the affiliate comparison sites paid by the vendors, and the analysts paid by the vendors. A model answering "what is the best help desk for a small team" is compressing a corpus that was marketing before the model ever saw it. Ads inside the answer window did not introduce commercial pressure into vendor recommendations. They put a label on a room that was never neutral.
The Sponsored Agent adds something genuinely new, though, and it is worth being precise about what. It is a salesperson with a chat interface, running on a model you did not choose, under instructions you cannot read, written by the party that wants the sale, sitting one click away from the assistant you were using to think. It is labelled, and I believe the label. But the German technology site MIXED made the point that matters: OpenAI's assurances are "statements of policy from the company selling the ads, not an audited finding." You would not let a vendor's rep sit inside your strategy meeting on the strength of their promise to stay quiet. The interface makes it very easy to do the equivalent without noticing you have.
Why founders are unusually exposed here
This is where the coaching side of my work earns its keep, because the mechanism is not technical. Automation bias, our tendency to accept an automated system's output over contradictory evidence and our own judgement, has been documented for decades; in Mosier and Skitka's classic aviation study, participants working with an automated aid that missed things achieved 59% task accuracy against 97% for the group without it. A 2026 paper in Philosophy & Technology reviewed newer experiments showing users following AI advice even when it contradicted both the available contextual information and their own prior assessment. Every one of those subjects was, presumably, trying to get the answer right.
Kahneman gave the sharper name for what happens inside a shortlist: WYSIATI, what you see is all there is. Once an assistant has produced five vendors, those five become the universe of the decision. Nobody asks about the sixth. The founder I opened with was not evaluating the market for support platforms; she was evaluating the intersection of three shortlists, and she had no idea how any of them had been generated. An ad card under one answer is a tiny nudge. The shortlist itself is the large one, and it was there before the ad.
And there is the founder-specific pressure. A vendor decision is an anxious decision: it has switching costs, it touches money and customers, and it interrupts the real work. Anxiety wants an authority to close it. A confident, fluent, instant answer is exactly that authority, which is why I keep sending people back to how to know when to trust an AI's confident answer. The problem is not that the assistant is wrong. It is that it is confident in a domain where its raw material is advertising, and you are in a hurry.
The Four-Room Protocol
What I actually do, and what I now ask every founder I coach to do for any software decision with real switching costs, is to keep four jobs in four separate rooms and never let the assistant be in more than one of them at a time. The rooms are Map, Score, Attack, Verify. The AI is useful in three of them. It is banned from the one that matters most.
| Room | The job | Where AI helps | What I never delegate |
|---|---|---|---|
| 1. Map | Build the long list: 10-15 candidates, not 5 | Ask for every vendor in the category with pricing model, hosting region, and API, in a table, no ranking. Run it on two models, one of them ad-free (a paid tier or Claude). | Accepting a top-five. A ranked list at this stage is the shortlist deciding for you. |
| 2. Score | Write the criteria and weights before seeing any recommendation | Read the 40-page terms and the pricing page; extract exit terms, data-export clauses, per-seat traps. Draft the scorecard from my constraints. | The weights. What matters to a twelve-person team in Tallinn is not in any corpus. |
| 3. Attack | Red-team the top two | "Argue that the top pick is the wrong choice for a team like this." "What would this vendor not want me to know before signing?" "Which criterion did I forget?" | Letting the same conversation that recommended it defend it. Fresh session, or a different model. |
| 4. Verify | Primary sources and a human | Nothing generative. At most, help me draft the questions for a reference call. | The pricing page, the status page, the changelog, the contract, and one conversation with an actual customer. These are never delegated. |
The order is the point. Most founders run Map and Score in one prompt ("what should I use for X?") and skip Attack and Verify because the answer arrived with enough confidence to feel finished. The protocol forces the criteria to exist before the candidates do, the only reliable defence against WYSIATI I know, and it forces the recommendation to survive a hostile pass in a room where the assistant has no memory of having made it.
The prompts, as I run them
Map. "List every vendor in [category] that a [size] team in the EU could reasonably use. For each: pricing model, minimum contract, data residency, native integrations with [my stack], and the date of the last major release. Table. Do not rank, do not recommend, and flag any entry you are not certain still exists." That last clause matters more than it looks; I covered the reason in does an AI model's training cutoff still matter if it can search the web. Even a current model with search, GPT-6 Astra or Claude Fable 5.1, will happily list a product that was acquired and shut down eighteen months ago, because the marketing corpus about it is still there.
Score. I paste the actual pricing pages and terms of the top six, not the model's memory of them, and ask for a table of exit costs: notice period, export format, whether my data is deleted on termination, and any clause that changes price on renewal. The model is excellent at this. It is reading, not recommending.
Attack. New session, ideally a different model, with the scorecard and the top two pasted in. "You are the founder who chose [runner-up] over [top pick] and was right. Make the case." Then: "Name three ways the scorecard is biased towards [top pick]." The general version of this move is in how to get AI to challenge your decisions and, for the two or three decisions a quarter that really matter, in how to pressure-test a big decision with AI before you commit.
Verify. Open the pricing page yourself. Open the status page and look at the last 90 days. Read the changelog and see whether the last entry is from this quarter or from 2024. Ask the vendor for two customers at your size and talk to one. This is the part no assistant, sponsored or not, can do for you. Which decisions should be kept away from AI altogether is a separate question I took up in when a founder should not use AI to make a decision.
Which assistant for which room
I get asked this constantly and the honest answer is less about model quality than about incentive. For Map and Attack I want either an ad-free surface or a paid tier, and I want two of them, because the disagreement between two long lists is more informative than either list. As of today that means a ChatGPT business tier or Claude for the mapping pass; a Free or Go account is fine for reading terms in Score, because the ad has nothing to attach to when the input is a contract you pasted. The research-tool comparison I keep current is in Perplexity vs Claude vs ChatGPT for executive research, and the September model releases are compared for decision work specifically in which new AI model a founder should use to pressure-test a big decision.
If you click into a Sponsored Agent, treat it as exactly what it is: a call with sales. Useful for facts about the product, useless as a judge of whether you need it, and everything it says goes into the Verify room, not the Score one.
Where the AI is genuinely better than I am
I want to be fair to the tool, because the protocol can read as distrust and it is not. In the Score room the assistant beats me every time. It reads the entire terms of service; I skim. It notices the auto-renewal clause on page 31 and the per-seat minimum that only applies from the second year. It converts four vendors' incompatible pricing pages into one table with my actual seat count and usage plugged in. It drafts a migration-risk memo from the export documentation in minutes that would have taken me an afternoon. Skipping AI here out of ad-suspicion pays a real cost to avoid an imaginary one.
And in the Attack room it is a better adversary than most humans I could ask, precisely because it has no stake in the answer, provided it is not the same conversation that made the recommendation. A model defending its own earlier output is the most sycophantic thing in software.
What this costs
Time, obviously. The protocol turns a ten-minute question into two or three hours spread over a few days, and for a $30-a-month tool with no data lock-in that is absurd. So the rule is switching costs: if leaving would take more than a day of work or involve customer-facing data, run the four rooms. If not, ask once, buy, and move on; over-processing small decisions is its own failure.
The other cost is that you lose the feeling of a settled answer. Three shortlists and a scorecard do not produce the clean "use this one" that the first prompt did. That feeling was never information. It was the fluency of the interface, and letting it go is most of the skill.
Where this could be wrong
My own exposure to ChatGPT ads is close to zero: I work on paid tiers and in Claude, so my picture of what a Free or Go user actually sees is built from other people's screenshots and Similarweb's sampling, not from daily experience. OpenAI's answer-independence principle may hold completely, in which case the direct effect of ads on the organic answer is nil and the whole issue reduces to the older, pre-ad problem of a marketing-shaped corpus, which I think is the larger risk anyway. Sponsored Agents are a US-only test with select advertisers as I write this; the format may change or fail. The automation-bias literature comes mostly from aviation and medicine, and founders choosing software are neither pilots nor clinicians, so the effect sizes may not transfer. And the protocol is built for the founders who bring decisions to a coach, who are by selection the ones already uneasy; the operator who asks once, buys, and is fine does not show up in my sample.
What I am confident about is the shape of the mistake. The founder with three screenshots was not asking the wrong tools. She was asking them the wrong question, in the wrong order, and then treating the confidence of the answer as evidence about the vendor. The ads did not create that. They just made it visible enough that she noticed and asked. The fix is not a different assistant. It is four rooms, criteria before candidates, and one human on the phone before you sign.
Sources
OpenAI — Reimagining advertising with AI (16 September 2026): Sponsored Agents tested with select US advertisers; "distinct from ChatGPT's independent answers"; HubSpot and Shopify as first CRM and ecommerce partners
OpenAI — Our approach to advertising and expanding access to ChatGPT: ads principles; "Ads do not influence the answers ChatGPT gives you"; ads on Free and Go; Plus, Pro, Business and Enterprise ad-free; Go at $8/month
OpenAI — ChatGPT Ads expands to Southeast Asia and Taiwan (23 September 2026): over 60 countries; $1 billion annualised run rate in under 200 days; tens of thousands of advertisers
Similarweb — AI Ads press release (17 August 2026): 26% of ChatGPT responses to Free and Go users carry a sponsored ad; nearly 30% of ad-eligible Google AI Mode queries show ads
Fello AI — ChatGPT pricing 2026: Free, Go, Plus and Pro compared (updated 25 September 2026): US ads from February 2026; self-serve Ads Manager from 5 May; 31 European markets from 24 August
MIXED — ChatGPT Ads now run in over 60 countries, but only Free and Go users see them (26 September 2026): "statements of policy from the company selling the ads, not an audited finding"
Anthropic — Claude is a space to think: Claude to remain ad-free; no sponsored links adjacent to conversations, no third-party product placements
The Decision Lab — Automation Bias (24 November 2025): Mosier and Skitka, 59% task accuracy with a faulty automated aid vs 97% without
What is Wrong With Automation Bias? (Philosophy & Technology, April 2026): users follow AI advice even when it contradicts contextual information and their own assessment
How stale is your AI? Release age and training cutoff for 20 models (September 2026): GPT-6 Astra 3 September, Claude Fable 5.1 1 September
Daniel Kahneman, Thinking, Fast and Slow (2011) — WYSIATI: what you see is all there is
