For a decision pressure-test, Meta's Muse Spark 1.3 (launched September 2, 2026, about $0.10 per million tokens blended) scores highest on Sir John Whitmore's coaching criteria because it is trained to ask clarifying questions and confirm before acting, while Claude Fable 5.1 ($10/$50 per million, 1M-token context) wins when the whole data room must be in view. GPT-6 Astra and Gemini 3.8 Flash are execution engines first.
The first ten days of September 2026 were the densest stretch of frontier releases I can remember: Anthropic shipped Claude Fable 5.1 on September 1, OpenAI released GPT-6 Astra on September 3, Google DeepMind and Meta both shipped on September 2 with Gemini 3.8 Flash and Muse Spark 1.3, and DeepSeek closed the window with V4.1-Flash on September 10. Every founder I coach asked some version of "which one do I switch to?" and every one of them was asking the wrong question, because they were not choosing a model for coding. They were choosing a thinking partner for the two or three decisions a quarter that actually matter: the hire, the price change, the round.
Sir John Whitmore's Coaching for Performance gives me a cleaner set of criteria for that job than any benchmark. Whitmore's equation is Performance equals Potential minus Interference, and a good coach reduces interference by raising awareness and responsibility rather than by telling you what to do, because "telling creates dependence". Translate that into model behaviour and you get four things to score: does it ask before it tells, does it hold the whole picture, does it leave the decision with you instead of acting, and how much noise does a session cost you in tokens, time and money. Here is how the four models a founder can reach this month come out, scored one to five from the labs' own launch documentation and from my own sessions where I have access.
| Model (launch, list price) | Asks before telling | Holds whole picture | Leaves decision with you | Interference (noise, cost of a 150k-token review) | Total /20 |
|---|---|---|---|---|---|
| Muse Spark 1.3 (Sep 2; ~$0.10/M blended) | 5 — trained to ask clarifying questions when prompts are ambiguous | 3 | 5 — confirms before consequential actions | 4 — ~25% fewer tokens than 1.2; about $0.02 a session | 17 |
| Claude Fable 5.1 (Sep 1; $10 in / $50 out per M) | 3 — answers unless instructed to interview you | 5 — 1M-token context, 128K output | 3 | 2 — ~1.7x the output tokens of Fable 5; about $2.00 first pass, $0.54 on cached re-reads | 13 |
| GPT-6 Astra (Sep 3; plan-gated, API price not in the launch post) | 3 — built for delegation and intent-following | 3 | 4 — 0% scope-exceeding on OpenAI's Hugging Face-incident eval vs 48% for GPT-5.6 Sol | 4 — 47% less time per task on OSWorld 2.0 | 14 |
| Gemini 3.8 Flash (Sep 2; $0.75 in / $3.75 out per M until Dec 31) | 2 — designed to "work harder", not ask | 3 | 2 — calls tools iteratively by default | 4 — about $0.15 a session, doubling to $0.30 from January 1, 2027 | 11 |
The scores surprised me. Muse Spark 1.3 is the cheapest model in the current top five, and Meta's launch post says almost nothing about raw intelligence; its headline claim is behavioural, that the model "asks clarifying questions when prompts are ambiguous, invokes help from the user when stuck, and confirms before taking consequential actions". That is a description of a coach, not an oracle. In a decision review, being interviewed is the product. When I gave it a pricing memo and asked it to help me think, it asked which customer segment I was most afraid of losing before it offered a view, which is exactly the Whitmore move of raising awareness first.
Claude Fable 5.1 wins a different job. Its 1M-token context and 128K-token output mean the whole data room, the board deck, last quarter's transcripts and the cap table sit in one window, and Anthropic's 75% cut to cache-read pricing, from $1.00 to $0.25 per million tokens, makes the second and third passes over that material cheap. The price is verbosity: independent measurement found Fable 5.1 emitting about 1.7 times the output tokens of Fable 5 at maximum effort, which in coaching terms is interference. I now run it at medium effort with an explicit instruction to ask three questions before writing anything, and the sessions got shorter and sharper.
GPT-6 Astra and Gemini 3.8 Flash are superb, and they are the wrong tool for this. OpenAI's own announcement pitches Astra on computer use and professional deliverables, scoring 72.6% on OSWorld 2.0 in about 40 minutes per task versus 65.7% in 75 minutes for GPT-5.6 Sol, and its alignment claim, going beyond authorized scope in 0% of cases versus 48% for Sol, is genuinely reassuring for delegation. But delegation is the opposite of a decision review. Google describes Gemini 3.8 Flash as a model that "works harder", executing extra reasoning steps and calling tools iteratively, which is what you want when the decision is made and the work begins. Astra's API pricing was not in the launch post I read, and its safety numbers have no third-party replication until at least OpenAI's DevDay on September 29, so treat its row as provisional.
Two failure modes apply to every row. The first is the one Whitmore warned managers about: the moment a model tells you the answer, you have outsourced responsibility, and the decision quality drops even if the answer is good. The second is that none of these models knows your business, and Whitmore's point that "the coach doesn't need expertise in the subject" cuts both ways; the value of the session is entirely in the quality of the questions, which means the prompt is the coaching skill, not the model. If you take one thing from the table, take this: pick Muse Spark 1.3 or Fable 5.1 for the review, write "interview me before you advise me" as the first line, and keep Astra and Gemini for the work that follows the decision.
Sources: Meta AI Research, "Introducing Muse Spark 1.3" (September 2, 2026); OpenAI, "GPT-6 Astra: A new generation of intelligence"; Google, "Introducing Gemini 3.8 Flash and 3.8 Flash Cyber" (September 2, 2026); Implicator, "Anthropic Fable 5.1 keeps $10/$50 price, cuts cache reads 75%"; Local AI Zone, "September 2026 AI Model Updates" (updated September 11, 2026).
Related: How to Find Your Passion · Best Self-Improvement Books · How to Make Better Decisions · Why Exploration Is Important for Success
