Yes, because the decision to search runs on the same stale weights. In a September 2026 harness of 16 models and 2,000 calls, GPT-6 Astra made the right search decision 80 of 80 times and Claude Fable 5.1 78 of 80, while GPT-5.6 Luna, told to search, still named a king who died August 28. Ed Miller’s Logic of Sports Betting names the risk: a stale line. Only 10 of 20 current models even publish a cutoff.
A founder asked me last week whether it still matters when a model stopped reading, now that every serious assistant has a search button. I said no, reflexively, and then spent an evening finding out I was half wrong. The half that is wrong is the half that costs money, because founders use these tools for exactly the facts that go stale: competitor pricing, who runs what, which regulation passed, which model is current.
The data comes from two pages by the same engineer, both on the Hacker News front page this month. The first, stale.jock.pl, checked September 16, 2026, lists release date and training cutoff for 20 current models from 8 labs, and only 10 of the 20 have a cutoff the lab actually publishes. GPT-6 Astra shipped September 3, 2026 with an April 30, 2026 cutoff; Claude Fable 5.1 shipped September 1 with a June 2026 cutoff; Claude Sonnet 5, released June 30, stopped reading in January 2026; Gemini 3.1 Pro, released February 19, 2026, carries a January 2025 cutoff. Gemini 3.8 Flash, Muse Spark 1.3 and DeepSeek V4.1-Flash, all September releases, publish no cutoff at all. The second page is the experiment: 16 of those models, each handed a single web_search tool, 40 questions, four system prompts, 2,000 calls at temperature zero for $6.53. Twenty questions were about events after every model’s cutoff, including the death of King Harald V of Norway on August 28, 2026; twenty were evergreen controls like why the sky is blue. The only thing measured was whether the model decided to search.
The results split the shelf. GPT-6 Astra made the correct decision 80 of 80 times across all four prompt conditions, GPT-5.6 Sol went 80 of 80 in three of four, Claude Opus 5 79 of 80 and Claude Fable 5.1 78 of 80, with zero false searches between them. Five of the sixteen told the author Harald V is the current king, and GPT-5.6 Luna did so under a system prompt that literally said its data ends in February and to call web_search for anything later. Claude Sonnet 5 with no system prompt gave three confident stale answers and zero searches, then went to zero misses once the one-line rule was added. At the other end, Meta’s Muse Glimmer searched the web for 18 of 20 evergreen questions, including the boiling point of water, and Grok 4.6 searched to check whether 1,000,003 is prime. Fable’s single miss was the most instructive: it confabulated a tennis result from real names rearranged, in clean prose, without checking.
Ed Miller and Matthew Davidow’s The Logic of Sports Betting is the lens I use, because they describe this exact market. Market-making books move their line the moment information arrives; retail books copy the line and are slow to update, and the bettor who spots the stale price at a retail book, a tactic Miller calls steam chasing, wins until they are banned. A model answering from weights is a retail book quoting last season’s line, and the founder acting on it is on the wrong side of that trade. Miller’s other rule is that the hold destroys your margin for error: one bad bet in five wipes out four good ones. One stale competitor price in five research sessions does the same to a strategy memo.
| Model (release / cutoff) | Moves the line: searched on post-cutoff facts (/5) | Low hold: no wasted searches on settled facts (/5) | Publishes its line: cutoff disclosed (/5) | Stale-line risk score /15 |
|---|---|---|---|---|
| GPT-6 Astra (Sep 3, 2026 / Apr 30, 2026) | 5 | 5 | 5 | 15 — research grade |
| Claude Fable 5.1 (Sep 1, 2026 / Jun 2026) | 4 — one fluent confabulation | 5 | 5 | 14 |
| Claude Opus 5 (Jul 24, 2026 / May 2026) | 4 — announced a search it never made | 5 | 5 | 14 |
| Claude Sonnet 5 (Jun 30, 2026 / Jan 2026) | 2 bare, 5 with the rule | 5 | 5 | 12 with a system prompt, 9 without |
| GPT-5.6 Luna (Jul 9, 2026 / Feb 16, 2026) | 1 — kept the dead king even when instructed | 4 | 5 | 10 |
| Muse Glimmer (Aug 10, 2026 / not published) | 4 | 1 — searched 18 of 20 evergreen questions | 0 | 5 |
The scores are mine, built from the harness author’s published counts and Miller’s three questions, and they carry his caveats. The evergreen dimension is a real cost, not a nicety: a model that distrusts its own weights pays in latency and tokens on every question a first-year student answers cold, and at Astra’s $10 per million input tokens the discipline is something you are paying for. The disclosure column matters more than it looks. Half the shelf will not tell you when it stopped reading, so you cannot write the always-search rule for them, and a book that hides where its line came from is a book Miller would not bet at.
The limits. This is one engineer’s harness, 40 questions, run once, with the tool never executed and the answer never graded, so it measures the decision to search and nothing else. Two misses may be harness artifacts from a 400-token ceiling. Product chat interfaces wrap these models in their own search logic, so the API behavior above is a floor, not a promise about the ChatGPT or Claude app you actually use. And the frontier result cuts against my own instinct: on the expensive models the decision to search is effectively trained in, and the cutoff matters less than it did a year ago. It matters most on the mid-tier models people route their bulk work to for cost.
What I do now is what the harness author does, with one addition. Date and published cutoff go in every system prompt, which is free. An always-search list, no judgment allowed, covers prices, versions, titles and anything with a name and a date in it. Research work goes to the model that over-searches; fixed offline work goes to the one that trusts itself, and I check the output. The addition is Miller’s: before I act on any fact a model gave me, I ask which book quoted the line and when it last moved. If the model cannot tell me its cutoff, the answer is a stale price until proven otherwise.
Sources: stale.jock.pl — AI training cutoff and release dates for 20 models (data checked Sept 16, 2026), Does AI training cutoff matter with web search? 16 models, 2,000 calls (September 2026), OpenRouter — GPT-6 Astra model page and pricing.
Related: How to Find Your Passion · Best Self-Improvement Books · How to Make Better Decisions · AI Coach App — Building It in 8 Hours
