In June I set out to learn a domain I had been faking for about a year: how modern retrieval and evaluation systems actually behave under load, as opposed to how the papers say they behave. I did what any reasonable person with a Claude subscription does. I had the model explain the concepts, generate examples, walk me through the tradeoffs. It was a genuinely pleasant three weeks. I felt sharp. I could hold my end of a conversation about it.

Then in August a client asked me a direct question about eval design in a meeting where I had no laptop open, and I discovered something unpleasant: I could recognise every correct answer, and produce almost none of them. I had built a very convincing feeling of knowing, sitting on top of nothing durable.

So I went and read the research properly, expecting the usual mush of contradictory findings. Instead I found something clean enough to change how I learn. The literature does not disagree about whether AI helps you learn. It disagrees about where you stand when someone measures you.

The result that should worry you

Start with the most direct test anyone has run. André Barcaui at the Federal University of Rio de Janeiro randomly assigned 120 undergraduates to learn a topic either with ChatGPT as a study aid or with traditional, non-AI methods, then sprang a surprise test on them 45 days later. The AI group averaged 5.75 out of 10; the non-AI group averaged 6.85 — an 11-percentage-point gap, which in a normal exam is close to a full grade.

Two details matter more than the headline. First, learning with AI was genuinely faster — that part of the promise is real, and Barcaui does not dispute it. Second, the traditional group's scores clustered toward the high end while the AI group's spread out. The crutch did not lower everyone equally; it made outcomes erratic. Barcaui's own explanation is the one I would give: unrestricted use "impaired long-term retention, likely by reducing the cognitive effort that supports durable memory."

That is my June experience with a sample size. I had optimised, without deciding to, for the feeling of understanding rather than the fact of it. I have written before about whether AI book summary apps actually help you learn, and it is the same failure with better packaging. The pattern is showing up at population scale too, in the September 2026 data showing homework scores up and learning down.

The result that should encourage you

Now the opposite finding, from a study just as rigorous. Gregory Kestin and Kelly Miller at Harvard ran a crossover randomised trial with 194 introductory physics undergraduates, published in Scientific Reports on 3 June 2025. Each student did some material with in-class active learning — already the gold standard, already better than lecturing — and some with a custom AI tutor called PS2 Pal.

The AI tutor won, and not narrowly. Median post-test scores were 4.5 with the AI tutor versus 3.5 with in-class active learning, with learning gains more than double relative to baseline, and students got there in a median of 49 minutes against a 60-minute class block. Engagement and motivation ratings were higher too.

But read how PS2 Pal was built, because that is the actual finding. It was instructed to be concise, to give step-by-step guidance without handing over full solutions, to force active thinking, to manage cognitive load, and to push a growth mindset — and it was supplied with correct solutions in advance so it would not hallucinate the physics. This was not "a chatbot." It was a pedagogy, implemented in a prompt.

The variable nobody names: distance

Here is where the two studies stop contradicting each other. The Harvard post-test was administered essentially at the end of the session. Barcaui's test came 45 days later with no tool in the room. Both measured real things. They measured them at different distances from the machine.

The study that nails this is the one people quote most and read least: "Generative AI Can Harm Learning" by Hamsa Bastani, Osbert Bastani and Alp Sungu, a field experiment with nearly 1,000 high-school students in Turkey, published in PNAS. Students got either a plain GPT-4 assistant (GPT Base), a GPT-4 tutor with pedagogical safeguards (GPT Tutor), or nothing.

On the practice problems — done with the tool in hand — GPT Base students performed 48% better than control and GPT Tutor students 127% better. Those are the numbers that get screenshotted. Then came the exam, taken without any AI, and the GPT Base cohort performed 17% worse than the control group. Not 17% less improved. Worse than students who never had the tool at all. The tutor safeguards mitigated that damage — GPT Tutor students did not end up behind — but note precisely what that means: the well-designed tutor's main achievement was avoiding harm, not multiplying learning.

So the honest summary of the field, as of today, is this. Close to the tool, AI makes you look dramatically more capable. Far from the tool, plain AI use makes you measurably less capable, and careful AI design gets you back to roughly where you would have been. The 127% was never learning. It was the tool's performance wearing your name.

The distance test

This gave me a rule I now apply to any learning setup, mine or a client's: ask at what distance from the machine your competence has been verified. Zero distance is a feeling. Real distance is evidence. Here is how the common setups score when you grade them that way rather than by how good they feel.

SetupDo you produce first?Is the answer withheld?Verified at distance?Verdict
"Explain X to me" in a normal chatNoNoNoFluency theatre — the Barcaui failure mode
Reading an AI summary of a book or paperNoNoNoFastest way to feel informed and stay ignorant
Deep-research report you read end to endNoNoNoExcellent for deciding, near-useless for learning
Study Mode / Learning Mode / Guided LearningYesMostlyNoNecessary, not sufficient — the Bastani floor
Socratic mode, then notes written from memoryYesMostlyWeaklyReal learning starts here
Socratic mode + unassisted retest 2-6 weeks laterYesYesYesThe only setup the 45-day evidence supports
Explaining it to a human who can push backYesYesYesStill the strongest test available

The third row deserves a defence, because I rely on it daily. A deep-research report is a superb instrument — for deciding. Reading one is how I form a view on a market in an afternoon, and pressure-testing that view is the subject of getting AI to challenge your decisions instead of agreeing with you. It simply is not learning, and the mistake is filing it as though it were.

Notice that the top three rows are what almost everyone actually does, and they share one property: at no point are you the one generating the answer. Every row that works is a row where you produce before the machine does. That is not an AI insight, it is the testing effect — the finding, replicated for over a century and popularised in Make It Stick, that retrieving information strengthens memory far more than re-reading it. AI did not repeal it. AI made it trivially easy to skip.

The study modes are real — and they are not the fix

The vendors have, to their credit, shipped exactly what the research recommends. ChatGPT's Study Mode is available even on the free tier and toggles from the tools menu inside a normal chat. Claude has Learning Mode, which works inside Projects, so it can tutor you Socratically against your own uploaded documents rather than generic knowledge — for a founder learning their own market or codebase, that is the most useful of the three. Gemini has Guided Learning. All three do the same core thing: ask guiding questions and withhold the finished answer. Daniel Nest's hands-on comparison of all three found the differences are real but modest; the category works as advertised.

Then comes the part that decides everything, and it is not technical. In June 2026 Stanford researchers reported on two school districts using an AI tutoring platform: even with scheduled time set aside, only about 61% and 53% of students used it at all, and average weekly use was 2.18 and 5.23 minutes against the 30 minutes a week the provider says is needed for measurable gains. Adding a human tutor alongside raised engagement by one minute and 4.4 minutes a week respectively. Stanford's companion brief, AI Tutoring is Not a Monolith, is blunt: the strongest evidence sits with AI tools built to support human tutors, not to replace them, and fully AI-led tutoring "does not yet meet the established evidence base." When students were left to work independently, 40-47% never used the platform at all.

Founders are not schoolchildren, but we fail the same way, faster. Socratic mode is slower and mildly humiliating — it makes you sit in not-knowing, which is precisely the discomfort that does the work. The escape hatch is one click away and unmonitored. I abandoned Learning Mode twice in July for exactly this reason, both times while telling myself I was short on time. That is the same muscle I described in breaking AI dependency without giving up AI at work, and it is why choosing an AI tutor as a busy professional is mostly a question about your own compliance, not the product.

What I actually run now

Four steps, roughly forty minutes a week, no new software. I rebuilt my eval-systems attempt on it in August and the difference was not subtle.

One: state the question, then answer it badly, before opening anything. Write your current best understanding in three or four sentences from memory. It will be embarrassing. That is the pre-test, and the gap it exposes is what makes the next twenty minutes stick — the same mechanism behind the pre-test effect, where guessing wrong before instruction improves later recall.

Two: use Socratic mode for the session, not chat mode. Claude Learning Mode inside a Project holding your real documents, ChatGPT Study Mode, or Gemini Guided Learning. If you catch yourself asking it to "just give me the answer," you have left the study session and entered a different activity — a legitimate one, but not this one.

Three: close the tab and write the explanation from memory. Not notes taken during the session, which are transcription. A blank-page explanation afterwards, in your own words, including what you could not reconstruct. This is the step everyone skips and the step every piece of evidence points at.

Four: schedule the unassisted retest at three weeks. One calendar entry: "explain X, no tools, 10 minutes." Three weeks is far enough that the Barcaui decay would have shown up. If you cannot reproduce it, you did not learn it, and you now know that cheaply instead of finding out in front of a client. I keep these in the same growth log described in this protocol for measuring personal growth with AI memory.

The whole design is one idea: move the measurement away from the machine. Everything else follows.

Where this could be wrong

Three honest limits. Barcaui's trial is 120 undergraduates learning one topic over 45 days — a real RCT, but small, and undergraduates studying for a presentation are not founders learning a market under commercial pressure, where motivation is much higher. The Harvard result was measured immediately, so we do not know what those physics students retained at 45 days; nobody has run the long-horizon version of that experiment, and it is the single most useful study someone could run right now. And the Stanford engagement data comes from children in grades 1-5, whose compliance problems are not identical to yours, though I would argue mine differ mainly in the sophistication of my excuses.

There is also a case for not learning some things. Ethan Mollick's argument in Co-Intelligence is that we are all still discovering where the human should sit in the loop, and for genuinely peripheral domains, staying dependent is a perfectly rational trade. I do not need to internalise tax law. But I do need to internalise anything I will be questioned on without a laptop, anything I have to make judgement calls about under time pressure, and anything I intend to have an opinion about in public. For that list, the 45-day test is the only score that counts. Deciding which list a topic belongs on is its own skill — I covered the founder version of it in how founders actually learn new skills with AI, and the judgment-preservation side in breaking free from AI dependency on judgment.

The uncomfortable summary: the tools got good enough that the feeling of learning became free, while learning itself stayed exactly as expensive as it was in 1885 when Ebbinghaus first plotted the forgetting curve. AI will happily sell you the feeling all day. The only defence is to keep measuring yourself somewhere the machine cannot reach.

Sources

André Barcaui — ChatGPT as a cognitive crutch: evidence from a randomized controlled trial on knowledge retention, Social Sciences & Humanities Open (n = 120; 5.75 vs 6.85 out of 10 on a surprise test 45 days later)
ScienceAlert (2 April 2026) — coverage of the Barcaui trial, including the forgetting-curve chart and score distributions
Kestin, Miller et al. — AI tutoring outperforms in-class active learning, Scientific Reports, 3 June 2025 (194 students; median 4.5 vs 3.5; 49 minutes vs a 60-minute block)
Bastani, Bastani & Sungu — Generative AI Can Harm Learning, PNAS (~1,000 Turkish high-school students; +48% and +127% on assisted practice, −17% for GPT Base on the unassisted exam)
Knowledge at Wharton — Without guardrails, generative AI can harm education
Stanford SCALE / National Student Support Accelerator — AI Tutoring is Not a Monolith (2026): the AI-human relational-intensity spectrum; 40-47% non-use when unsupervised
K-12 Dive (18 June 2026) — AI tutor access alone doesn't equate to student gains: 2.18 and 5.23 minutes of average weekly use against a 30-minute threshold
Education Week (August 2026) — When does AI help most with tutoring?
Mary Burns, Brookings (27 January 2026) — What the research shows about generative AI in tutoring
Daniel Nest, Why Try AI — I tested three different AI "study" modes (ChatGPT Study Mode, Claude Learning Mode, Gemini Guided Learning)
Peter Brown, Henry Roediger & Mark McDaniel, Make It Stick (2014) — the testing effect and the pre-test effect
Ethan Mollick, Co-Intelligence (2024) — where the human belongs in the loop

Share this post