In 2025 a research team published in PNAS the results of an experiment with nearly a thousand high school students in Turkey. They were split into groups: one had a GPT-4 based assistant while practising mathematics, one had a version with safeguards that offered hints instead of answers, and one had nothing.
While they still had the tool, the AI groups did markedly better: a 48% improvement with the plain assistant and 127% with the guarded one.
Then the tool was withdrawn and they sat an exam alone. The group that had used the unguarded assistant performed 17% worse than the students who had never had access.
Not worse than they had been with it. Worse than people who never had it.
The group with safeguards — hints, not solutions — largely escaped that damage. That detail matters more than the headline number, because it tells you what is actually doing the work.
The evidence is not one-sided, and the mixture has a pattern
Read only headlines and you get two opposite stories: AI is revolutionising education, or AI is destroying our ability to think. The evidence sits in between — but not in a lazy both sides way. It has a shape.
The meta-analyses so far are broadly positive: a synthesis of 35 experimental studies with over four thousand participants found a moderately positive effect on learning outcomes. Personalised AI tutoring does work reasonably well.
But one other study is more revealing. Fan and colleagues had 117 university students write an essay under four conditions: with ChatGPT, with a human expert, with writing analytics tools, or unaided. The ChatGPT group improved their essay scores most of all. Yet when the researchers measured knowledge gain and transfer to a new task, there was no significant difference between the groups.
Put the two findings together and you get a sharp summary: AI reliably improves the work, and does not reliably improve the worker.
Better essay. Correct exercises. No better writer.
Why this happens
The mechanism is no mystery. It is a set of long-established learning principles, and AI happens to press on exactly the vulnerable point of each.
It removes the step where you generate the answer. Roediger and Karpicke showed in 2006 that retrieving an answer strengthens memory far more than rereading it. This is among the most robust findings in the field. And the default mode of using AI is: you ask, it answers, you read. The single most valuable step is skipped.
It erases desirable difficulties. The Bjorks demonstrated that certain obstacles during study improve long-term retention: having to recall, having to wait, having to work it out. AI is designed to remove obstacles — including the useful ones.
And it amplifies the illusion of understanding. This is the dangerous part. Fisher, Goddu and Keil (2015) found that merely searching the internet inflates people's estimates of the knowledge inside their own heads — even on topics unrelated to what they searched. The line between I know this and I can reach this blurs.
An AI-written explanation is usually coherent, well structured and easy to read. And fluency is precisely the cue our self-assessment wrongly trusts. You finish reading, it feels clear, and clarity gets mistaken for comprehension.
Researchers have a name for the result: metacognitive laziness — learners handing over not just the task, but the monitoring and evaluation of their own learning.
Offloading memory is not new
To be fair: humans have offloaded memory for a very long time.
Daniel Wegner described transactive memory in the 1980s: in a family or a working group, each person holds part of the knowledge and everyone knows who to ask. You do not remember your sibling's phone number because your phone does. You do not remember the recipe because the book does.
In 2011 Sparrow, Liu and Wegner measured this for the internet: when people believe information will remain available, they remember the content less well — and remember where to find it better. Memory does not vanish; it shifts from content to index.
So what is new about AI? Two things. Previously you still had to read and synthesise what you found; now you do not. And previously search returned fragments you had to assemble; now it arrives finished, coherent and plausible — including when it is wrong.
So what is still worth remembering?
This is the real question, and the answer is neither everything nor nothing. Some kinds of knowledge cannot be offloaded for structural reasons, not moral ones.
First: anything that has to run in real time. A conversation moves in a few hundred milliseconds per turn. You cannot look up a word mid-sentence. No app helps when the other person has finished asking and is looking at you. For language learners this is the clearest answer of all: vocabulary, sound and grammatical reflex have to be in your head, because there is no time to fetch them from anywhere else.
Second: whatever you need to evaluate the answer. A language model can be confidently and fluently wrong. The only person who catches that is someone who already knows enough about the subject. Without background knowledge you cannot check the machine — you can only trust it, which turns learning into delegation.
Third: whatever you need in order to have the question at all. You cannot look up what you do not know exists. Knowledge in your head is what tells you which questions are worth asking. Someone who knows radicals looks at an unfamiliar character and asks what the left-hand element means; someone who does not sees only a block of strokes.
Fourth: whatever comprehension depends on. There is a classic study by Recht and Leslie (1988): weak readers who knew baseball understood and remembered a passage about baseball better than strong readers who did not. Comprehension is not a content-free general skill; it depends on what you already know. This applies directly to AI — you understand its explanation only as far as your background knowledge allows.
Fifth: whatever forms your schemas. In cognitive load theory, experts do not have wider working memory than novices; they have large pre-organised chunks in long-term memory, so each chunk occupies a single slot. That is the only way past the narrow channel. Offloading knowledge does not widen the channel — it just leaves it empty.
And what genuinely matters less now
Honesty runs the other way too. Some things no longer deserve room in your head.
Reference facts you use a handful of times in a lifetime and have leisure to look up. Exact figures, dates, tables. The detailed syntax of a tool you rarely touch. Long lists whose order does not matter.
What they have in common: used rarely, never urgent, and not the foundation for understanding anything else. Those are the three tests.
For language learners the line is unusually clear
Language learning may be the field where the answer is sharpest, because so much of its content is the kind that cannot be offloaded.
You can have AI translate a passage, but you cannot have it listen for you while someone is speaking to you. You can look up a character, but at three hundred characters a page, looking up each one is no longer reading. You can ask how a word is pronounced, but nobody can train your mouth on your behalf.
There is a real irony here worth naming: AI has made understanding a language nearly free, while using one costs exactly what it always did. The gap between those two is where the learning actually happens.
How to use it so it helps rather than replaces
The study at the top gives the key clue: the group that only received hints escaped the damage. What matters is not whether you use AI, but who is doing the heavy lifting.
Have it test you rather than teach you. Ask me ten questions on this and mark them has far more learning value than explain this to me. That is how you put the retrieval effect back in.
Produce first, submit second. Write your sentence, then ask what is wrong with it. Translate the passage, then compare. Guess the character from its radical, then look it up. This order preserves the generation step that the reverse order destroys.
Ask for hints, not answers. A prompt as simple as do not give me the answer, just point me to the next step recreates exactly the guardrail that protected those students.
Distrust the feeling of ease. An explanation that makes you nod immediately is not evidence that you understood it. The evidence is whether you can still reconstruct it tomorrow with the screen closed. Test that, often.
And use it for what it is genuinely excellent at: feedback. A tool that will always mark your sentence, flag your grammar and suggest a more natural phrasing is something only a private tutor could once provide. That is real value, and it lives in the direction where you produce and the machine comments — not the other way round.
What remains
The question what is left to remember when a machine can answer anything assumes that learning is the accumulation of answers. If that were true, it would indeed be obsolete.
But it never was. Knowledge in long-term memory is not a backup store for when the network drops. It is the material you think with. It sets how much of an explanation you can follow, how reliably you spot an error, which questions occur to you at all, and what you can say when someone is waiting for you to answer.
Machines are very good at answering questions. They do not produce the person who can ask them. That part still has to be built, one piece at a time, inside your own head.