Bizarus.
HomeBlog › Cognitive debt is a real finding. It is not the finding the headlines took from it.
Argument

Cognitive debt is a real finding. It is not the finding the headlines took from it.

Three 2025 studies are being read as proof that AI erodes thinking. Read carefully, they show something narrower and more useful: people stop doing whatever step a tool removes, and whether that matters depends entirely on whether that step was where the thinking was. The finding is about offloading, not intelligence, and it argues for tools that keep the load-bearing reasoning with the human rather than for abstinence.

18 August 2026 Bizarus

In a lab at MIT, 54 people wrote essays while wearing EEG caps. One group used ChatGPT, one used a search engine, one used nothing but their own heads. The preprint that came out of it reported that the ChatGPT group showed the weakest and least distributed brain connectivity of the three, reported the lowest sense of owning what they had written, and, most tellingly, often could not quote a sentence from an essay they had finished minutes earlier. The authors called what accumulated "cognitive debt."

The study traveled fast, and it traveled as a verdict: this is your brain on AI, and the brain is losing. That reading is understandable and it is wrong, or at least far more confident than anything the evidence supports. The gap between what these studies measured and what they are taken to mean is worth closing carefully, because the correct reading points somewhere more useful than either alarm or reassurance.

Three studies, one mechanism

The MIT essay work is not alone. In a separate study of 666 people, Michael Gerlich found a negative correlation between frequent AI use and performance on critical thinking measures, with the effect running through cognitive offloading and hitting younger users hardest. And in a survey of 319 knowledge workers, researchers at Microsoft and Carnegie Mellon found that the more people trusted an AI tool's output, the less critical thinking they reported doing about it, while people with higher confidence in their own judgment did more.

Three different methods, three different populations, one recurring word: offloading. When a tool will do a step, people let it. That is the finding these studies actually share, and it is narrower than the headline.

Offloading is not the same as decline

The slide happens in a single move. "People stopped doing a step the tool now does" becomes "people got worse at thinking." Those are not the same claim, and the difference is the whole argument.

We have been here before. In 2011 Betsy Sparrow and colleagues described the Google effect: when people expect to be able to look something up later, they remember the fact less well and remember where to find it better. That was framed at the time as a loss too. Mostly it was not. It was a reallocation. We stopped memorizing what we could reliably retrieve and kept the index instead, which is what a working memory system is supposed to do. Long division went to calculators and the people who needed real mathematical fluency built it anyway. Offloading is not a modern accident. It is one of the main ways human cognition has ever scaled, and every wave of it has arrived with the same worry attached.

So the strongest version of the skeptical case is not that these studies are wrong. It is that they are measuring a very old and mostly benign thing, and dressing it in EEG traces does not make it new.

The finding everyone skipped

The answer to that case is sitting inside the Microsoft and Carnegie Mellon paper, in its least quoted result. Critical thinking did not disappear from the workers who used AI. It moved. The effort shifted from gathering information to verifying it, from solving the problem to integrating the machine's answer, from doing the task to supervising it.

That single observation does two things at once. It defeats the "makes you stupid" reading, because thinking that relocates has not been destroyed. And it puts a hard condition on the optimistic reading, because a relocation is only safe if the person actually does the work at the new location. The same study found that confidence in the tool predicts they will not. The verification step is exactly the one people skip when they trust the output, which means the safety of offloading depends on the one behavior the tool most reliably erodes.

The variable is not the tool

Put those together and the useful question is not whether AI helps or harms thinking. It is whether the step being offloaded was load-bearing.

The essay experiment is the clean case. Constructing the argument was not a means to writing the essay; it was the point of writing the essay. Offloading it offloads the entire objective, which is why the students could not quote themselves: there was nothing to recall because the thinking never happened in their heads. Compare that to offloading a citation lookup you would have done mechanically anyway, where nothing of value is lost. Most real tasks sit between those poles, and the distinction genuinely blurs in the middle. But it has teeth at the edges, and the edges are where the damage is.

This is the honest limit of the evidence, and it should be stated plainly rather than buried. These studies are largely correlational and lean on self-report. The MIT work is a small, non-peer-reviewed preprint that has already drawn methodological criticism; the Gerlich paper has since carried a published correction. None of them establishes that AI use causes durable cognitive decline, and anyone citing them as if they had is overreaching in the opposite direction. The most defensible claim they jointly support is modest: when a tool removes an effortful step, people skip the effort, and effort that is skipped is not encoded.

What that asks of a tool

Modest is not the same as unimportant, because the claim has a direct design consequence. If the harm comes from frictionlessly removing the load-bearing step, then a tool meant to help people think should do close to the opposite of what most are optimized for. It should keep the load-bearing step visible and slightly effortful rather than smoothing it away. It should surface what needs checking instead of hiding it inside a fluent answer, since the fluent answer is precisely what suppresses the verifying the whole arrangement depends on.

This is the line Bizarus tries to hold, and it is worth being honest that it is a design commitment rather than a solved problem: the measure of a research tool is not how much thinking it removes but where it leaves the thinking that remains. A tool that quietly does the reasoning and hands over a confident conclusion is not saving the researcher effort. It is relocating the effort to a verification step the researcher has been trained, by the tool's own confidence, to skip.

The MIT participants who could not quote their own essays had not lost the ability to think. They had simply never done the thinking, because nothing in front of them asked for it. That is not a verdict on human cognition. It is a fact about what the tool was built to want, and unlike a verdict, it can be built differently.

AI & ResearchEvidence & Evaluationcognitive offloadingcritical thinkingcognitive debttransactive memorytool design
← All Blog
© 2026 Bizarus AI