Why MIT’s Brain Scans Should Change How You Grade
The cognitive damage isn’t coming from AI. It’s coming from what we still measure.
In 2025, a team at the MIT Media Lab put EEG sensors on the heads of college students and asked them to write essays. Some used ChatGPT. Some used Google. Some used nothing but their own minds. The researchers measured neural connectivity across the brain as the students worked.
The students who wrote without AI showed the strongest, most distributed neural activity. The ChatGPT group showed the weakest. After four months, the AI-dependent writers performed measurably worse on subsequent cognitive tests, and they could not even recognize sentences from essays they had themselves submitted weeks earlier. The researchers, led by Nataliya Kosmyna, called the effect cognitive debt.

The study has been read mostly as a warning about AI. Faculty have read it as evidence that AI use in classrooms must be limited, banned, or detected. CTL directors have circulated it as a justification for tighter policies. State systems have cited it in faculty handbook revisions.
That reading is wrong, or at least incomplete. The cognitive debt did not accumulate because the students used AI. It accumulated because they were graded on producing text.
Read the methodology again. The students were not running verification protocols. They were not auditing AI claims against primary sources. They were not running comparative outputs through multiple models to test for hallucinations. They were doing the one task at which AI has now become structurally superior to early-career writers: generating fluent prose on an assigned topic.
That task is the dominant unit of assessment across most of the humanities and large parts of the social sciences. The five-paragraph essay. The reflection paper. The literature review. The discussion board post. These were never wrong assignments before 2022. They were proxies. They measured something we actually cared about, the ability to think through an argument, marshal evidence, weigh counterclaims, by asking for the artifact that ability tends to produce.
The artifact is now uncoupled from the ability. A student can produce the artifact with no underlying ability at all, and the artifact will often look better than what most undergraduates can write on their own.
This is the actual implication of the Kosmyna findings, and it is not a policy implication. It is a measurement implication. If the unit of assessment is “generate text on a topic,” AI has already collapsed the validity of the measurement, and the students who comply with the assignment as written are the ones accumulating the cognitive debt. The AI policy on your syllabus is a downstream patch on an upstream problem.
The lever is the rubric, not the prohibition.
Consider a writing assignment under two different grading schemes.
Under the conventional rubric, the weights are roughly: thesis and argument 30 percent, evidence and analysis 30 percent, organization 20 percent, mechanics and citation 20 percent. Every category rewards the production of the text. A skilled AI user earns full marks. A first-generation student who wrote it herself, with all the rough edges that come from real thinking, earns a B-minus.
Under a verification-weighted rubric, the weights look more like this: verified sources with traceable URLs and direct quotes 25 percent, identified counter-evidence with reasoning for or against 20 percent, recomputed claims and manually checked logic 20 percent, cross-model comparison with documented discrepancies 15 percent, original synthesis the AI did not produce 20 percent.
The artifact looks similar on the surface. The student still hands in an essay. But the points are now distributed across cognitive operations that AI cannot perform on its own and that a student cannot fake without doing the work. A student who pastes ChatGPT output into the assignment scores near zero on four of the five categories, because the model does not know which of its own sources are fabricated, which of its arguments have credible opposition, or where its numbers came from. The student has to do that work, and the work itself is the learning.
This is what the guidebook calls the AI Audit. It is not a clever assignment add-on. It is a structural answer to the measurement problem the brain scans diagnosed.
The argument I am making is narrower than the one usually made about Kosmyna. I am not arguing students should not use AI. I am not arguing you need to detect or police it. I am arguing that the cognitive debt the study identified was a predictable consequence of grading an artifact that AI now produces better than the student. Change what you grade, and the cognitive debt stops accumulating, because the operations that earn the grade are the operations that exercise the brain.
This shift has a second effect that is harder to see at first. It changes which students benefit most. Under the generation rubric, the student with the best prose voice wins, and prose voice is heavily correlated with cultural capital and prior schooling. Under the verification rubric, the student willing to do unglamorous, slow, careful work wins, and that population looks very different. At a community college like Lansing, where most students are working adults, first-generation, or career-switchers, the verification rubric tends to reveal capability that the generation rubric was hiding.
The MIT result is not a story about AI. It is a story about what happens when an assessment instrument loses its validity. The cognitive debt was the cost of running an invalid instrument on a population. The students paid it.
If you are revising a syllabus for next term, this is the question to start with. Not “will I allow AI?” The question is: “What percentage of the points in this course currently reward an artifact that AI can now produce on its own?” If the answer is above forty percent, the brain scans are describing your course.
The fix is not philosophical. It is mechanical. Rebalance the rubric. Add the verification work. Put points on the operations that cannot be outsourced. The students who used to coast on prose voice will struggle. The students who used to be invisible will surface. And the cognitive activity in the room will start to look more like what the MIT scanners measured in the no-AI group than in the ChatGPT group.
The brain is a muscle. It builds when it is used. The grade is the lever that determines whether it gets used at all.
Dr. Szymon Machajewski is the author of The Learn-It-All Educator: A Guidebook for Training Brains, Not Replacing Them with AI. The free OER edition, Chapters 1 through 4, is available on Zenodo.com under a CC BY 4.0 license. The Complete Edition is available on Amazon. For institutional bulk pricing and faculty common-read inquiries, contact press@dataii.com.

