When Thibault Schrepel designed his "Law of AI" course at Vrije Universiteit Amsterdam, he carried a common assumption into the classroom: that handing students an AI chatbot without any instruction would likely cause more harm than good. Two years of controlled testing pushed him toward the opposite conclusion. In his own words, "I was wrong."
The experiment he ran is notable less for its headline than for its structure. Rather than surveying students about their attitudes toward AI, Schrepel split them into groups with different rules and then measured what they actually produced. The same setup ran twice — first in 2024 with 66 students, then again in 2025 with 164 participants — which let him observe not just whether AI helped, but whether the way it was introduced mattered over time.
How the Three-Group Experiment Worked
Every participant faced the same core task: working in small teams of four or five, they had 20 minutes to improve a provision of the EU AI Act. Submissions were graded on substance, clarity, proportionality and innovation — criteria that reward genuine legal reasoning rather than surface-level polish. All students also sat the same assessments: a multiple-choice exam and a take-home exam that asked them to revise a separate provision of the regulation.
The only variable was how they were allowed to work:
- No AI: This group was barred from using ChatGPT entirely and had to rely on discussion and its own legal knowledge.
- Unguided AI: These teams received ChatGPT-generated revision suggestions embedded directly into the text. They could keep using the tool, but nobody explained how to prompt it or how to evaluate what it returned.
- Trained AI: This group received hands-on instruction in legal prompt engineering, along with explicit training in checking AI output for consistency and accuracy.
The No-AI Group Hit a Wall Called "Idea Exhaustion"
The students working without AI finished last in both years of the study. Their edits skewed superficial: rephrasing sentences for clarity, trimming redundancy, tightening wording. Only a small number of subgroups attempted substantive legal improvements to the provision. More striking was the timing. After roughly ten to fifteen minutes, many teams ran out of things to try — a pattern Schrepel labels "idea exhaustion."
The ban did produce one benefit that the researcher acknowledges. Without a chatbot to lean on, students argued with each other more, hashing out the text through genuine discussion. The trade-off, however, was a ceiling on how far that discussion could travel within the time limit.
Unguided AI Use: Fluent Output, Little Scrutiny
If the no-AI group suffered from too few ideas, the unguided group suffered from too little skepticism. Students in the middle condition largely accepted what ChatGPT proposed, frequently justifying a change on the grounds that the AI's version simply "sounded better." In some cases, participants swapped out terms such as "shall" and "individual" for AI-suggested alternatives without demonstrating any grasp of the legal consequences of doing so — a substitution that carries real weight in statutory drafting.
Every single subgroup in this condition retained at least one misleading or legally extraneous term from ChatGPT's output. The problem was not that the tool produced useless text; it is that the text arrived polished enough to discourage questioning.
Structured Training Produced Real Dialogue With the Tool
The third group behaved differently. These students pushed back on the model, testing alternate phrasings and interrogating the substantive legal questions underneath the assignment rather than accepting the first plausible draft. Their scores reflected it: in 2024, the trained group outperformed the other two by a wide margin, with the advantage showing up most clearly on the more demanding take-home exam.
The Training Edge Nearly Disappeared in Year Two
Then the pattern shifted. When the experiment was repeated in 2025, the gap had almost closed. All three groups landed at roughly the same level of performance, including the students who had no access to AI at all.
Schrepel's explanation is straightforward: familiarity. As chatbots became a normal part of daily life, students arrived in the classroom already knowing how to work with them. Formal instruction still had value, but it no longer conferred the same advantage it had a year earlier, when the tool was less embedded in routine academic work.
That finding complicates any simple policy reading. It suggests the benefit of training is partly a function of baseline skill — valuable when a tool is novel, less differentiating once it becomes ambient.
What Educators Should Take From the Results
Several practical observations emerge from the two-year run:
- Removing AI from a time-limited exercise did not raise the quality of student work; in this course, it lowered it.
- Allowing AI without instruction left students vulnerable to confident-sounding errors they lacked the expertise to catch.
- Teaching students to interrogate AI output — checking consistency, verifying accuracy, prompting deliberately — produced measurably better results, at least while the skill was uncommon.
- The knowledge gap around AI literacy is narrowing on its own as students gain everyday experience with these tools.
The broader lesson is that the meaningful distinction may not be AI versus no AI, but guided versus unguided use. The trained group's advantage came from treating the chatbot as something to be challenged rather than obeyed, and that habit, not the tool itself, appears to have driven the improvement in outcomes.
It is worth keeping the study's scope in perspective. The findings come from a single course, a single instructor and one discipline, with relatively modest cohort sizes. Legal drafting is also a domain where precision matters in ways that may not map cleanly onto other subjects.
Still, the two-year arc gives the results unusual weight. Schrepel set out expecting unguided AI to backfire, and the data refused to cooperate with that expectation. His conclusion is a useful corrective for institutions weighing outright bans: prohibition removed a source of ideas without eliminating the need for judgment, while instruction in how to use the tool responsibly delivered the biggest measurable gain.
This article is based on reporting by The Decoder. Read the original article.
Originally published on the-decoder.com








