ChatGPT and Critical Thinking: What OpenAI's 1,053-Student Bocconi Study Actually Found

A preregistered randomized controlled trial on 1,053 Bocconi students measures ChatGPT against causal-reasoning training for the first time — the results surprise in both directions.

ChatGPT and Critical Thinking: What OpenAI's 1,053-Student Bocconi Study Actually Found
Table of contents

On August 27, 2026, OpenAI published some of the most consequential education research of the year: a preregistered randomized controlled trial conducted with Bocconi University in Milan that finally puts hard numbers on the question every student, teacher, and parent keeps asking. Does using ChatGPT on real coursework erode critical thinking — or raise the quality of the work? The answer, laid out in a 1,053-student experiment and a public research paper titled "Training novices to think, or giving them LLMs? Evidence from an RCT," is genuinely surprising in both directions. ChatGPT lifted students' grades by nearly a full point on a five-point scale. Critical-thinking training didn't lift grades at all — yet it changed how students think in ways the graders' rubric couldn't capture. And the study's sharpest finding isn't about AI at all: it's about what our grading systems reward. This article walks through the experiment in detail, what the numbers do and don't prove, the counter-evidence from other large studies, and what it all means for how you should use AI in your own studies. (Last updated: August 31, 2026)

What exactly happened: an experiment run inside real classrooms

This is not a survey, and it is not an opinion essay. The design is a preregistered 2×2 factorial randomized controlled trial — the researchers locked their analysis plan before collecting data, which rules out after-the-fact cherry-picking. It ran in November 2025 across 13 sections of the same first-year management course at Bocconi University, one of Europe's leading business schools, covering 1,053 first-year undergraduates across three bachelor's programs (economics & management, economics & finance, and management). Randomization happened at the classroom level within each program, producing four groups:

  • Control (249 students): completed the task with no assistance.
  • Causal-reasoning training (256 students): received a short training module on causal reasoning — the skill of linking causes to effects through explicit mechanisms and identifying the conditions under which an idea might fail. The training was a game with examples, questions, and feedback, and it never mentioned AI.
  • ChatGPT access (197 students): got access to ChatGPT Edu running GPT-4o during the task.
  • Both combined (351 students): deliberately oversampled by the researchers to measure the interaction between the two with extra statistical power.

The task itself was deliberately real: 45 minutes, during class time, with no advance announcement (so participation simply reflected attendance). Students wrote consultant-style recommendations of at most 180 words on how to increase alumni awareness and usage of the university's merchandise shop — two established criteria in the marketing literature. Solutions were graded by twenty trained master's students, three raters per submission, and separately compared against the recommendations of three domain experts. Automated text analysis measured the number and variety of ideas, signs of causal reasoning, and similarity to expert solutions. Compliance was high: 90.5% of students in the GPT groups reported actually using it.

Experimental design: four randomized groups and one real marketing task


Source: Official research paper (OpenAI PDF)

Finding one: ChatGPT raised evaluated performance by nearly a full point

Students with ChatGPT access scored almost a full point higher on the five-point evaluation scale — a large effect by education-research standards. Their submissions contained more ideas (an increase of more than half a standard deviation in the number of core ideas — roughly two additional ideas on average), followed clearer logic, and came closer to the recommendations written by the three marketing experts. In the study's own framing, AI helped novices produce work that looked more professional.

Crucially, the researchers decomposed this effect. After controlling for textual features — coherence, idea count, idea diversity, and other text properties — about half of the ChatGPT advantage survived. That means the gain was not cosmetic polish alone: the substance of the recommendations genuinely improved. OpenAI's post stresses a nuance that matters for the cheating debate: students were not simply handing the assignment over. They still had to decide what to ask, evaluate the responses, and choose what went into the final submission. Those are themselves skills — and if you are comparing AI tools for study, our earlier breakdown of which assistant fits students best covers the practical differences.

Estimated effects on evaluation scores and similarity to experts


Source: Official research paper (OpenAI PDF)

Finding two: critical-thinking training changed the mind, not the grade

Here is the study's second surprise. Students who received the causal-reasoning training did not score higher — if anything, their grades came in slightly lower. But the text analysis showed the training had changed how they reasoned, measurably:

  • Mechanism identification rose by about 0.5 standard deviations — they explained more clearly why their ideas might work.
  • Falsification logic rose by about 0.8 standard deviations — the largest behavioral effect in the entire experiment — meaning they more often spelled out when their ideas might fail.
  • Idea diversity increased: their solutions bundled more varied ideas, and the training group was the only condition with a clear positive effect on between-solution diversity — producing ideas that stood apart from what classmates wrote.

The combined group — training plus ChatGPT — inherited both profiles: grades on par with the ChatGPT-only group, causal-reasoning markers at the level of the training group, and the diversity benefit intact. There was no evidence the combination lifted grades above either single treatment, but nothing cancelled out either. The researchers' practical conclusion: causal reasoning is a skill worth cultivating precisely because its effects survive — and remain independent of — LLM use.

The paradox at the heart of the study: rubrics reward what AI fakes best

The deepest result in this paper is not about ChatGPT; it is about grading. When the researchers examined which textual features correlated with higher scores, they found that more ideas and greater coherence predicted higher grades — while falsifiability, mechanistic detail, and divergence from peers correlated with lower scores. Put bluntly: the traditional rubric rewarded clear, well-structured, standard answers, and systematically failed to notice whether a student had an idea nobody else had.

This is why The Decoder's August 30, 2026 analysis, "The skills that earn top grades are the ones AI can fake best," landed so hard. If the qualities grading systems reward — polish, structure, completeness — are exactly the qualities current LLMs reproduce effortlessly, then grades increasingly measure the tool, not the learner. The paper's authors conclude that evaluation systems need to build originality, reasoning, and consideration of multiple approaches explicitly into their criteria, rather than relying on conventional polished output as a proxy for understanding.

Quick comparison: the four groups side by side

Group N Score (5-point scale) Number of ideas Causal reasoning (mechanisms / falsification) Idea diversity
Control 249 Baseline Baseline Baseline Baseline
Causal training only 256 No gain (slightly negative) Modest increase +0.5 to +0.8 SD Highest — most distinct ideas
ChatGPT only 197 +~1 full point +>0.5 SD No effect Negligible effect
Training + ChatGPT 351 On par with ChatGPT group Increase Both effects together Training benefit retained

Source: the official research paper (August 2026).

What this means for you: five practical takeaways

1 — Use AI as a supervisor, not a ghostwriter. The students who benefited still made the decisions: what to ask, which answer to accept, what entered the final submission. Adopt that as a rule: nothing goes into your assignment that you cannot understand and defend orally in front of your instructor.

2 — Train the "why" and "when it fails" reflex. The training that moved students' reasoning wasn't about AI at all. After any AI-drafted answer, write two lines: why this idea works, and under what conditions it breaks. That single habit produced the experiment's largest reasoning gains.

3 — Understand the rubric before investing in originality. If your instructor grades purely on structure and polish, deep originality may go unrewarded in the short term — even though employers increasingly prize it. Read the assessment criteria first, then decide where to spend your energy.

4 — Access is no longer the barrier. The experiment used GPT-4o, which is available today in ChatGPT's free tier with usage limits and more broadly on paid plans; students comparing options can check our guide to the student discount and budget-friendly alternatives and our Perplexity Pro comparison for students. If your coursework is research-heavy, our review of the best AI tools for academic research maps the landscape.

5 — Follow evidence, not panic. The education debate around AI runs hot. This study is a model of calm reading: check the design, check the numbers, check what the design cannot answer. For how AI tools are actually being integrated into classrooms, see our coverage of Khan Academy's new Gemini-powered Khanmigo features. For a wider look at OpenAI's research and publishing direction this month, read our explainer on OpenAI's new Intelligence Age blog and the Strategic Futures team.

The counter-evidence: what this study does not tell you

Honesty requires presenting the other side, which The Decoder's analysis assembled well:

  • A 30-month study of more than 26,000 Chinese students: homework performance improved with AI use, but exam scores fell — and on later entrance exams, heavy users scored 18–24% lower in the long run. The striking exception: students who kept their study time equal to non-users avoided comparable declines.
  • An analysis of over 500,000 US college grades: the steepest grade inflation occurred in writing- and programming-heavy courses — a pattern consistent with widespread AI-assisted submissions.
  • Controlled removal experiments: once AI was taken away, users underperformed those who had never used it, with the steepest declines among those who had used it mainly to get direct answers. Even roughly ten minutes of answer-machine-style use measurably eroded problem-solving.

The common thread across all the evidence: the danger is not the tool — it is the tool replacing your thinking rather than supporting it.

The study's own limitations, stated plainly

  • No follow-up test without ChatGPT. The design never measured performance after the tool was removed, so improved graded output cannot be distinguished from actual learning. The authors write it explicitly: whether students "acquired and retained" knowledge or merely "procured" output is something "we cannot say."
  • Narrow scope. One university, first-year students only, a single 180-word marketing task, and classroom-level (not individual) randomization.
  • A declared conflict of interest. OpenAI supplied the technology under study and collaborated on the research; several authors are or were OpenAI employees or contractors. The preregistration and blinded grading protect the methodology, but independent replication will matter.

What should universities actually do?

The paper sends a double message to institutions. Banning the tools is neither realistic nor supported by these results — the students who used ChatGPT actively, questioning and selecting, produced the best combination of output quality and reasoning depth. The real lever is assessment design: reward originality, causal reasoning, multiple approaches, and oral defense of ideas — the things a model cannot fake on a student's behalf. And the experiment's quiet revolution is that a single short training session, teachable at scale, produced the largest measurable change in how students think. Curricula should teach the question "why" before they teach any tool.

Frequently asked questions

Did the study prove that ChatGPT harms critical thinking?

No. It found no direct harm during the task, but it also did not measure long-term effects. What it showed: ChatGPT improved output quality, causal-reasoning training improved depth and diversity of thinking, and the two work along different, complementary dimensions. Long-run evidence from other studies suggests harm appears when AI replaces study time rather than supporting it.

Why didn't critical-thinking training raise the grades?

Because the rubric used two standard marketing criteria (raising awareness and usage) that reward clear, well-structured answers. The training's effects appeared in dimensions the rubric did not score: falsifiability, mechanisms, and idea diversity. The researchers noted these very qualities correlated negatively with the grade — a measurement problem, not a skill problem.

Which model did the study use, and can students access it?

The experiment used ChatGPT Edu with GPT-4o during a supervised session. GPT-4o is available today within ChatGPT's free tier with limits, and more extensively on paid plans; our student-pricing guide covers the cheapest legitimate routes.

Do results from an Italian university generalize to students elsewhere?

The core mechanisms — AI lifting polish and substance for novices, training lifting reasoning depth — are not nationality-bound. But caution applies: the task was a marketing case at a single European institution with its own grading culture, and the study itself lists single-site design as a limitation.

How is this study different from the viral "AI makes you dumber" studies?

Most viral claims come from opinion surveys or very short experiments on closed tasks. This is a preregistered RCT with a true control group, a real-world deliverable, dozens of trained human graders, and comparison against expert benchmarks — far higher on the evidence hierarchy, while still carrying the limitations above.

Should universities ban ChatGPT based on these findings?

The study does not support bans. Active use — asking, evaluating, selecting — combined with reasoning training produced the best overall profile. The evidence-aligned recommendation is to permit the tools while redesigning assessments to reward originality and reasoning.

Sources


Start your journey with Truescho: explore fully funded opportunities to study AI and beyond at top universities, or try our Apply-For-Me service so your time goes to what actually compounds: learning that stays.