OpenAI-Bocconi RCT Finds ChatGPT and Critical-Thinking Training Deliver Different Benefits
A randomized study of 1,053 first-year students found that ChatGPT improved conventional task performance while causal-reasoning training increased idea diversity and mechanism-based thinking.
A randomized test of AI assistance and reasoning training
OpenAI Economic Research and Bocconi University researchers published results on August 27, 2026 from a preregistered randomized controlled trial examining a practical education question: does giving novices access to an LLM improve their work in the same way as training them to reason more carefully?
The study included 1,053 first-year undergraduates in economics, management and finance. Classes were assigned to one of four conditions: ChatGPT Edu access using GPT-4o, causal-reasoning training, both interventions, or neither. Students then completed a real-world merchandising case and produced short recommendations.
Because the design manipulated both AI access and reasoning training, the researchers could separate the effects of each intervention and examine what happened when they were combined.
ChatGPT improved conventional performance
Students with ChatGPT access produced work that scored higher on the study's standard five-point evaluation rubric. The paper estimates an improvement of about 0.86 points relative to the control group's estimated score of 2.09. Their responses also contained more ideas, showed stronger textual coherence and were closer to recommendations produced by domain experts.
The researchers report that coherence and the number of ideas explain part, but not all, of the performance gain. Their interpretation is that ChatGPT improved not only presentation but also the substance of solutions on this well-defined task.
This result should not be generalized to every educational activity. The task had explicit criteria, the students were novices, and the model was operating within a domain where it had relevant knowledge. The study does not show that AI access automatically improves deep understanding or performance on problems outside a model's capability range.
Causal-reasoning training changed the kind of thinking
The causal-reasoning intervention produced a different effect. Students trained to reason about mechanisms and falsifiable conditions generated a wider range of ideas and were more likely to explain why an idea should work and under what conditions it might fail.
Those changes did not translate into a higher score on the conventional rubric. The researchers argue that this exposes an assessment problem: a rubric optimized for standard objectives can reward polished, conventional solutions while failing to value idea diversity or explicit causal reasoning.
Importantly, the causal-reasoning benefit did not disappear when ChatGPT was available. Students who received both interventions retained the broader idea diversity associated with reasoning training while also showing the performance benefits associated with AI access.
Why the result matters
The study complicates the common framing that education must choose between AI tools and teaching students to think independently. In this experiment, the interventions were complementary rather than substitutes.
For schools and universities, the more difficult implication concerns assessment. If AI can help novices generate polished, expert-like answers to well-specified tasks, final-answer quality alone may reveal less about what a student understands. Assignments may need to evaluate reasoning, originality, assumptions, evidence and the ability to compare alternatives, not just whether a response matches a standard solution.
For employers, the result may also be relevant to junior knowledge work. AI can narrow a conventional performance gap on structured tasks, while human reasoning training may contribute value in dimensions that standard evaluation systems do not automatically reward.
The paper is one experiment in a specific context, not a universal verdict on AI in education. Its strongest contribution is the randomized evidence that LLM assistance and cognitive-skill training can change different dimensions of novice work—and that assessment design determines which of those gains are visible.
This article is built from the source material below. Open the originals for full context and the latest updates.