Better answers, broader thinking: What students gain from ChatGPT and critical-thinking training
In a randomized experiment with more than 1,000 students, access to ChatGPT and training in causal reasoning improved student work in complementary ways.
What happens when students use ChatGPT on a real-world assignment? Do quality improvements come at the expense of originality?
A new experiment from researchers at Bocconi University, in collaboration with OpenAI Economic Research, found distinct and complementary effects from ChatGPT access and critical-thinking training. Access to ChatGPT improved the quality and coherence of students’ work, while an exercise in causal reasoning—a form of critical thinking—led students to generate more unique ideas. Students who received both ChatGPT access and the training showed both effects.
The results also highlight the importance of a holistic approach in assessing student progress. As AI makes it easier for students to produce polished answers, assignments may need to adapt to measure other desired qualities such as originality.
During the experiment, more than 1,000 first-year undergraduate students at Bocconi University worked on a real-world business case developing marketing recommendations for the university’s merchandise store.
Students were randomly assigned by class period into one of four groups that received: access to ChatGPT (GPT‑4o), training in causal reasoning, both, or neither. Causal reasoning is a specific form of critical thinking related to linking cause and effect and explaining why a given solution may or may not work. The training students received was unrelated to AI, instead teaching students causal reasoning concepts through an exercise that involved a game, examples, questions, and feedback.
Student submissions were evaluated by trained human graders using a five-point rubric. Separately, researchers used automated text analysis to measure each submission’s number and variety of ideas, signs of causal reasoning, and similarity to submissions from three experts.
The students who had access to ChatGPT scored almost a full point higher on the five-point scale. Their answers included more ideas, followed clearer logic, and were more similar to recommendations written by experts. In other words, AI helped novices produce work that looked more professional. Importantly, students weren’t simply handing over their assignments to ChatGPT. They still had to decide what to ask, evaluate the responses, and choose what went into their final submission.
The critical-thinking exercise produced a more unexpected result. Students who completed the exercise explained more clearly why their ideas might work and when they might fail, but did not score higher on the grading rubric, which only measured how well the recommendations addressed two standard marketing goals: increasing awareness and use of the university store.
Text analysis revealed another benefit that the rubric did not capture. Across the group, students who completed the exercise produced a wider range of ideas that were more distinct when compared to what their peers produced. This matters because a traditional rubric, like the one used to grade the students in this experiment, can reward a clear, well-structured answer while overlooking whether a student came up with an idea that no one else did.
It can be tempting to frame the education debate as a choice of whether students should learn to think for themselves or learn to use AI.
This experiment highlights that both are valuable in different ways. AI access helped students produce answers that were more polished, idea-rich, and more logically coherent. Critical-thinking training encouraged them to develop a wider range of original ideas, question assumptions, and explain why their ideas should work.
Together, they are complementary.
Students who received both ChatGPT access and the critical-thinking exercise showed the benefits of each. Their idea variety matched that of students who completed only the exercise. Their rubric scores and number of ideas were similar to those of students with ChatGPT access alone. Their work also showed stronger logical coherence and more evidence of looking for explanations and questioning assumptions. Overall, this group showed gains across the widest range of measures.
The experiment’s randomized design allowed researchers to dig into these different factors, separating the effects of ChatGPT access and the critical-thinking exercise from the effect of combining them. That makes the experiment a particularly useful contribution to a rapidly growing body of research on the impact of AI on students and how to best structure and support their learning.
The difference between the impact of the exercise on students and what was captured by the traditional-style rubric points to a broader challenge facing schools.
If AI can help students produce polished, expert-like work, then looking only at the final answer tells us less about what a student actually understands.
Many educators are already grappling with the implication that assignments and evaluations may need to change. This follows a multi-decade pattern of technology and education evolving together in service of supporting the skills students need in modern society.
These results highlight the importance of rewarding students for producing work that reflects originality, reasoning, and consideration of multiple approaches, not just the most conventional or polished answers.
The takeaway: AI helped students make their answers better. Critical-thinking training helped make their ideas broader. The two play complementary roles in preparing students for the future.

