Industry

A study gave students ChatGPT and their grades went up. That is the problem

A trial of 1,053 students found ChatGPT lifted marks 0.86 points. The group taught to reason scored lower. The rubric explains both results.

Multi Chats Team
September 26, 2026 · 6 min read

A randomised trial gave first-year undergraduates access to ChatGPT and graded their work. The group with access scored almost a full point higher on a five-point scale. The group taught a thinking skill instead scored slightly lower.

Both numbers came from the same marking scheme. That is the finding worth your time.

The paper is "Training novices to think, or giving them LLMs? Evidence from an RCT", dated August 2026 and circulated as CEPR Discussion Paper 21882 on 27 August 2026. Every figure below comes from the paper itself, read on 12 September 2026.

What the trial did

  • 1,053 first-year undergraduates, across 13 class sections of one introductory management course.

  • Three bachelor programmes: economics and management, economics and finance, management.

  • Run in November 2025, in class, under proctor supervision. One session, roughly 45 minutes.

  • Four arms, randomised at class level: Control (249 students), Causal (256), GPT (197), Causal and GPT (351).

  • The causal arms played a training game on causal reasoning, with worked examples and feedback after every answer. The other arms answered the same 12 questions on the same case with neither.

  • Students in the AI arms "could use ChatGPT Edu during the task (GPT-4o)". That is the model, named as the paper names it.

Then everyone got the same problem. The university's merchandising division wants alumni to notice and buy its branded goods. Students received operational data on sales channels, pricing and an alumni base of more than 144,000. They wrote a recommendation of about 180 words. The main text keeps the university anonymous; the appendix names Bocconi.

How it was marked

  • Two attributes, taken from a standard marketing framework: awareness and usage.

  • Each scored 1 to 5 on a Likert scale, with written anchors for every level.

  • Twenty trained evaluators. Three read every response. 3,159 ratings in all.

  • A response's score is the average across its three evaluators.

  • The trial was pre-registered with the AEA registry before data collection and cleared by the university's ethics committee.

The grades went up

Access to ChatGPT raised the awareness and usage score by 0.834 points with no controls, and 0.862 with balance controls. The paper's own framing: 0.86 against an estimated control score of 2.09. Both estimates are significant at the 1% level. On a scale that runs from 1 to 5, that is a lot.

Two figures sit alongside it:

  • The GPT arms put 2.25 more ideas into each recommendation.

  • Their texts landed closer to three expert-written model solutions, by 0.026 to 0.044 on the paper's similarity measure.

The paper's own summary of it: "LLMs make novices resemble experts on standard tasks on which LLMs are well trained."

The group taught to think scored lower

The training did work, on its own terms. Against control it raised mechanism identification by about 0.55 standard deviations and falsifiability by about 0.85. It produced the strongest effect on diversity of ideas inside a single answer. It was the only treatment with a clear positive effect on distance from everyone else's answers.

Its effect on the mark was 0.101 points down without controls, and 0.111 down with them. The second estimate is significant at the 1% level.

More original, slightly worse grade. Same task, same graders, same week.

The rubric is doing this

The paper pulls the score apart with a lasso regression. Coherent logic pushes the mark up hard. Falsifiability pushes it down, by 0.109 to 0.159 points. Mechanism identification pushes it down further, by 0.166 to 0.242. Distance from everyone else's answers carries the heaviest negative weight in the table.

So the two habits the training installed are the two the marking punished. The paper is explicit: the causal treatment's negative effect vanishes once those variables go in.

The authors put it in their own words: "The rubrics we used in the study penalizes distance from the conventional answer." And in the conclusions: "Distance from the standard, in other words, is not rewarded by the rubric, even where it may reflect a more original answer."

What the study does not measure

Nobody sat these students down afterwards without the tool. The authors state the gap themselves, in the conclusions:

"These results leave open a question our design cannot settle. We observe the performance of learning, and we can decompose it: About half of the LLM advantage operates through clearer, richer, better organized text, and about half through the substance of the recommendations. Whether the second component reflects knowledge that participants acquired and retained, or output they procured without acquiring anything, we cannot say."

That is a limit on what one 45-minute session can show, and the people who ran it wrote it in.

Who wrote it, and what is never stated

Read the author list before the abstract. Twelve names:

  • Bocconi University, SDA-Bocconi and the ION Management Science Lab: Betti, Camuffo, Fumagalli, Gambardella, Mariani, Pandey, Salvucci, Simic.

  • OpenAI: Asirvatham. Aaron Chatterji is listed as OpenAI and Duke University.

  • Two footnotes on the title page. Rachel Brown, listed as an independent researcher, "contributed to this work during her employment at OpenAI". Steve Ramos, of the University of California, Berkeley, "contributed to this work in his capacity as a paid contractor for OpenAI".

Some other facts of the same kind. The tool under test was ChatGPT Edu. The PDF sits on OpenAI's own content delivery network. Human evaluators scored performance, while the causal-reasoning scores, the idea extraction and the embeddings behind the similarity and diversity measures all came from OpenAI models. There is no acknowledgements section and no funding statement, so who paid is not stated anywhere in it.

Why weigh that as a reader? Four of the twelve have a stated OpenAI connection, and the headline result flatters a product OpenAI sells. Two things in the design survive that reading. It was pre-registered before any data came in. And the least flattering finding, that the marking scheme rewards conventional answers, is in the paper rather than left out of it.

What we would take from it

  • A higher mark says the answer matched what the rubric rewards. It says nothing about whether the student could produce that answer again, alone.

  • Originality costs marks when the scale has no line for it. The paper's own recommendation is that assessment has to ask for diversity if it wants any.

  • Fluency and a longer list of ideas carry much of the gain, though about half of the effect survived every textual control. The substance improved too.

  • One session, one well-defined business case, one model. Nothing here reaches proofs, essays or an exam you sit without a laptop.

Two neighbouring questions, kept separate here. On integrity, the usable line is whether you still do the thing being assessed. On revision, being tested beats reading a summary that felt clear at the time.

MultiChats carries 60 models from 18 providers under one subscription, and a second model is often the quickest way to find out that a confident answer was wrong. None of that changes what a mark measures. The grade in this trial rose because the text got cleaner and the list of ideas got longer, and the people who ran it cannot tell you whether anything stuck.