Skip to main content
Abstract digital art representing GPT-4o's advanced AI, with glowing neural networks and data streams.

Editorial illustration for GPT-4o's Unidentified Edge Persists in Analysis

GPT-4o Outperforms Students in University Test

GPT-4o's Unidentified Edge Persists in Analysis

4 min read

Bocconi University ran an experiment last November that most professors would rather not think about too hard. Researchers split 1,053 freshmen across 13 sections of an introductory management course into four groups, then asked each student to write marketing recommendations for the school's merchandise shop in 180 words or less. Some got access to GPT-4o.

Some got a short lesson on causal reasoning, covering things like falsifiability and whether a proposed action would actually produce the outcome it claimed. Some got both. Some got neither.

The results split cleanly along a line that should worry anyone grading student work by hand. One intervention moved grades. The other didn't, but changed how students thought.

That gap raises an uncomfortable question about what business school assignments have been rewarding all along, and whether the traits that earn a 5 out of 5 are just the traits a language model happens to be good at faking. The study didn't test whether students learned anything once the chatbot was taken away.

What makes a good student paper, and which parts of that can AI improve? A randomized experiment with 1,053 freshmen at Bocconi University found that GPT-4o helped students earn significantly better grades on a business assignment.

Why this matters

The Bocconi study lands right where a lot of AI-in-education debates get sloppy: assuming better output means better learning. It doesn't, or at least this experiment can't tell us that it does. Grades went up, the researchers controlled for the obvious variables, and GPT-4o still left a residue of advantage the authors pin on raw content quality, not student understanding.

That's worth sitting with if you build tools for classrooms or evaluate talent using written work. If a rubric rewards polish, argument structure, and idea diversity, GPT-4o can supply all three without the student knowing more than they did before. The causal reasoning lesson is the more interesting thread: it didn't move grades, but it pushed students into weirder, less templated solutions.

That suggests grading rubrics and AI capability are now optimizing for the same narrow band of "quality," which is a problem for anyone using grades, or GPT output, as a proxy for competence. No follow-up test without ChatGPT means we still don't know if anyone learned anything. That gap should worry educators more than the grade bump itself.

Common Questions Answered

What was the Bocconi University experiment design for testing GPT-4o's impact on student grades?

Bocconi University researchers split 1,053 freshmen across 13 sections of an introductory management course into four groups and asked each student to write marketing recommendations for the school's merchandise shop in 180 words or less. Some students received access to GPT-4o while others received a short lesson on causal reasoning covering concepts like falsifiability and whether proposed actions would produce intended outcomes. This randomized experiment allowed researchers to measure the impact of AI assistance on student performance and grades.

Did GPT-4o provide a measurable advantage in the Bocconi study, and what was the source of that advantage?

Yes, GPT-4o helped students earn significantly better grades on the business assignment in the Bocconi study. The researchers controlled for obvious variables and found that GPT-4o left a residue of advantage that the authors attributed to raw content quality rather than improved student understanding. This distinction is important because it suggests the AI improved the output quality without necessarily enhancing the students' learning or reasoning abilities.

Why does the distinction between better grades and better learning matter according to the article?

The article argues that the Bocconi study reveals a critical flaw in AI-in-education debates that assume better output automatically means better learning. The study demonstrates that grades went up with GPT-4o assistance, but this doesn't necessarily indicate that students developed deeper understanding or improved their actual skills. This distinction is particularly important for those building educational tools or evaluating talent through written work, as it raises questions about whether AI-assisted improvements reflect genuine learning or merely better-looking submissions.

What specific skills did GPT-4o appear to improve in student papers according to the Bocconi experiment?

According to the article, GPT-4o improved the skills that earn top grades, which are notably the same skills that AI can fake best. The experiment found that the AI's advantage came from raw content quality improvements rather than from enhancing students' causal reasoning or critical thinking abilities. This suggests that GPT-4o excelled at producing polished, grade-worthy content in areas where AI is particularly capable of mimicking quality.

LIVE13:12AI Agents Solved Tasks but Couldn't Track Time Accurately