Generative AI in higher education: Mapping the behaviour of students and machines
Source PublicationScientific Publication
Primary AuthorsRojas, Guerra, Frez et al.
"Working with an AI peer is like dancing a tango with a robot. If you lead clearly and respond to its steps, you create a smooth routine. If you just stand there and expect it to drag you across the floor, you will both trip over."

The study claims that active, step-by-step collaboration with an artificial intelligence significantly improves a student's ability to solve complex maths problems. Yet, to appreciate this finding, we must immediately pivot to the historical difficulty of mapping this genome. For decades, charting the exact sequence of human DNA was a massive, frustrating task. Scientists had millions of data points but struggled to understand how they all interacted. Today, scientists evaluating Generative AI in higher education face a remarkably similar obstacle. They have vast amounts of chat logs, but making sense of how students and machines actually work together remains incredibly difficult.
To understand the analytical methods used here, we can look at how biologists map physical traits. Researchers often contrast 'gene markers' with 'GC content'. Gene markers are specific, identifiable DNA sequences that flag a known physical trait or disease. They are precise but narrow, offering a highly targeted view of a specific function. On the other hand, GC content measures the overall percentage of guanine and cytosine bases across a broad stretch of DNA. This provides a wide, structural overview of genome stability, but it lacks the fine detail needed to identify specific biological mechanisms. The old method of analysing student data relied on broad metrics like GC content—looking only at total time spent or final grades. This broad view has massive blind spots. The new method acts more like gene markers, tracking specific conversational turns to identify exactly where a student's logic succeeds or fails.
Generative AI in higher education: The new method
Researchers tested 30 university students using a web platform. The system paired each student with a large language model to solve calculus problems. Unlike older, optional tutoring software, this system split the information. The student and the AI had to talk to each other to find the answer. The researchers measured the chat logs using an automated scoring pipeline.
They found eight distinct types of student behaviour. Some acted as Co-Regulators, carefully checking the machine's logic. Others were Low-Effort Guessers, blindly asking for the final answer. The data showed a stark contrast. Students in the high-interaction group solved both maths tasks 71 per cent of the time. Those in the low-interaction group succeeded only 19 per cent of the time.
This new method of tracking individual chat prompts is highly efficient for spotting specific learning gaps. However, based on this specific cohort of 30 undergraduates solving calculus problems, it still harbours potential blind spots. The automated scoring pipeline demonstrated lower agreement with human coders than humans did with one another, indicating that the machine still misses certain subtleties in student interaction. While the system measured clear differences in success rates, it suggests that human-AI teamwork is far from uniform. Simply dropping a chatbot into a university programme could fail if the task does not force the student to engage. Designing the right feedback loop remains essential.