The Brief
Creativity & Culture 4 min read

Only 39% Could Tell the Difference. That Says a Lot.

NAVION

Share

A study involving more than 1,600 adults has produced a result that cuts against one of the most common assumptions in the current debate about artificial intelligence and creative work: that AI-generated writing is obviously inferior, and that people can easily spot it. The data suggests neither is true.

What the Numbers Actually Show

Researchers Sydney Sears and Deena Weisberg at Villanova University in Pennsylvania designed a careful experiment. They recruited 1,682 adults, selected to reflect the US population in terms of gender, age, and race, and asked each participant to read a single story. Half read a story written by a human, sourced from a literary journal. The other half read a story generated by ChatGPT 4.0.

Each group was then split again. Some participants were told the truth about what they had read. Others were told the opposite. This design allowed the researchers to separate two distinct questions: how good is the writing, and how much does the label matter?

The quality scores, measured on a scale from -3 to 3, told a clear story. Participants who read AI-generated stories rated them at an average of 1.54. Those who read human-written stories gave an average score of 0.97. The AI stories also scored higher on how absorbing they felt: 1.42 versus 1.00 for human-written work. One additional pattern emerged across all conditions: stories were rated slightly higher when participants believed a human had written them, regardless of the actual source.

The Identification Problem

The researchers ran two further studies to test whether people could reliably distinguish AI output from human writing. In the first, only 39 percent of participants correctly identified which story was which. In the second, conducted several months later, 52 percent were correct. The researchers note they have no definitive explanation for the gap between the two results, though they suggest that society may be gradually adapting to recognize AI-generated content. Either way, the performance in both tests hovers close to what one would expect from random guessing.

This is what most coverage of AI writing tends to skip over. The debate often assumes a clear perceptual divide between human and machine output, one that readers can navigate intuitively. These results challenge that assumption directly.

Claire Hardaker at Lancaster University points to a broader discomfort at work here. Her own research on people’s ability to identify AI-generated speech, text, and music found that music provokes the strongest emotional reaction when people discover they have been deceived. The creative domain, in other words, is not just a technical question. It touches something people feel protective about.

Rodney Jones at the University of Reading offers a more structural explanation for why AI writing scored higher in this particular study. The human stories were drawn from a literary journal, a genre that tends toward complexity and demands more effort from the reader. AI, asked to produce stories with similar traits, likely defaulted to something more accessible and easier to process. Jones puts it plainly: literary writing is not very popular, because it is hard work. People tend to prefer content that flows without friction.

What This Means Beyond the Experiment

The finding that people prefer AI writing, at least in this context, does not settle the larger question of what AI creativity actually is or what it is worth. Weisberg herself draws a distinction. She describes AI output as more “vanilla” than human writing, and attributes the quality gap in the scores to that accessibility rather than to any deeper artistic achievement. She plans future experiments across different types of writing to test whether the pattern holds.

Her broader view is worth sitting with. She argues that AI will be capable of producing work with genuine artistic value, but that this value may need to be measured on a different scale than the one applied to human creativity. The two are not the same thing, even if they can produce outputs that are difficult to tell apart.

This is the tension the study surfaces without fully resolving. Cultural resistance to AI-generated art is real and documented. The study itself notes that a short story awarded the Commonwealth prize and published in Granta was accused by some of being AI-generated, something the competition explicitly prohibited, though an investigation was inconclusive. Artists have raised objections to AI-generated art appearing at major auction houses, arguing that models trained on existing images represent a form of unauthorized use of human work.

Those concerns are not addressed by preference scores. People can simultaneously prefer something and object to how it was made. The experiment measures one thing: how readers respond to text in isolation, without knowing its origin. It does not measure whether that preference would survive full transparency, or whether it would hold across genres, lengths, and contexts.

In Short

Most people in this study rated AI-generated stories higher than human-written ones, and fewer than half could correctly identify which was which. The preference appears linked to readability rather than depth. AI writing tends to be smoother and easier to process, which scores well in short-form tests but may not translate to every kind of creative work. The deeper question, what it means for human creativity when machines can produce indistinguishable output at scale, remains genuinely open.

Based on reporting from New Scientist.

Written by

NAVION