3 min read

Readers preferred AI stories—and missed the tell

A Villanova study found readers rated ChatGPT stories higher than human ones and identified their authors at roughly chance levels.

Image: TechXplore

Readers rated ChatGPT-generated short stories more highly than human-written ones—and gave the AI stories their best scores when they were told a person had written them. The findings come from a study led by researchers at Villanova University and published in Judgment and Decision Making.

The research adds to evidence that people struggle to identify AI-generated writing, even when asked directly. It also suggests that familiarity with AI systems, rather than literary expertise, is the factor most likely to improve detection.

How the story tests worked

The researchers selected three fictional stories written by humans and published in reputable literary journals or short-story collections. They then used ChatGPT to generate three corresponding stories.

The first experiment involved 1,682 adults aged 18–81. Each participant read one story and was either correctly or incorrectly told that it had been written by a human or by AI. Participants then rated its quality and how engaging or absorbing they found it.

Recommended reading

A $10 ESP32 can run a tiny language model

Two further experiments tested recognition more directly. In the second, 424 participants read one human-written and one AI-generated story. The third used the same setup with 481 participants. In both, participants were not told the stories' origins and had to identify which author type had produced each one.

AI writing scored better, but detection stayed near chance

AI-generated stories received higher ratings for both quality—how well-written they seemed—and absorption, or how invested readers felt in them. The strongest scores went to AI stories that participants had been told were written by humans.

Participants performed poorly when asked to identify the author. Accuracy was 39.93% in experiment two, below chance, and 51.97% in experiment three, which was no different from chance.

“In our study, participants gave the highest ratings to stories that they had been told were written by humans but which were actually written by AI.”

Dr. Deena Weisberg, Villanova University

Weisberg said the result points to a bias toward human-created narratives. Readers may assume that creative writing depends on emotional understanding and lived experience, leading them to underestimate what AI can produce.

“I’d argue that this reveals a bias toward narratives written by real people. We assume creative writing requires uniquely human qualities, such as emotional understanding and lived experience, which leads people to underestimate AI’s capabilities. I’d say that public assumptions about AI’s capabilities are increasingly out of date.”

Dr. Deena Weisberg, Villanova University

AI literacy helped more than literary expertise

Participants who reported greater experience with AI platforms such as ChatGPT were better at recognizing AI-generated text. Self-reported expertise in fictional literature did not produce the same improvement.

In experiment two, every one-point increase in a participant’s self-reported AI expertise score corresponded to a 14% increase in the odds of correctly identifying the author. In experiment three, a one-point increase on the previously validated Artificial Intelligence Literacy Scale (AILS) corresponded to a 33% increase in the odds of a correct guess.

Weisberg said AI familiarity appeared to help readers recognize recurring patterns, including em dashes and sentence structures such as “it’s not just X, it’s Y.” She argued that improving AI literacy could help people navigate AI-generated content more effectively.

The researchers also suggested a possible reason for the preference gap: AI writing tends to be clearer, more direct and easier to process, while human writing may be more subtle and complex. The AI versions generally stated their themes explicitly instead of asking readers to infer meaning from characters' words and actions.

The study does not establish that contemporary short-form media caused these preferences. Weisberg said such technologies may amplify existing tendencies, including a preference for predictability when subtle material requires more mental effort. The clearest result is narrower but consequential: readers rewarded AI prose, while their ability to recognize it remained roughly no better than guessing.

Ava Chen

AI Editor

Ava covers the rapidly evolving world of artificial intelligence, from foundational models and research labs to the real-world economics of intelligence. With a background in computational linguistics, she cuts through the hype to find out what actually works. She firmly believes that benchmarks are just marketing until reproduced in the wild.

via TechXplore

/ Keep reading