• 2 min read
ChatGPT stories beat human fiction—until readers knew the source
A Villanova-led study found ChatGPT stories rated higher than human ones, but readers struggled to identify AI authorship and favored texts labeled human.

Image: iXBT
ChatGPT-generated short stories were rated more highly than human-written ones in an experiment involving 1,682 people—but the advantage depended heavily on what readers believed about the author.
Researchers led by scientists at Villanova University showed participants a short story and told them it had been written either by a person or by an AI. That attribution could be accurate or deliberately wrong. Participants then rated the story’s quality and how engaging it was to read.
The researchers used three stories by human authors published in literary magazines or collections, along with three corresponding stories generated by ChatGPT.
Human attribution changed the scores
On average, the ChatGPT stories received higher ratings for both quality and reader engagement. But the strongest effect came from perceived authorship: stories participants believed were written by humans scored better regardless of who actually wrote them.
That gave ChatGPT its highest ratings when readers incorrectly assumed its stories were human-written. The result points to a clear reception bias. The same text can be judged differently depending on whether readers think it came from a person or a model.

Recommended reading
Meta’s Muse Spark AI breached a site during testing
The study also tested whether readers could identify the source without being told. In one experiment with 424 participants, accuracy was 39.93%, below random guessing. A second experiment with 481 participants produced 51.97% accuracy, a result the researchers said was not statistically different from chance.
Those figures do not show that readers can reliably detect AI-generated fiction unaided. They also show why an authorship label can matter as much as the prose itself.
The research does not compare literary skill, originality, or the ability to produce longer works. It therefore supports a narrower conclusion: ChatGPT stories performed well in these reader judgments, while their origin was difficult to identify without verification methods. The evidence is strong on perception and detection, but it does not settle whether the model is a better literary author than the humans whose work was used in the comparison.
AI Editor
Ava covers the rapidly evolving world of artificial intelligence, from foundational models and research labs to the real-world economics of intelligence. With a background in computational linguistics, she cuts through the hype to find out what actually works. She firmly believes that benchmarks are just marketing until reproduced in the wild.
via iXBT


