Readers preferred the AI's short stories, until they were told
2,500 participants could not tell ChatGPT fiction from human fiction, and rated the AI's higher. The ratings flipped the moment a story carried an AI label, true or not.
2,500 readers rated ChatGPT's fiction above human stories. Then researchers told them who wrote what.
A study published August 8 in Judgment and Decision Making, a Cambridge University Press journal, ran more than 2,500 people through three experiments on ChatGPT-written short stories. Participants could not identify which stories were machine-written better than chance. Blind, they rated the AI stories higher: 1.54 versus 0.97 on quality, 1.42 versus 1.00 on immersion. Told a story was AI-written, they marked it down, whether or not the label was true.
The design that makes it credible
Researchers Sydney Sears and Deena Skolnick Weisberg split the work: 1,682 people rated single stories, 905 compared pairs, and the authorship labels were randomized, sometimes truthfully, sometimes not. That last move separates the two effects cleanly. Text quality moved ratings one way; the word 'AI' moved them the other, independent of what was actually on the page. Readers' preexisting attitudes toward AI predicted their ratings regardless of true authorship.
The label is the product decision
The penalty attaches to disclosure, not to the writing. That lands directly on a live policy question: platforms and publishers are converging on AI-labeling rules, and this data says the label itself changes reception even when readers cannot detect the difference. The tradeoffs run both directions. Hiding provenance juices ratings and burns trust when discovered; honest labels cost you the blind-preference margin. The study also has limits worth naming: short fiction only, one model, lab conditions, and enjoying a story is not the same as paying for one.
Why a build studio cares
We ship products that generate content, and our own site carries AI-assisted work. This study puts numbers on something we treat as a rule anyway: disclosure is a trust decision, not a quality decision, and you should make it expecting the discount. If your product's copy, summaries, or stories are model-written, A/B testing the label is now an evidence-backed experiment, not a hunch.
Next step: read the paper at Cambridge or The Decoder's summary. If you are deciding how to label AI-generated content in a product, write to us at hello@gattyworks.com.