Which AI writes fiction best
Writing fiction with AI is an engineering project. We break that problem apart and benchmark models task by task.
Entries are scored by models and by people as the task requires; the human scoring is always double-blind. One to two issues will be released per month.
Every issue publishes the full data, the scoring rules and the limits of the test. We hope it helps in your writing.
All reports
Each issue tests something different. Every report carries the findings plus all the material the models produced.
Xianxia / high-concept / period head-to-head
Three prose benchmarks — xianxia, high-concept and period — spanning 90 works and 306 double-blind ratings. The focus: how genre-picky each model is, how far the beat sheet pulls the prose apart, and whether version upgrades and pricier siblings are worth the money.
One beat sheet, ten drafts — period fiction
10 models take the same beat sheet and each write three opening chapters of period fiction set in 1979 Beijing. Double-blind scoring covers the overall ranking, five abilities, period detail, consistency and cost — every chapter and every comment is public.
One beat sheet, ten drafts — high-concept
10 models, one human-written beat sheet, three high-concept openings each. Double-blind scoring covers the overall ranking, five axes, AI tells, consistency and cost — every text and comment is public.
One beat sheet, ten drafts — xianxia
Xianxia with dark Journey to the West elements: the protagonist wakes as the Turtle Prime Minister's least-favoured thirteenth son.
Outline, beat sheet and character notes were identical; 10 models each wrote 3 independent first chapters, scored double-blind by human reviewers. Includes the leaderboard, five-axis breakdown, consistency, cost analysis and a look at AI tells — every chapter and beat sheet is readable in full.
Benchmarking the whole writing workflow, from Plan onward
Martial arts in the late Qing: the protagonist wakes up as Runtu and trains in the fighting arts. One prompt, one set of follow-up constraints, three runs each for 18 models: Plan elicitation → planning documents → beat sheet → chapter one. “Writing ability” is split into four separately scored axes — elicitation, structure, instruction following and prose — with time and cost layered on top. The four are independent, and no model wins all of them.
How we test
Controlled inputs, automated collection, double-blind human review, principled data cleaning.
One shared input
Whatever the writing scenario, every model gets the same input and generates its runs independently, with no cross-contamination.
Human blind review
Writing quality can't be left to model graders — a human reading the prose tells you more. Our reviewers never see which model wrote what.
Clean, then count
Outliers are removed on a stated rule, reviewers whose scores show no discriminating power are dropped wholesale, and only then do we count. Sample sizes and definitions are fully public.
Join the review panel
We're always recruiting people who read a lot of web fiction. Every human score and comment in these issues comes from experienced readers and reviewers in the SoloEnt community.
Join the panel →What should we benchmark next?
Tell us which genre, which model or which writing scenario you want to see benchmarked next.