Which AI writes fiction best
Writing fiction with AI is an engineering project. We break that problem apart and benchmark models task by task.
Entries are scored by models and by people as the task requires; the human scoring is always double-blind. One to two issues will be released per month.
Every issue publishes the full data, the scoring rules and the limits of the test. We hope it helps in your writing.
All reports
Each issue tests something different. Every report carries the findings plus all the material the models produced.
One beat sheet, ten drafts — xianxia
Xianxia with dark Journey to the West elements: the protagonist wakes as the Turtle Prime Minister's least-favoured thirteenth son.
Outline, beat sheet and character notes were identical; 10 models each wrote 3 independent first chapters, scored double-blind by human reviewers. Includes the leaderboard, five-axis breakdown, consistency, cost analysis and a look at AI tells — every chapter and beat sheet is readable in full.
Benchmarking the whole writing workflow, from Plan onward
Martial arts in the late Qing: the protagonist wakes up as Runtu and trains in the fighting arts. One prompt, one set of follow-up constraints, three runs each for 18 models: Plan elicitation → planning documents → beat sheet → chapter one. “Writing ability” is split into four separately scored axes — elicitation, structure, instruction following and prose — with time and cost layered on top. The four are independent, and no model wins all of them.
How we test
Controlled inputs, automated collection, double-blind human review, principled data cleaning.
One shared input
Whatever the writing scenario, every model gets the same input and generates its runs independently, with no cross-contamination.
Human blind review
Writing quality can't be left to model graders — a human reading the prose tells you more. Our reviewers never see which model wrote what.
Clean, then count
Outliers are removed on a stated rule, reviewers whose scores show no discriminating power are dropped wholesale, and only then do we count. Sample sizes and definitions are fully public.
Join the review panel
We're always recruiting people who read a lot of web fiction. Every human score and comment in these issues comes from experienced readers and reviewers in the SoloEnt community.
Join the panel →What should we benchmark next?
Tell us which genre, which model or which writing scenario you want to see benchmarked next.