SoloEnt Novel Benchmark

Which AI writes fiction best

Writing fiction with AI is an engineering project. We break that problem apart and benchmark models task by task.
Entries are scored by models and by people as the task requires; the human scoring is always double-blind. One to two issues will be released per month.
Every issue publishes the full data, the scoring rules and the limits of the test. We hope it helps in your writing.

0Reports published
0Models benchmarked
0Anonymous works
0Valid ratings

How we test

Controlled inputs, automated collection, double-blind human review, principled data cleaning.

1

One shared input

Whatever the writing scenario, every model gets the same input and generates its runs independently, with no cross-contamination.

2

Human blind review

Writing quality can't be left to model graders — a human reading the prose tells you more. Our reviewers never see which model wrote what.

3

Clean, then count

Outliers are removed on a stated rule, reviewers whose scores show no discriminating power are dropped wholesale, and only then do we count. Sample sizes and definitions are fully public.

Score with us

Join the review panel

We're always recruiting people who read a lot of web fiction. Every human score and comment in these issues comes from experienced readers and reviewers in the SoloEnt community.

Join the panel →
Requests

What should we benchmark next?

Tell us which genre, which model or which writing scenario you want to see benchmarked next.