One beat sheet, ten drafts · Xianxia
The genre is dark Journey to the West: the protagonist wakes up in the Dragon Palace death row as the least-favoured thirteenth son of the Turtle Prime Minister. Outline, chapter beat sheet and character notes were identical for everyone; each model generated three chapters independently, with no sight of the others. All 117 reviewer ratings were double-blind — the people scoring saw only a four-digit random code, never the model behind it.
Place a bet before you look
Go on instinct: which of these ten writes web fiction best? Pick one, then see how good your instinct is.
The leaderboard reshuffles with every axis
The ranking uses the reviewers' own overall score — a field they filled in directly, not a weighted average of the five axes. Tap an axis below to watch the board reshuffle: the model on top is not on top of everything.
The 0.18 — 0.20 spread between the three top-tier models is not statistically significant, so this report groups models into tiers instead of pinning down exact places.
Five axes: where each model wins and where it collapses
“Logic problems” and “AI tells” are better when lower, so both are converted into quality scores (10 − raw) — on all five axes, higher is better. Pick two models and compare.
Radial scale runs 2 — 8 (every model's actual scores fall inside that band)
Consistency: same input, three runs
Each dot is one generation, the bar is the spread across the three runs and the tick is the mean. The tightest range is 0.22, the widest 2.56 — more than a tenfold difference. Hover a dot for the exact score.
Steady isn't good
Claude Fable 5 has the smallest range in the field at 0.22, but all three chapters are stuck at 4.7 — 4.9. The problem is systematic; running it again will not fix it.
Only one is both high and steady
GPT-5.6 Sol is 3rd on the mean with a range of 0.23 — every delivery lands on the same line. Claude Opus 5's best chapter, 6.23, would place in the field's top 12; its worst, 3.67, is dead last.
Where the money goes: 163× the price, 1.4× the quality
Cost per chapter runs along the x-axis (log scale), overall score up the y-axis. If you got what you paid for, the dots would line up on the diagonal. They don't.
Costs assume a fixed token count and exclude retries, over-long outputs and human polishing time.
Second-placed Gemini 3.6 Flash costs a quarter of what third-placed GPT-5.6 Sol costs per chapter, and scores 0.18 higher. Eighth-placed Claude Opus 5 costs 60× what sixth-placed Hy3 costs, and scores 0.34 lower. Rank and price are all but unrelated.
Every model has AI tells
Six of the ten score above 5.5 on AI tells (lower is better). Several reviewers independently named the same handful of sentence patterns. The detector below counts exactly the patterns they named — paste in some Chinese prose of your own and see. (It matches Chinese text only.)
The sample is a passage written for this demo, not an entry. The detector does literal matching, not model attribution — reviewers judged “AI tells” from the reading experience as a whole, and this only highlights the traces they kept naming.
Characters who know too much
Having read the outline, the model lets a character say things he has no way of knowing, and his behaviour stops making sense. Expert and senior reviewers flagged this independently, again and again, as the structural problem that hurt the reading experience most.
Openings and endings give it away
Several reviewers noticed the same thing: the middle holds up, but the scene-setting at the start and the summing-up flourish at the end are where the machine shows.
Missing where the cheat power goes
The beat sheet called for the protagonist's cheat power to trigger at the end. Several models treated it as ordinary plot movement, not registering that this is where web fiction plants its hook.
Not even first place escapes
Gemini 3.1 Pro takes first on four of the five axes but only 4th on AI tells. The best-written entry is not the one that sounds least like a machine.
So which one should I use
Answer two questions and you'll get the matching recommendation from section eight.
1. Cost sensitivity
2. After generation
Pick both and the matching model and reasoning appear here.
These costs are API calls only. Reviewers said of the top tier that “a light polish and it's publishable”, and of the bottom that it is “completely unsuited to web fiction”. Price the polishing hours in and the real total-cost gap between models runs opposite to the API-price gap.
In the reviewers' words
With thanks to the SoloEnt community reviewers, who patiently wrote 148 comments in all. Everything below is as submitted during scoring, unedited; tap any comment to read the chapter it is about.
When the blind review closed, Fable 5 and Opus 5 had scored far below what we expected, and that started an argument.
The beat sheet we handed out was strict, and it was not perfect: an AI wrote it, and it pinned down more detail than it needed to.
Fable 5 and Opus 5 followed the beat sheet too obediently, which is why the prose came out badly — and which says something about how rigid these models are.
This is exactly what strong coding models look like on creative writing: they lack the spark the work demands.
All 30 texts, in the open
This is exactly what the reviewers read. Open any entry for the full text, its six scores and every word the reviewers wrote about it.
Method and limits
How the scoring worked, and what this result cannot be used to claim.
How the data was handled
Double-blind
Reviewers saw an anonymous code and nothing else — no model names, no one else's scores, no view of anyone's progress but their own. The code-to-model mapping existed only on the admin side.
Outlier removal
Judged column by column. A score is dropped only if it is both more than 2.5 × MAD from that entry's medianandmore than 2 points away in absolute terms. 28 of the 30 entries triggered at least one removal.
Reviewer screening
Every reviewer's scores were checked for internal consistency; those showing no discriminating power were excluded wholesale and counted in no statistic. Eight sets were excluded this time.
Sample size
Three versions per model, 3 — 5 reviewer ratings per entry, plus one internal expert rating.
Where this result stops
The sample is small
Three chapters per model. The gaps between the three top-tier models are not statistically significant, hence tiers rather than exact places.
Don't quote a single axis on its own
Some ratings show little separation on individual axes, which weakens the per-axis breakdown. The overall ranking is unaffected.
Only chapter one was tested
Every judgement rests on a 3,000-character opening. Pacing over a book, paying off setups and character arcs all need more length than that, and are not covered here.
“AI tells” is the most subjective axis
Scores on this axis run high and scatter widely; reviewers do not necessarily agree on what counts as a machine trace. It is also the only axis that reorders the field, so read it with particular care.
Appendix: the outline all ten models were given
Positioning
Journey to the West: Dragon Palace Graveyard, the Long Game of Survival — dark Journey to the West / play-it-safe survival / court intrigue / mortal-tier progression.
Logline: “In this man-eating Journey to the West, I, Guiyuan, do not aspire to Buddhahood. I only intend to outlive every one of you, scatter your ashes, and inherit what you leave behind.”
Chapter one: the Dragon Palace death row, jailer No. 9527
Guiyuan, transmigrated into the body of the Turtle Prime Minister's illegitimate son, wakes in a cold, damp death row, is bullied by his superior Third Crab and forced to deal with the corpse of a lethally venomous sea snake. Three scenes:
- Waking in the death row(oppressive / cold): what seeps from the walls is not water but black death-vapour. Cold, wet and dark are to be foregrounded; no long interior monologue — the inner state shows through what he notices around him.
- The superior's bullying(conflict / contrast): the original owner of the body was both insecure and proud; the new occupant kneels and grovels, calling himself “this lowly one”. Playing weak, the better to kill in one strike later.
- Down into the depths(atmosphere / setup): through the corridor of raving prisoners to Block D and the sea-snake corpse leaking green venomous fog.
Closing hook: the moment his fingertip touches a scale, a cold current bores into him and a grinding rumble sounds in his skull. This is where the beat sheet called for the cheat power to trigger — and where several models dropped it.
Target length was about 3,000 characters. One model wrote 5,500.
Thanks to this issue's reviewers
武行悟道alikezklastworm流萤King长庚笑沧桑海獭慕容复国无语在天子家喝茶瑞天阁主zsqSandra
Our review panel is still recruiting people who read a lot of web fiction — come join us.
Join the panel →What should we benchmark next?
Tell us which genre, which model or which part of the process you want benchmarked next.