Picking an AI model for novel writing, the internet is full of claims — “Claude has the best prose,” “DeepSeek has the densest foreshadowing,” “Kimi is in a league of its own for ultra-long-context continuation.” Which of these are backed by actual testing, and which are marketing? This article pulls together all publicly available raw benchmark data as of September 2026, evaluates models across the five dimensions that actually matter for novel writing, and gives you an actionable answer.
The weights are set based on the real demands of long-form serialized fiction: long-form generation 30% · long-context recall 25% · story logic and consistency 25% · literary quality/style 15% · general reasoning 5% (coding ability is excluded — it’s essentially irrelevant to writing novels).
The Bottom Line First: Overall Rankings
| Rank | Model | Overall Score | One-Line Rationale |
|---|---|---|---|
| 🥇 | Claude Opus 5 | 92 | #1 on the longform writing leaderboard (86.3 points, lowest slop at 5.6), creative writing Elo 2120, multi-needle long-context recall 93% @128K — the only model with no weak spots |
| 🥈 | GPT-6 Astra | 88 | Ceiling for general capability (GPQA 96), creative writing Elo 2163 tops the field (provisional sample), but longform output skews slop-heavy and repetitive |
| 🥉 | Claude Fable 5 / 5.1 | 87 | Style on par with Opus 5 (Elo 2152), #2 on EQ-Bench’s emotional intelligence leaderboard, but burns through quota fast |
| 4 | Kimi K3 | 82 | Strongest for Chinese: EQ creative writing 2070 (top non-US model), #3 on the EQ4 emotional leaderboard, unmatched reputation for 1M-context continuation |
| 5 | GLM-5.3 | 81 | Dark horse: EQ creative writing 2064, longform 81.8 ties GPT-5.6, 0.992 on the AIME 2026 leaderboard |
| 6 | GPT-5.6 Sol | 79 | Strong writing but heavy “AI flavor” (slop 16.9) — a widely acknowledged readability weakness in the community |
| 7 | DeepSeek V4 Pro | 78 | 384K single-pass output, longest in the field + best-in-class reputation for Chinese foreshadowing; weak on overly dense prose (slop 19.7) |
| 8 | Muse Spark 1.3 (Meta) | 77 | King of cost-performance, longform 82.8 ties GPT-6 Astra |
| 9 | Qwen3.8-Max | 76 | #1 on LongBench v2 long-document understanding (66.3), MRCR 8-needle @256K 92.9 |
| 10 | Gemini 3.1 Pro | 74 | Meticulous worldbuilding, but recall drops 50 points past 128K, and its Chinese prose feel is weak |
Overall scores are normalized estimates from five-dimension evidence at the stated weights — useful for ranking, not precise measurements. Entries marked * are provisional EQ-Bench samples and their rankings may drift.
The one-line answer: if budget allows, go with Claude Opus 5. For mass-production Chinese serialized fiction, a combination of Kimi K3 + DeepSeek V4 Pro + GLM-5.3 gets you 80% of the effect at a fraction of the cost.
Dimension 1: Long-Form Generation (Weight 30%)
1.1 EQ-Bench Creative Writing Longform (The Most Relevant “Chapter Writing” Leaderboard)
This is currently the only public leaderboard dedicated to “writing long chapters” (judged by Claude Sonnet 4.6, scored out of 100, with deductions for slop — “AI-flavored prose” — and repetition). Data source: creative_writing_longform.js, snapshot dated 2026-09-07, covering 134 models.
| Rank | Model | Total Score | Avg Chapter Length (tokens) | Slop↓ | Repetition↓ |
|---|---|---|---|---|---|
| 1 | claude-opus-5 | 86.3 | 6264 | 5.64 | 5.0 |
| 2 | claude-fable-5-1 * | 85.3 | 5777 | 7.59 | 5.4 |
| 3 | claude-fable-5 | 83.0 | 6295 | 8.31 | 4.4 |
| 4 | gpt-6-astra * | 82.8 | 5845 | 9.07 | 6.3 |
| 4 | muse-spark-1.3 * | 82.8 | 6253 | 10.73 | 4.8 |
| 6 | claude-opus-4-7 | 81.8 | 5552 | 9.06 | 4.6 |
| 6 | GLM-5.3 * | 81.8 | 5928 | 7.09 | 4.6 |
| 8 | gpt-5.6-sol | 81.7 | 6881 | 11.98 | 6.3 |
| 9 | muse-spark-1.2 * | 81.5 | 7258 | 11.43 | 4.5 |
| 10 | claude-opus-4-8 | 80.8 | 5460 | 9.39 | 3.8 |
| 11 | claude-sonnet-4-6 * | 79.9 | 6893 | 10.56 | 5.5 |
| 11 | ox-alpha * (stealth model) | 79.9 | 6302 | 7.43 | 4.0 |
| 13 | kimi-k3 | 79.6 | 7296 | 9.67 | 4.9 |
| 14 | Kimi-K2.6 | 78.5 | 6649 | 18.92 | 4.6 |
| 15 | gpt-5.4 | 78.3 | 8192 | 12.45 | 4.8 |
| 16 | claude-sonnet-5 | 78.3 | 5138 | 13.53 | 5.6 |
| 17 | gpt-5.5 | 78.2 | 8812 | 16.89 | 5.3 |
| 18 | gpt-5.6-terra | 78.0 | 7482 | 15.41 | 7.4 |
| 19 | GLM-5.2 | 77.9 | 5316 | 16.51 | 4.5 |
| 20 | claude-opus-4-6 * | 77.7 | 6189 | 16.86 | 4.7 |
| 21 | gemini-3.8-flash * | 76.8 | 7061 | 27.69 | 5.2 |
| 22 | DeepSeek-V4-Pro | 75.6 | 7516 | 21.83 | 3.9 |
| 24 | Kimi-K2.5 | 74.9 | 6315 | 17.28 | 4.8 |
| 27 | GLM-5.1 | 73.5 | 6271 | 24.94 | 4.4 |
| 31 | GLM-5 | 70.9 | 8065 | 24.04 | 4.4 |
| 35 | grok-4.20-beta | 68.5 | 8370 | 27.41 | 5.2 |
| 37 | gemini-3.1-pro-preview | 68.2 | 7433 | 38.12 | 4.9 |
| 40 | DeepSeek-V3.2 | 66.8 | 6496 | 41.24 | 4.9 |
| 41 | Qwen3-Max-2025-09-24 | 66.2 | 4403 | 49.19 | 7.7 |
| 42 | DeepSeek-V4-Flash | 66.0 | 5505 | 19.59 | 7.0 |
| 44 | DeepSeek-V4-Flash-0731 | 61.2 | 6221 | 31.44 | 7.6 |
| 46 | DeepSeek-R1 | 59.5 | 4035 | 56.44 | 6.4 |
| 51 | Qwen3.8-27B * | 53.3 | 6293 | 46.34 | 8.9 |
Note who’s NOT on the leaderboard: Qwen3.8-Max, Qwen3.8-2.4T-A95B, and grok-4.5/4.6 did not participate in the longform evaluation — any longform scores you see online for these models are fabricated.
1.2 Single-Pass Output Token Limit (Determines Whether a Chapter Can Be Written in One Go)
| Model | Max Output | Context Window |
|---|---|---|
| DeepSeek V4 family (Pro/Flash/V4.1) | 384K (longest in the field) | 1M |
| Claude Opus 5 / Fable 5.1 / Sonnet 5 | 128K | 1M |
| GPT-6 Astra | 128K | 1.05M |
| GLM-5.2 / 5.3 | 128K | 1M |
| Qwen3.8-2.4T-A95B (open weights) | 262K reasoning / 131K body text (official recommended config) | 262K native, extendable to 1M |
| Kimi K3 | Not officially disclosed; community reputation is “seamless continuation with 2 million Chinese characters of prior text fed in” | 1M |
DeepSeek’s 384K output limit means it can theoretically write 200,000 Chinese characters in a single pass — you probably won’t use it that way, but being able to write a complete 10-20K-character chapter without truncation and continuation stitching is a real, tangible advantage.
Dimension 2: Literary Quality / Style (Weight 15%)
2.1 EQ-Bench Creative Writing v3 (Main Creative Writing Leaderboard, Elo-Based)
Snapshot dated 2026-09-07, 133 models. * = provisional sample.
| Rank | Model | Elo | Writing Score /20 | Slop↓ | Repetition↓ |
|---|---|---|---|---|---|
| 1 | gpt-6-astra * | 2163.9 | 16.80 | 8.41 | 3.54 |
| 2 | claude-fable-5-1 * | 2152.7 | 16.95 | 8.16 | 3.64 |
| 3 | claude-opus-5 | 2120.6 | 17.07 | 6.59 | 4.31 |
| 4 | kimi-k3 | 2070.6 | 16.85 | 9.70 | 3.73 |
| 5 | GLM-5.3 * | 2064.1 | 17.04 | 8.42 | 3.23 |
| 6 | gpt-5.6-sol | 1963.4 | 16.78 | 11.68 | 3.41 |
| 7 | ox-alpha * | 1960.6 | 16.89 | 9.83 | 3.05 |
| 8 | claude-fable-5 | 1934.6 | 16.81 | 10.28 | 3.92 |
| 9 | muse-spark-1.1 | 1916.1 | 16.54 | 12.11 | 3.49 |
| 10 | claude-opus-4-7 | 1907.1 | 16.57 | 11.09 | 4.00 |
| 11 | muse-spark-1.3 * | 1905.5 | 16.70 | 10.73 | 3.66 |
| 12 | gpt-5.6-terra | 1850.3 | 16.56 | 12.40 | 3.02 |
| 13 | gpt-5.5 | 1843.5 | 17.01 | 13.10 | 2.48 |
| 14 | Qwen3.8-2.4T-A95B * | 1840.9 | 16.72 | 12.26 | 3.50 |
| 15 | gpt-5.4 | 1835.6 | 16.89 | 12.20 | 2.71 |
| 16 | claude-opus-4-8 | 1835.2 | 16.66 | 13.16 | 3.68 |
| 17 | muse-spark-1.2 * | 1835.2 | 16.44 | 12.98 | 3.31 |
| 18 | gpt-5.6-luna | 1825.8 | 16.58 | 11.80 | 4.00 |
| 19 | claude-sonnet-4-6 | 1804.2 | 16.50 | 9.90 | 4.06 |
| 23 | GLM-5.2 | 1752.8 | 16.44 | 13.11 | 3.94 |
| 24 | gemini-3.8-flash * | 1749.7 | 16.54 | 22.60 | 3.12 |
| 25 | Kimi-K2.6 | 1721.3 | 16.67 | 13.30 | 3.77 |
| 26 | Qwen3.8-27B * | 1668.4 | 15.50 | 12.12 | 4.19 |
| 27 | Kimi-K2-Instruct | 1662.7 | 16.40 | 15.51 | 3.40 |
| 32 | GLM-5.1 | 1589.2 | 16.26 | 23.56 | 3.76 |
| 33 | grok-4.5 | 1576.0 | 16.25 | 17.73 | 3.52 |
| 35 | grok-4.20-beta | 1570.7 | 14.51 | 15.85 | 4.50 |
| 36 | DeepSeek-V4-Flash | 1555.7 | 16.29 | 20.92 | 4.30 |
| 37 | DeepSeek-V4-Pro | 1552.1 | 16.45 | 19.66 | 3.21 |
| 39 | DeepSeek-V3.2 | 1511.2 | 16.28 | 23.19 | 4.06 |
| 40 | DeepSeek-R1 (baseline) | 1500.0 | 15.68 | 31.21 | 4.62 |
| 48 | DeepSeek-V4-Flash-0731 | 1438.4 | 15.98 | 24.59 | 5.45 |
| 53 | gemini-2.5-pro-preview-06-05 | 1418.9 | 16.16 | 28.73 | 4.90 |
Notably absent (important): qwen3.8-max itself, Qwen3-Max, grok-4.6, and grok-4.3 did not participate in this leaderboard.
Key readings: GLM-5.3 (2064) sits a full 311 Elo above GLM-5.2 (1752) — a gap created purely by post-training on the same base model; Zhipu really invested in this generation’s post-training. Kimi K3 (2070) is the only non-US model in the top 10. The entire DeepSeek V4 family scores low on style (the 1550 tier), corroborating the community criticism of “overly dense prose” — on EQ-Bench’s radar chart, V4-Pro’s weakest dimension is precisely “Avoids Purple Prose,” while its strongest is “Creativity.”
2.2 LMArena Creative Writing Category (Blind Human Voting, 2026-09-11)
The top 12 is dominated by Claude and Gemini: 1. claude-fable-5 · 2. claude-opus-4-6-high · 3-4. gemini-3.7/3.8-flash-high · 5. claude-opus-4-7-high · 6. claude-fable-5.1-max · 7. gemini-3-pro · … Highest-ranked Chinese models: glm-5.3-max #17, qwen3.8-max #18, kimi-k3-max #25.
Note the divergence between Arena and EQ-Bench: Gemini Flash’s high-thinking tier does very well in Arena’s quick blind votes, but on EQ-Bench’s careful evaluation its slop score hits 22.6 — charming in short chats ≠ readable in longform.
2.3 Chinese “AI Flavor” Deep Dive (Community-Tested Consensus)
Conclusions cross-validated across four independent Chinese-language sources:
- Lightest AI flavor: Doubao (colloquial/internet-savvy tone), Claude (literary quality)
- Greatest depth: DeepSeek, Kimi
- No model produces Chinese output that escapes the need for polish — the difference is only in how heavy the baseline AI flavor is
Dimension 3: Long-Context Recall (Weight 25%)
Writing serialized longform fiction, the model must remember the foreshadowing you planted 50 chapters ago.
3.1 MRCR v2 Multi-Needle Recall (By OpenAI, the Closest to “Cross-Referencing a Story Bible”)
| Scenario | Model | Score |
|---|---|---|
| 8-needle @128K | Claude Opus 4.6 | 93.0% (leading) |
| 8-needle @128K | Claude Sonnet 4.6 / Gemini 3.1 Pro | 84.9% |
| 8-needle @128K | GPT-5.5 | 74.0% |
| 8-needle @1M | Claude Opus 4.6 | 76-78% |
| 8-needle @1M | Gemini 3 Pro | 26.3% (a 50-point cliff from 128K→1M) |
| 4-needle @256K | GPT-5.2 | 98% (first near-perfect score at this scale) |
| 8-needle @256K | Qwen3.8-Max | 92.9% (official figures) |
| 8-needle @256K | Qwen3.7-Max | 86.7% |
| MRCR 1M | DeepSeek-V4-Pro-Max | 83.5% (official) |
3.2 Fiction.LiveBench (120K-Token Novel Reading Comprehension — the Closest Match to “Remembering the First 100 Chapters”)
Snapshot from 2026-05 (predates the H2 2026 flagships; reference the landscape only):
| Rank | Model | Score |
|---|---|---|
| 1 | o3 | 100.0 |
| 2 | GPT-5.2 | 96.9 |
| 2 | Grok 4 | 96.9 |
| 4 | Gemini 2.5 Pro | 90.6 |
| 5 | Qwen3 235B | 68.8 |
| 10 | MiniMax M2 | 59.4 |
| 13 | Claude 3.7 Sonnet | 53.1 |
| 16 | Kimi K2 Instruct | 40.6 |
| 19 | Claude Opus 4.5 | 37.5 |
| 21 | DeepSeek R1 | 33.3 |
⚠️ Note: Claude models score abnormally low on this leaderboard (likely because reasoning-model methodology is favored), which contradicts MRCR’s conclusions. Reading both leaderboards together, the takeaway is: OpenAI/xAI are strong at “locating and recalling novel text,” while Claude is strong at “multi-thread cross-referencing recall.” Also, this leaderboard’s official site is currently inaccessible, and no H2 2026 flagships are listed.
3.3 LongBench v2 (8K–2M Character Long-Document Understanding)
| Rank | Model | Score |
|---|---|---|
| 1 | Qwen3.8-Max | 66.3 (official + third-party agree) |
| 2 | Claude Opus 4.5 | 64.4 |
| 3 | Qwen3.5-397B-A17B | 63.2 |
| — | DeepSeek-V4-Pro (official self-reported) | 51.5 |
| — | DeepSeek-V4.1-Flash (official self-reported) | 45.2 |
3.4 Claimed vs. Actually Reliable Context by Vendor (2026 Real-World Consensus)
| Model | Claimed | Actually Reliable Range |
|---|---|---|
| Claude Opus 5 / Fable | 1M | ~200-400K |
| GPT-6 Astra | 1.05M | Single-needle to 1M at 96% |
| Gemini 3.1 Pro | 1M (3 Pro claims 10M) | Excellent ≤128K, cliff past 256K |
| DeepSeek V4 family | 1M | Competitive single-needle @1M, discounted on downstream tasks |
| Grok 4 Fast | 2M | Few independent tests |
| Kimi K3 | 1M | Community-tested reputation for ultra-long continuation is in a league of its own |
Industry consensus: 50-70% of the claimed context is the reliable range. For chapters under 100K characters plus a story bible, every top-tier model suffices; once total context exceeds 300K characters, Claude (multi-needle) and GPT/Grok (single-needle localization) are the most stable.
Dimension 4: Story Logic and Consistency (Weight 25%)
There is no public third-party quantitative leaderboard for this dimension — every claim online about “staying coherent over a million characters” or “character consistency” comes from writing platforms’ marketing copy (one platform claims “1.3 setting contradictions per 300K characters vs. ChatGPT’s 4.7” — no third-party verification, not credible). Reliable evidence comes from same-prompt community testing in the Chinese-speaking world:
Chinese Community Same-Prompt Test (2026-09, a Traditional-Chinese blogger ran 12 models on the same web-fiction outline)
| Model | Test Conclusion |
|---|---|
| Claude Opus 4.8 Max / 4.7 | Strongest at foreshadowing weave and full-chapter layout; 4.7 produces the most natural prose |
| DeepSeek V4 | Misses not a single foreshadowing thread, but “like espresso — lacks breathing room” |
| Doubao | Chapter titles show the best Chinese pun aesthetics, thinks for only 16 seconds, but the plot tends to drift after 30K characters |
| GPT-5.5 Pro | Research-tier 7-minute thinking actually over-interprets instructions |
| Gemini 3.x | Meticulous worldbuilding, comes with built-in conceptual analysis |
| Grok 4.x | The only model to pull off text + images in one pass, but its Chinese prose is mediocre |
Division-of-Labor Consensus Among Chinese Web Fiction Authors (Tahou platform’s 10-model test, 2026-09)
Claude handles prose polish and emotional scenes, DeepSeek handles logic deduction and setting error-checking, Kimi handles ultra-long prior-text continuation, Doubao handles fragmentary inspiration.
Official Writing Data (DeepSeek V4 Technical Report §5.4.1, arXiv 2606.19348)
The only vendor to publish official data treating “Chinese writing” as a core scenario:
| Comparison | Instruction-Following Win Rate | Writing Quality Win Rate |
|---|---|---|
| DS V4-Pro vs Gemini-3.1-Pro (creative writing) | 60.0% | 77.5% |
| DS V4-Pro vs Claude Opus 4.5 (complex multi-turn constrained writing) | 45.9% vs 52.0% | — |
DeepSeek itself officially acknowledges that Claude Opus still leads on the hardest constrained-writing tasks. This triangulates with the leaderboard data and community reputation.
Dimension 5: General Reasoning (Weight 5%)
Almost no differentiating power for novel writing, so a quick table suffices:
| Benchmark | #1 | Worth Noting |
|---|---|---|
| GPQA Diamond | GPT-6 Astra 96.0 | Saturated — 24 models ≥90% |
| AIME 2026 | GLM-5.2 family 0.992 | Note: all 26 models are vendor self-reported, 0 independently verified |
| MathArena Competition Overall | GPT-6 Astra 90.7% by a wide margin | Claude Opus 5 second at 72.1% |
| HLE | Claude Fable 5.1 0.650 | DeepSeek-V4-Pro-0813 0.600 (open-source #1) |
| Artificial Analysis Intelligence Index | Fable 5.1 = GPT-6 Astra = 53 | GLM-5.3 45, Kimi K3 44, Qwen3.8-Max 40 |
| LMArena Overall Elo | Fable 5.1 = GPT-6 Astra ≈ 1520 | Kimi K3 1506, Qwen3.8-Max 1506, GLM-5.3 1505, DeepSeek-V4.1-Flash 1503 |
Key Model Clarifications: Three Version Numbers You’ll Probably Confuse
❌ “DeepSeek-V4-0831” Does Not Exist
The 0831 date belongs to DeepSeek-V4-Flash-Vision-Exp (a multimodal experimental version) — it was announced on August 21, but its HuggingFace weights repository was only created on August 31; “0831” is the repository creation timestamp, not a text-model snapshot. DeepSeek’s official changelog lists only two August releases: 08-13 (V4-Pro GA) and 08-21 (Vision-Exp).
✅ The Real DeepSeek V4 Version Timeline
| Version | Release Date | Positioning |
|---|---|---|
| V4 / V4-Pro / V4-Flash (preview) | 2026-04-24 | Initial preview |
| V4-Flash-0731 | 2026-07-31 | Flash official release |
| V4-Pro-0813 | 2026-08-13 | Pro official release (the one with HLE 0.600, open-source #1) |
| V4-Flash-Vision-Exp | Announced 2026-08-21 | Multimodal experiment |
| V4.1-Flash | 2026-09-10 (yesterday) | New-architecture Flash, GPQA 90.9 / officially claimed to surpass V4 Pro across the board |
⚠️ Important Timing: V4 Pro Will Be Downgraded on September 14
Announcement on DeepSeek’s official pricing page: starting 2026-09-14 12:00 (Beijing time), API requests to deepseek-v4-pro will be routed to V4.1-Flash and billed at Flash pricing, on the grounds that V4.1-Flash has comprehensively surpassed V4 Pro. V4.1 Pro has not yet been released. If you’re currently using v4-pro, watch for output behavior changes in three days.
❌ “qwen3.8-max-0902” Does Not Exist
In Alibaba Cloud’s official documentation, dated snapshots only go up to qwen3.7-max-2026-05-20 / qwen3.7-max-2026-06-08; qwen3.8-max is a rolling alias with no date suffix. Its open-source counterpart is Qwen3.8-2.4T-A95B (in HF’s official words: “Qwen3.8-Max is the official version based on Qwen3.8-2.4T-A95B”). Any scores online citing “0902” are unverifiable.
Qwen3.8 / GLM-5.x Version Timelines
| Model | Released | Context | Output Limit | Notes |
|---|---|---|---|---|
| qwen3.8-max | 2026-08 | 1M (262K native) | 262K reasoning / 131K body text (recommended) | = official closed-source version of 2.4T-A95B + vision + non-thinking mode |
| Qwen3.8-2.4T-A95B | 2026-08-12 | 262K native, extendable to 1M | Same as above | Open weights, 2.4T total / 95B active MoE |
| Qwen3.8-27B | 2026-08-14 | 262K, extendable to 1M | Same as above | 27B dense pocket cannon, natively multimodal |
| GLM-5.2 | 2026-06-16 | 1M | 128K | 744B-A40B MoE, MIT open-source |
| GLM-5.3 | 2026-08-25 | 1M | 128K | Same base as 5.2, gains purely from post-training |
| GLM-5.3-Flash | End of 2026-08 | 1M | 128K | New 320B/18B base, first natively multimodal GLM-5, priced at roughly 1/10 of 5.3 |
(Note: glm-5.2-fast-preview cannot be found in Zhipu’s official documentation; it’s suspected to be a temporary alias used by intermediary relay channels and cannot be verified against official sources.)
Final Selection Recommendations
Single-Model Champion: Claude Opus 5
Ranked #1 in longform, top 3 in style, strongest multi-needle recall, ceiling-level reputation in the Chinese community — no weaknesses under the five-dimension weights. If budget allows, use it for your key chapters.
Multi-Model Division of Labor by Scenario (Practical Setup for Chinese Serialized Longform)
| Stage | Recommended Model | Rationale |
|---|---|---|
| Outline / worldbuilding brainstorm | Gemini 3.1 Pro / GLM-5.3 | Meticulous settings, built-in conceptual analysis |
| Key chapter drafting | Claude Opus 5 | Longform 86.3 points + lowest slop |
| Mass-production chapter drafting | Kimi K3 / GLM-5.3 | EQ 2070/2064, far cheaper than Claude |
| One-shot ultra-long single chapter | DeepSeek V4 Pro / V4.1-Flash | 384K output ceiling |
| Setting audit / foreshadowing check | DeepSeek V4 Pro | #1 reputation for foreshadowing density; its dense prose happens to suit audit work |
| Ultra-long prior-text continuation | Kimi K3 | Unmatched reputation for continuation with 2 million characters of prior text fed in |
| Polish to reduce AI flavor | Claude Opus 5 / Kimi K3 | The lightest-AI-flavor tier |
Three Pitfall Warnings
- There is no credible third-party data for “staying coherent over a million characters.” The only sources are writing platforms’ marketing copy. In practice, rely on a process of “story bible + a DeepSeek audit every N chapters” as your safety net — don’t expect any model to remember everything on its own.
- Leaderboard scores marked * will drift (GPT-6 Astra, GLM-5.3, and ox-alpha are all provisional samples) — re-check EQ-Bench monthly.
- On September 14, DeepSeek V4 Pro gets downgraded to V4.1-Flash — if your pipeline has v4-pro hardcoded, test V4.1-Flash’s style differences in advance (its technical report contains no writing-specific data; it’s an unknown).
Data sources: eqbench.com (creative_writing.js v1.0.91 / longform v1.0.9, snapshot 2026-09-07), lmarena.ai, llm-stats.com, artificialanalysis.ai, matharena.ai, writingbench.hf.space, neosignal.io (Fiction.LiveBench mirror), api-docs.deepseek.com, huggingface.co/deepseek-ai | Qwen | zai-org, arXiv 2606.19348 / 2408.07055 / 2506.18841, maplefeather.com, tahou.com. All figures marked “official self-reported” are vendor claims that have not been independently verified; version numbers with no verifiable evidence are explicitly marked as non-existent.