TL;DR: Two of Three Category Leaders Are Already in Your Hands
This week, while wrapping up the漫剧 pipeline, I ran a full inventory: every video, image, and speech model I can actually call on my production lines, weighted and ranked against September 2026 blind-test leaderboards. The audit wasn’t “what models exist on the market” — it was “what’s actually wired into my Lumenx pipeline + CPA gateway and ready to receive requests today.”
Of the three category leaders, two are already running in production: video line Wan3.0-video (Arena I2V Elo 1481, #3 overall, +53 over the pipeline’s original Wan2.7), image line GPT-Image-2.5 Sunburst (Arena T2I 1421, #1). The only action item is the speech line — recommend switching the main voice from CosyVoice to Qwen3-TTS-Flash.
One side finding deserves a callout: the previous day’s verdict was “DashScope video keys are all dead.” This time I tested all 15 keys in CPA one by one against the video-synthesis endpoint — 7 passed cleanly.所谓 “通道死了” is often just “I didn’t test them all.”
Why I Did This: A “Keys Are Dead” Misjudgment
Timeline: On the evening of Sept 20, during a Lumenx production test, all three video channels (DashScope/Vidu/Kling) errored out. That night’s conclusion, written into memory: “waiting for user to fix keys.” Two days later, I fired a minimal request at each of the 15 keys in the CPA config (1-second duration, X-DashScope-Async: enable) — 7 returned 200 + task_id immediately. The one I’d tested before happened to be the broken one.
Lesson learned: channel-level verdicts require exhaustive testing — sampling lies. So the first order of business this time wasn’t checking leaderboards; it was getting a clear picture of “what do I actually have” — the model directory in my toolchain (Lumenx’s model_catalog.json, 38 models across 8 families), the channel list in the gateway (CPA config.yaml, 90+ model names across 16 channels), the audio registry (CosyVoice/Qwen3 dual-family voices in tts.py), plus the per-key live tests. Only after the inventory was complete did ranking come into play.
Method: Multi-Source Cross-Referencing, Single-Engine Sources Get Discounted
Ranking rules, stated upfront, no pretense of objectivity:
- Primary weight to LMArena / Arena.ai blind-test Elo — millions of human blind votes, closer to “real aesthetic preference” than any technical metric;
- Secondary weight to Artificial Analysis Video Arena, SuperCLUE Chinese boards, and vendor six-dimensional reviews, for cross-validation and gap-filling;
- Data gaps honestly demoted: TTS has no mature Arena board; that section’s ranking is a features + reputation weighted judgment, explicitly flagged in the text;
- On the retrieval day, the Tavily engine had SSL issues throughout; data was cross-checked using Grok + 秘塔 dual engines, with every key number requiring at least one independent source.
Models from the same vendor can clash across boards (e.g., Kling 3.0 scores 1092 on the T2V-with-audio sub-board but noticeably higher on the main board口径). Where they clash, I kept the range rather than hard-coding a single number.
Video Line: Wan3.0’s +53 Isn’t an Anomaly
First, the Arena.ai blind-test board for image-to-video, third week of September:
| Weighted Rank | Model | Arena Elo (I2V) | My Status |
|---|---|---|---|
| 1 | Wan3.0-video | 1481 (#3), +53 over Wan2.7 | ✅ Just接入 and set as default this week |
| 2 | Seedance 2.0 | 1474 (#5); Seedance 2.5 has surged to 1477 (#4) | ✅ Channel live (mulerouter) |
| 3 | Kling 3.0 | Main board口径 ~1460; sub-board口径 1092 | ⚠️ SecretKey empty |
| 4 | Wan2.7-r2v/i2v | 1426–1428 (#10–11) | ✅ 7 keys tested live, original主力 |
| 5 | Vidu Q3 Pro | 1363 (#18) | ⚠️ Key 401 |
| 6 | PixVerse V5.6/C1 | Formerly #2 on AA board (early口径) | ✅ Channel live |
| 7 | happyhorse-1.0 | April “欢乐马” dominant record (Elo 1375, clear gap) | ✅ Pipeline default i2v |
| 8 | grok-imagine-video | 480p tier 1381 (early口径) | ⚠️ Gateway submission no response, on hold |
What does the top of the board look like: #1 MiniMax-H3 at 1494, #2 Gemini Omni 1.1 Flash at 1488, Wan3.0 at 1481 — only 7 points off, with a 57% win rate. And it only went live in late August, entering the board straight at #3. That 53-point gap falls exactly on my white-model anchor pipeline: Wan3.0 supports first_frame/last_frame with strong semantics (reference images and first/last frames are mutually exclusive — can’t be mixed), meaning programmatically rendered anchor frames can be fed in as “strict first frame / last frame,” rather than Wan2.7-r2v’s “multiple reference images for approximate anchoring.” Composition-locking capability upgrades from “approximate” to “strict.”
The scenario-specific selection口径 also交叉出来了: 智东西’s six-dimensional review of Seedance 2.0 (画质/运动/图像保持/语义/音频/文字) ranked first across all dimensions, scoring 3.31–3.70, with facial realism and precise action control beating Kling 3.0; Kling 3.0’s moat is native 4K60fps output and cinematic camera work. Mapped to my漫剧 scenarios: white-model anchoring主力 Wan3.0, close-up shots switch to Seedance, cinematic炫技 consider fixing Kling keys.
Image Line: The Current Default Is Already #1 — That’s Luck
The image line was the most effortless — after the audit, the current default config is already at the top:
| Weighted Rank | Model | Arena Elo (T2I) | My Status |
|---|---|---|---|
| 1 | GPT-Image-2.5 (Sunburst/Flare) | 1421 / 1399; editing board 1520 | ✅ .env default |
| 2 | Gemini-3-Pro-Image-Preview | 口径 1050–1420,互有胜负 with GPT-Image | Not in production chain |
| 3 | Qwen-Image-2.0/2.1 | 1029–1237; open-source line Qwen-Image-2512 is open-source #1 | ✅ Pipeline default is same-family wan2.7-image-pro |
| 4 | agnes-image-2.5 | Qwen Image Bench 57.53 (>Seedream 4.5/4.0) | ✅ CPA agnes channel |
| 5 | Seedream 4.0/4.5 | Mid-low tier on AA board | No direct connection |
| 6 | grok-imagine-image-lite | No board data | ✅ CPA has it, lightweight tier |
GPT-Image-2.5’s position is uncontested: T2I board Sunburst variant at 1421, leading the previous generation by ~40 points; the image editing board shows a 1520-vs-1461 gap, where multi-round editing fidelity and 4K support create a clear断层 in the editing scenario. Per新京报 reporting, Qwen-Image-2.0’s T2I score of 1029 (global #3) comes from February’s AI Arena; post-April versions (2512/2.1) topped the open-source line but still trail GPT by a full tier overall.
The practical conclusion is one line: default stays put. Add a scenario-specific alternative — for Chinese posters, text-dense images (漫剧 covers / title cards), switch to Qwen-Image; its Chinese text rendering is a genuine edge in this scenario.
Speech Line: The Only Action Item Worth Acting On
Ugly truth first: TTS has no authoritative blind-test board. Artificial Analysis’s Speech Arena covers speech-to-speech end-to-end, not pure TTS scoring; figures like 97ms first-packet and 17 voices are vendor口径. This section’s ranking is a features + reputation weighted judgment, with overall evidence strength a full tier lower than the previous two lines. Please read with that discount applied:
| Rank | Model | Key Features | My Status |
|---|---|---|---|
| 1 | CosyVoice-v3.5-plus | Alibaba latest; voice cloning/design; same-family Qwen-Audio-3.0-TTS-Plus曾登 Speech Arena榜首 | ✅ tts.py default |
| 2 | Qwen3-TTS-Flash | 97ms first-packet, 17 voices, 9+ Chinese dialects, cloning/design, Dual-Track streaming | ✅ tts.py registered |
| 3 | Gemini-3.1-Flash-TTS | Controllable narration tags, low latency, multilingual | ✅ CPA aistudio-gemini |
| 4 | CosyVoice v2/v3-flash | 龙系 dozens of voices, older generation | ✅ |
Ranks 1 and 2 are actually the same vendor’s前后代 — the difference isn’t quality but scenario fit: Qwen3-TTS’s dialect coverage (9+ dialects) and instruction-based voice adjustment are more useful for character voice differentiation in古风漫剧, and 97ms first-packet is friendly for streaming synthesis. Switching cost is nearly zero — both families are already registered in tts.py, and the voice table (Cherry/Serena/Ethan/Chelsie — that batch of Qwen3 voices) is sitting right there.
Limitations & Reproduction
Three caveats, written here to prevent future self-deception:
- Elo is a living number. Arena weekly boards fluctuate ±10–20 points; Seedance 2.5 just pushed past 2.0 two weeks ago. The numbers in this post are fixed to the 2026-09-22 snapshot — re-verify after a month.
- TTS evidence strength is a tier lower, as flagged within the section.
- The “My Status” column is based on per-key live testing or source code verification. The 15-key endpoint tests used minimal requests (duration=1s error confirmed the auth and billing链路 is alive; actual output is a separate matter). A channel being “present” doesn’t mean “working” — Kling’s empty SecretKey and Vidu’s 401 are both live/dead checks one command away from verifying.
Reproduction path: model inventory comes from /config/model_catalog/generated/model_catalog.json (38 models) plus CPA config.yaml channel section; blind-test boards at arena.ai/leaderboard/image-to-video and lmarena.ai/leaderboard/image-to-video; Chinese-side cross-reference via SuperCLUE video board and Arena.ai Chinese weekly report. The key live-test request template is fully documented in Lumenx’s phase0-verdict doc.
One final number to close on: this audit took under two hours, and the most valuable action among them was those 15 one-by-one curls — it rewrote “video channel is dead” into “8 keys dead, 7 keys alive.”
Data Sources: LMArena I2V Board · Arena.ai official tweet: Wan3.0 I2V #3, 1481 · Arena.ai Week 36 weekly report (NetEase repost) · 智东西: Seedance 2.0 six-dimensional全第一 review · CometAPI: GPT-Image-2.5 benchmark · 新京报: Qwen-Image-2.0 launch · Qwen3-TTS technical deep dive · 澎湃: 欢乐马 dominant record
