Featured image of post Alibaba Unveils Qwen4 Training Progress at Cloud Expo: Future Models to Reach 5-10 Trillion Parameters, Next-Gen Video Model Launching in November

Alibaba Unveils Qwen4 Training Progress at Cloud Expo: Future Models to Reach 5-10 Trillion Parameters, Next-Gen Video Model Launching in November

Alibaba reveals Qwen4 training underway with future versions scaling to 5-10T parameters; next-gen video generation model arrives in November.

Alibaba Unveils Qwen4 Training, Plans 5-10T Parameter Scale for Future Models;

Next-Gen Video Model Launching in November

Key Announcements at a Glance

Key Announcements at a Glance
Key Announcements at a Glance|News screenshot

  • Release Timeline: Qwen4 training is underway; next-generation video generation model launches November 2026
  • Scale & Versions: Qwen4.5, Qwen5 and beyond will expand to 5-10T parameters (5-10 trillion)
  • Inference Models: Qwen3.8-Max and Qwen3.8-Flash live with optimizations; Qwen3.8-Flash achieves 96% throughput increase per instance
  • Capability Benchmark: Qwen3.8-Max scores 45 points on Artificial Analysis, matching Claude and GPT-4o in the top tier
  • Device Deployment: Qwen Intelligence full-stack smartphone solution released, optimized for on-device Agent tasks

From Self-Training to Chip-Level Co-Design

At the Cloud Computing Week (Yunqi Conference), Alibaba showcased comprehensive progress in Recursive Self-Improvement (RSI) technology. The standout demonstration is Qwen3.8-Max achieving human-zero-intervention iteration, autonomously building training pipelines, constructing datasets, designing experiments, and identifying defects. Over one month, it executed 33 effective迭代 rounds.

In inference optimization, Qwen3.8-Max demonstrated cross-architecture adaptation by optimizing Qwen3.8-Flash’s inference framework on T-Head’s newly launched GPU without prior experience on the hardware. This required no manual intervention yet delivered near-doubling of throughput (96% increase)—a dramatic contrast to traditional model-hardware migrations that typically demand weeks of human tuning.

Chip-model co-design achieved similar automation: Qwen3.8-Max processed just one bus specification document, ranover 60 hours, invoked EDA tools more than 10,000 times, and reduced physical chip area by 42% across front-end, verification, and backend stages. This confirms RSI now spans training, inference, and silicon design.

Video & Multimodal: Supporting Director-Level Creation

Alibaba’s next video generation model, launching November, targets professional creators with four key improvements:

  • Longer Duration: Maintains consistency over extended sequences
  • Better Control: Precise specification of characters, scenes, and camera angles
  • Full Narratives: Shifts from single-shot generation to understanding complete story structure
  • Smarter Cognition: Models director-level creative reasoning, including shot sequencing and narrative logic

Alongside video, Alibaba also unveiled updates for speech, image, world, music, and multimodal models. ATH Tech VP and Taotian Group CTO Zheng Bo stated: a native unified multimodal model will emerge within three years, dissolving modality boundaries.

Device Deployment: AI-Powered Smartphones

Device Deployment: AI-Powered Smartphones
Device Deployment: AI-Powered Smartphones|News screenshot

Alibaba’s model deployment extends to endpoints. Qwen Intelligence, the smartphone-focused full-stack solution, provides an Agent platform enabling phones to reliably execute cross-app complex tasks—such as calendar retrieval, route planning, and meal booking in sequence.

Model VersionParametersKey CapabilityTraining Autonomy
Qwen3.8-MaxUndisclosed33 iterative rounds with zero human involvement, 45 Artificial Analysis scoreHigh (RSI + post-training)
Qwen3.8-FlashUndisclosed96% throughput increase, T-Head GPU compatibilityAutonomous inference optimization
Qwen4Training ongoingNew-architecture foundation, unknown scaleRSI early integration
Qwen4.5/Qwen55-10TFuture scaling target, multimodal foundationParameter expansion focus

Who Should Act Now?

  • Developers/Enterprises needing high-throughput inference should evaluate Qwen3.8-Flash; the 96% gain directly reduces per-query latency and infrastructure cost
  • Content Creators requiring long-form, director-grade video control should wait for November’s upload; current models lack narrative coherence at scale
  • Smartphone Manufacturers/App Developers targeting agent-like task automation should integrate Qwen Intelligence or engage Alibaba for device-level support

Final Thought

The industry is shifting from “parameter races” to “system co-evolution”—when models can autonomously design chips and optimize inference, the productivity leap dwarfs gains from parameter scaling alone. Alibaba’s RSI path, if sustained, may redefine large-model engineering paradigms globally.