
Alibaba has rolled out the beta version of Wan3.0, an advanced video generation model that supports 30-second video length and multimodal reference inputs. This innovative tool represents a significant advancement in generative AI, integrating high-fidelity video creation, precise visual consistency, and comprehensive multimodal inputs into a single model designed to streamline professional workflows. As Wan3.0 started public beta testing, users can apply for model testing on Alibaba Cloud's AI development platform Model Studio and AI-native cloud platform Qwen Cloud.
Mainstream AI video generators typically produce clips lasting only a few seconds to 15 seconds, which is the maximum clip duration of Wan2.7-Video, the preceding Wan video generation model. Wan3.0's native support for longer sequences up to 30 seconds per clip allows creators to execute complex camera movements and continuous, unbroken shots. The model also introduces an intelligent duration feature that recommends the optimal video length based on user prompts, alongside video extension features to help expand narrative timelines.
30-second video clips generated by Wan3.0
A major differentiator for Wan3.0 is its broad multimodal input capability. The model can process text, image, video, and audio inputs simultaneously, while also supporting web pages and documents like PDFs and PowerPoint presentations. This feature effectively allows users to convert static, text-heavy data directly into dynamic video content.
To address the visual drifting and distortion common in AI-generated media, Wan3.0 features high-precision visual continuity. The model renders highly realistic human faces with synchronized micro-expressions, produces natural multilingual voice outputs, and accurately present stable software user interfaces and motion graphics.
Beyond visual stability, Wan3.0 is designed to precisely replicate fine details from reference inputs, with strict control over character and product details, spatial relationships, and voice consistency. Instead of generating rough resemblances, it accurately replicates characters, props, audio, spatial layouts and styles from reference inputs while keeping layouts and audio stable. This precision is combined with natural movement and emotional expressions to turn standard AI clips into immersive, dramatic stories.
The model is designed to support a wide range of industries, from streamlining production for filmmakers, creating short dramas and social media content, to helping businesses easily turn text and images into marketing and educational videos. Additionally, Wan3.0 can serve as a powerful tool for tech developers by generating realistic simulation videos to train self-driving cars and robotics systems.
First introduced in July 2023, Alibaba's Wan series of visual generation models have undergone continuous upgrades to make image and video creation more realistic, easier to control, and more creator-friendly.
This article was originally published on Alizila written by Shao Xiaoyi and Karen Zhang.
Introducing Qoder Code Security: Security From the First Line of Code
Alibaba Cloud Smart Studio Self-Service Edition Now Live Internationally
1,493 posts | 508 followers
FollowAlibaba Cloud Community - December 16, 2025
Alibaba Cloud Community - April 7, 2026
Alibaba Cloud Native Community - March 6, 2025
Alibaba Cloud Community - June 23, 2026
ApsaraDB - January 16, 2026
Alibaba Cloud Community - February 28, 2025
1,493 posts | 508 followers
Follow
Alibaba Cloud Model Studio
A one-stop generative AI platform to build intelligent applications that understand your business, based on Qwen model series such as Qwen-Max and other popular models
Learn More
Qwen
Full-range, open-source, multimodal, and multi-functional
Learn More
Token Plan
Build more, spend less. One plan, every modality.
Learn More
Alibaba Cloud for Generative AI
Accelerate innovation with generative AI to create new business success
Learn MoreMore Posts by Alibaba Cloud Community