×
Community Blog Alibaba Unveils Wan3.0 with Twice as Long Video Outputs from a Richer Variety of Inputs

Alibaba Unveils Wan3.0 with Twice as Long Video Outputs from a Richer Variety of Inputs

Latest model supports multimodal reference inputs enabling static text-heavy data conversion into dynamic video contents up to 30 seconds

1

Alibaba has rolled out the beta version of Wan3.0, an advanced video generation model that supports 30-second video length and multimodal reference inputs. This innovative tool represents a significant advancement in generative AI, integrating high-fidelity video creation, precise visual consistency, and comprehensive multimodal inputs into a single model designed to streamline professional workflows. As Wan3.0 started public beta testing, users can apply for model testing on Alibaba Cloud's AI development platform Model Studio and AI-native cloud platform Qwen Cloud.

Mainstream AI video generators typically produce clips lasting only a few seconds to 15 seconds, which is the maximum clip duration of Wan2.7-Video, the preceding Wan video generation model. Wan3.0's native support for longer sequences up to 30 seconds per clip allows creators to execute complex camera movements and continuous, unbroken shots. The model also introduces an intelligent duration feature that recommends the optimal video length based on user prompts, alongside video extension features to help expand narrative timelines.


30-second video clips generated by Wan3.0

A major differentiator for Wan3.0 is its broad multimodal input capability. The model can process text, image, video, and audio inputs simultaneously, while also supporting web pages and documents like PDFs and PowerPoint presentations. This feature effectively allows users to convert static, text-heavy data directly into dynamic video content.

To address the visual drifting and distortion common in AI-generated media, Wan3.0 features high-precision visual continuity. The model renders highly realistic human faces with synchronized micro-expressions, produces natural multilingual voice outputs, and accurately present stable software user interfaces and motion graphics.

Beyond visual stability, Wan3.0 is designed to precisely replicate fine details from reference inputs, with strict control over character and product details, spatial relationships, and voice consistency. Instead of generating rough resemblances, it accurately replicates characters, props, audio, spatial layouts and styles from reference inputs while keeping layouts and audio stable. This precision is combined with natural movement and emotional expressions to turn standard AI clips into immersive, dramatic stories.

The model is designed to support a wide range of industries, from streamlining production for filmmakers, creating short dramas and social media content, to helping businesses easily turn text and images into marketing and educational videos. Additionally, Wan3.0 can serve as a powerful tool for tech developers by generating realistic simulation videos to train self-driving cars and robotics systems.

First introduced in July 2023, Alibaba's Wan series of visual generation models have undergone continuous upgrades to make image and video creation more realistic, easier to control, and more creator-friendly.


This article was originally published on Alizila written by Shao Xiaoyi and Karen Zhang.

0 0 0
Share on

Alibaba Cloud Community

1,493 posts | 508 followers

You may also like

Comments