Wan3.0 is the new generation of Alibaba's Wan video generation model family, now available on Alibaba Cloud Model Studio. It generates up to 30 seconds of video in a single pass, accepts text, images, audio, video, and — for the first time — documents (PPT, PDF, DOC, XLS, and more) as input, and delivers a major leap in real-world fidelity: lifelike, diverse human faces and precise reference-to-video consistency. API pricing starts at $0.05 per second for 480P output, with 720P and 1080P tiers available.
If you build with video generation APIs — for film and short-form content, advertising, design and creative work, or destination marketing — this release changes what a single API call can produce: not a clip, but a complete story.
| Spec | Details |
|---|---|
| Max single-generation duration | 30 seconds (with smart duration recommendation and video extension) |
| Input modalities | Text, image, audio, video, documents |
| Supported document formats | doc, xls, ppt, pdf, txt, key, pages, numbers, md |
| Document limits | Max 1 file or link, ≤100MB, ≤50 pages |
| Output resolutions | 480P, 720P, 1080P |
| API pricing | $0.05/s (480P) · $0.10/s (720P) · $0.20/s (1080P) |
| Video editing | Modify visuals, plot, and dialogue without full regeneration |
| Availability | Alibaba Cloud Model Studio |
Most AI video generation models produce clips of a few seconds — enough for a shot, not a story. Wan3.0 extends single-generation length to 30 seconds, and the difference is more than a bigger number. Longer duration means room for narrative pacing, continuous camera moves, and complex cinematic language like one-take sequences. Instead of generating a shot, you're telling a complete story.
Two supporting features make the longer canvas practical:
Video has become a universal way to communicate, and Wan3.0 is designed to be a universal expression layer for information — not just a text-to-video model. Alongside full multimodal reference across text, images, audio, and video, Wan3.0 is the first model in the family to read documents, spreadsheets, and web pages as structured input.
Supported input formats include doc, xls, ppt, pdf, txt, key, pages, numbers, and md, with a limit of one file or link per request, up to 100MB and 50 pages. Pair a document with a prompt, and everyday office material converts directly into video:
This works because of Wan3.0's prompt engineering and instruction-following capabilities: the model treats your prompt as the control center that interprets the input, orchestrates the content, and shapes the final expression — turning wildly different source material into coherent video.
Upload a café's brand introduction document, add one sentence, and get a complete brand film.
| Reference file | Prompt |
|---|---|
| Café brand introduction document | Turn this brand story into a warm-toned brand TVC. Make it emotionally resonant — the kind of film that makes people want to visit the café and stay a while. |
Video generation is shifting from content novelty to production tooling, and production demands more than sharp frames — it demands believability. Wan3.0 targets accurate reproduction of both the physical world and digital information.
Wan3.0's text-to-video mode is built for diverse, lifelike faces rather than the glossy, interchangeable look of typical AI characters. It renders facial features and skin detail with sharper sensitivity, keeps emotional expression natural and restrained, and links micro-expressions with body language. Even in group scenes, different characters carry distinct, nuanced emotions.
Consistency is the metric production teams care about most. In reference-to-video mode, Wan3.0 reproduces reference details precisely and holds them stable across key dimensions:
For digital information, Wan3.0 also raises the aesthetic bar: harmonious color, fluid information delivery, and charts, data, and interface content that carry real visual polish.
Video editing capabilities, first introduced in Wan2.7, carry forward into Wan3.0. You can modify visuals, plot, and dialogue on generated videos — including flexible story re-creation — without regenerating from scratch. Iteration becomes a revision pass, not a do-over.
During internal testing, we identified areas where Wan3.0 still has room to grow: audio texture and on-screen text rendering accuracy are improving but not yet where we want them. Both are active areas of refinement for upcoming versions.
Wan3.0's upgrades are grounded in real industry needs, and the model is designed to slot back into real creative and production workflows.
The combination of 30-second duration, real-scene fidelity, and precise consistency improves the experience of creating AI films, short dramas and animated series, music and dance videos, and lifestyle vlogs — helping creators finish work at dramatically lower cost.
Advertising demands high quality and deep customization. Wan3.0's everything-to-video capability breaks past static image-and-copy formats, letting creative concepts render directly as film. Across consumer electronics, beauty, automotive, FMCG, 3C, apparel, and software, Wan3.0 covers the full creative range from product showcase to brand storytelling.
For UI interaction demos, software feature animations, and data visualization, Wan3.0 preserves the aesthetics and texture of the original design while elevating it into motion that tells a story — skipping the tedious animation production step so ideas are seen immediately.
With Wan3.0's real-scene reproduction, destination films no longer require large-scale location shoots. Natural landscapes, cultural landmarks, and local cuisine — city image films and digitized cultural scenes can be produced at low cost and high speed, bringing faraway places to the screen.
Wan3.0 video generation is billed per second of generated video, by output resolution:
| Resolution | Price per second |
|---|---|
| 480P | $0.05 |
| 720P | $0.10 |
| 1080P | $0.20 |
At these rates, a full 30-second 1080P generation costs $6.00**, and a 30-second 480P draft costs **$1.50 — practical for iterating on drafts at 480P and finishing at 1080P.
The Wan model family has shipped 8 iterations from Wan1.0 to Wan3.0, evolving from asset generation to full creation, and from content generation to commercial delivery. Every extension of generation length and every new reference modality has pushed the model closer to faithfully reproducing the real world.
Wan3.0 marks a new generation of that trajectory. Looking ahead, the same capabilities — understanding richer information and generating more complete, physically consistent output — point toward embodied AI, autonomous driving, and industrial simulation, where simulating and reconstructing the physical world will find new answers.
Wan3.0 is available on Alibaba Cloud Model Studio. You can review the model documentation, explore capabilities, and start generating today:
Wan3.0 is the latest generation of Alibaba's Wan video generation model family, available on Alibaba Cloud Model Studio. It generates up to 30-second videos in a single pass from text, images, audio, video, or documents, with lifelike human rendering and strong reference-to-video consistency.
Wan3.0 generates up to 30 seconds of video in a single generation. A smart duration feature recommends the right length based on your prompt, and video extension lets you continue a story beyond the initial generation.
Beyond text, image, audio, and video, Wan3.0 accepts documents in doc, xls, ppt, pdf, txt, key, pages, numbers, and md formats. Each request supports a maximum of 1 file or link, up to 100MB and 50 pages.
Wan3.0 is priced per second of generated video: $0.05/s (480P),** **$0.10/second at 720P, and $0.20/second at 1080P on Alibaba Cloud Model Studio.
Wan3.0 extends single-generation length to 30 seconds (with smart duration and extension), adds document input for "everything to video," and significantly upgrades realism — diverse lifelike faces and fine-grained reference consistency for characters, props, spaces, and style. The video editing capability introduced in Wan2.7 carries forward.
Wan3.0 supports modifying visuals, plot, and dialogue on generated videos, including flexible story re-creation, so you can iterate on content without regenerating from scratch.
Wan3.0 is available through Alibaba Cloud Model Studio. Visit the Wan3.0 model page to view documentation and start generating.
Audio texture and on-screen text rendering accuracy are still improving. Both are being actively refined and will continue to mature in subsequent versions.
From a single prompt to a full 30-second brand film — and from a PPT to a product ad — Wan3.0 turns any input into video. Try Wan3.0 on Alibaba Cloud Model Studio today.
To sharpen your results, explore our Wan3.0 Model Release page for more demos.
One Key, One CLI — Manage Your Alibaba Cloud Model Studio Token Plan from Terminal or AI Agent
1,501 posts | 510 followers
FollowAlibaba Cloud Community - August 7, 2026
Farruh - April 8, 2025
Alibaba Cloud Community - February 27, 2025
Alibaba Cloud Community - April 7, 2026
Alibaba Cloud Community - January 5, 2026
Alibaba Cloud Project Hub - August 19, 2025
1,501 posts | 510 followers
Follow
Alibaba Cloud Model Studio
A one-stop generative AI platform to build intelligent applications that understand your business, based on Qwen model series such as Qwen-Max and other popular models
Learn More
Token Plan
Build more, spend less. One plan, every modality.
Learn More
Qwen
Full-range, open-source, multimodal, and multi-functional
Learn More
Alibaba Cloud for Generative AI
Accelerate innovation with generative AI to create new business success
Learn MoreMore Posts by Alibaba Cloud Community