×
Community Blog Wan3.0: 30-Second AI Video Generation from Any Input

Wan3.0: 30-Second AI Video Generation from Any Input

Wan3.0 is the new generation of Alibaba's Wan video generation model family, now available on Alibaba Cloud Model Studio.

Wan3.0 is the new generation of Alibaba's Wan video generation model family, now available on Alibaba Cloud Model Studio. It generates up to 30 seconds of video in a single pass, accepts text, images, audio, video, and — for the first time — documents (PPT, PDF, DOC, XLS, and more) as input, and delivers a major leap in real-world fidelity: lifelike, diverse human faces and precise reference-to-video consistency. API pricing starts at $0.05 per second for 480P output, with 720P and 1080P tiers available.

If you build with video generation APIs — for film and short-form content, advertising, design and creative work, or destination marketing — this release changes what a single API call can produce: not a clip, but a complete story.

Key Takeaways

  • Wan3.0 generates 30-second videos in a single generation — up from short clips — with smart duration recommendation and video extension.
  • "Everything to video": beyond text, image, audio, and video inputs, Wan3.0 reads documents — doc, xls, ppt, pdf, txt, key, pages, numbers, and md (1 file or link, ≤100MB, ≤50 pages).
  • Real-world fidelity: diverse, lifelike human faces (no more same-face AI people), natural micro-expressions, and stable consistency for characters, props, spaces, and style in reference-to-video mode.
  • Video editing built in: modify visuals, plot, and dialogue without regenerating from scratch.
  • API pricing: $0.05/second (480P), $0.10/second (720P), $0.20/second (1080P) — available on Alibaba Cloud Model Studio.

Wan3.0 Specs at a Glance

Spec Details
Max single-generation duration 30 seconds (with smart duration recommendation and video extension)
Input modalities Text, image, audio, video, documents
Supported document formats doc, xls, ppt, pdf, txt, key, pages, numbers, md
Document limits Max 1 file or link, ≤100MB, ≤50 pages
Output resolutions 480P, 720P, 1080P
API pricing $0.05/s (480P) · $0.10/s (720P) · $0.20/s (1080P)
Video editing Modify visuals, plot, and dialogue without full regeneration
Availability Alibaba Cloud Model Studio

30-Second Videos: From a Single Shot to a Complete Story

Most AI video generation models produce clips of a few seconds — enough for a shot, not a story. Wan3.0 extends single-generation length to 30 seconds, and the difference is more than a bigger number. Longer duration means room for narrative pacing, continuous camera moves, and complex cinematic language like one-take sequences. Instead of generating a shot, you're telling a complete story.

Two supporting features make the longer canvas practical:

  • Smart duration recommendation — the model analyzes your prompt and automatically recommends an appropriate video length, so a simple product beat doesn't get padded to 30 seconds and an ambitious narrative isn't cut short.
  • Video extension — extend an existing generation to continue the storyline, letting plots and ideas grow naturally across segments.

Everything to Video: Documents Become Films

Video has become a universal way to communicate, and Wan3.0 is designed to be a universal expression layer for information — not just a text-to-video model. Alongside full multimodal reference across text, images, audio, and video, Wan3.0 is the first model in the family to read documents, spreadsheets, and web pages as structured input.

Supported input formats include doc, xls, ppt, pdf, txt, key, pages, numbers, and md, with a limit of one file or link per request, up to 100MB and 50 pages. Pair a document with a prompt, and everyday office material converts directly into video:

  • A product PPT becomes a brand film or product launch ad.
  • A training deck becomes video courseware.
  • Spreadsheet data becomes animated, dynamic charts.
  • A business report becomes a narrated video briefing.

This works because of Wan3.0's prompt engineering and instruction-following capabilities: the model treats your prompt as the control center that interprets the input, orchestrates the content, and shapes the final expression — turning wildly different source material into coherent video.

Prompt Showcase: Coffee Brand Doc to TVC

Upload a café's brand introduction document, add one sentence, and get a complete brand film.

Reference file Prompt
Café brand introduction document Turn this brand story into a warm-toned brand TVC. Make it emotionally resonant — the kind of film that makes people want to visit the café and stay a while.

Real-World Fidelity: Lifelike Faces and Reference Consistency

Video generation is shifting from content novelty to production tooling, and production demands more than sharp frames — it demands believability. Wan3.0 targets accurate reproduction of both the physical world and digital information.

No More Same-Face AI People

Wan3.0's text-to-video mode is built for diverse, lifelike faces rather than the glossy, interchangeable look of typical AI characters. It renders facial features and skin detail with sharper sensitivity, keeps emotional expression natural and restrained, and links micro-expressions with body language. Even in group scenes, different characters carry distinct, nuanced emotions.

Consistency Across Characters, Props, Space, and Style

Consistency is the metric production teams care about most. In reference-to-video mode, Wan3.0 reproduces reference details precisely and holds them stable across key dimensions:

  • Characters — facial features, hairstyle and hair color, body shape, clothing, and accessories stay consistent at fine granularity, so characters never break the illusion.
  • Props — multi-angle appearance, hardware structure, logos, and material details remain stable shot to shot.
  • Space — character blocking and camera perspective are handled with correct spatial relationships.
  • Style — cinematic tone is rendered accurately, and multi-style projects avoid style bleed between looks.

For digital information, Wan3.0 also raises the aesthetic bar: harmonious color, fluid information delivery, and charts, data, and interface content that carry real visual polish.

Edit Videos Without Starting Over

Video editing capabilities, first introduced in Wan2.7, carry forward into Wan3.0. You can modify visuals, plot, and dialogue on generated videos — including flexible story re-creation — without regenerating from scratch. Iteration becomes a revision pass, not a do-over.

A Note on Current Limitations

During internal testing, we identified areas where Wan3.0 still has room to grow: audio texture and on-screen text rendering accuracy are improving but not yet where we want them. Both are active areas of refinement for upcoming versions.

Industry Use Cases: Where Wan3.0 Fits Your Pipeline

Wan3.0's upgrades are grounded in real industry needs, and the model is designed to slot back into real creative and production workflows.

Film, Episodic, and Creator Content

The combination of 30-second duration, real-scene fidelity, and precise consistency improves the experience of creating AI films, short dramas and animated series, music and dance videos, and lifestyle vlogs — helping creators finish work at dramatically lower cost.

Advertising and Marketing

Advertising demands high quality and deep customization. Wan3.0's everything-to-video capability breaks past static image-and-copy formats, letting creative concepts render directly as film. Across consumer electronics, beauty, automotive, FMCG, 3C, apparel, and software, Wan3.0 covers the full creative range from product showcase to brand storytelling.

Design and Creative Work

For UI interaction demos, software feature animations, and data visualization, Wan3.0 preserves the aesthetics and texture of the original design while elevating it into motion that tells a story — skipping the tedious animation production step so ideas are seen immediately.

Tourism and Destination Marketing

With Wan3.0's real-scene reproduction, destination films no longer require large-scale location shoots. Natural landscapes, cultural landmarks, and local cuisine — city image films and digitized cultural scenes can be produced at low cost and high speed, bringing faraway places to the screen.

Wan3.0 API Pricing

Wan3.0 video generation is billed per second of generated video, by output resolution:

Resolution Price per second
480P $0.05
720P $0.10
1080P $0.20

At these rates, a full 30-second 1080P generation costs $6.00**, and a 30-second 480P draft costs **$1.50 — practical for iterating on drafts at 480P and finishing at 1080P.

From Wan1.0 to Wan3.0: Eight Iterations Toward a World Simulator

The Wan model family has shipped 8 iterations from Wan1.0 to Wan3.0, evolving from asset generation to full creation, and from content generation to commercial delivery. Every extension of generation length and every new reference modality has pushed the model closer to faithfully reproducing the real world.

Wan3.0 marks a new generation of that trajectory. Looking ahead, the same capabilities — understanding richer information and generating more complete, physically consistent output — point toward embodied AI, autonomous driving, and industrial simulation, where simulating and reconstructing the physical world will find new answers.

How to Access Wan3.0

Wan3.0 is available on Alibaba Cloud Model Studio. You can review the model documentation, explore capabilities, and start generating today:

Frequently Asked Questions

What is Wan3.0?

Wan3.0 is the latest generation of Alibaba's Wan video generation model family, available on Alibaba Cloud Model Studio. It generates up to 30-second videos in a single pass from text, images, audio, video, or documents, with lifelike human rendering and strong reference-to-video consistency.

How long can Wan3.0 videos be?

Wan3.0 generates up to 30 seconds of video in a single generation. A smart duration feature recommends the right length based on your prompt, and video extension lets you continue a story beyond the initial generation.

What file formats does Wan3.0 accept as input?

Beyond text, image, audio, and video, Wan3.0 accepts documents in doc, xls, ppt, pdf, txt, key, pages, numbers, and md formats. Each request supports a maximum of 1 file or link, up to 100MB and 50 pages.

How much does the Wan3.0 API cost?

Wan3.0 is priced per second of generated video: $0.05/s (480P),** **$0.10/second at 720P, and $0.20/second at 1080P on Alibaba Cloud Model Studio.

How is Wan3.0 different from Wan2.x?

Wan3.0 extends single-generation length to 30 seconds (with smart duration and extension), adds document input for "everything to video," and significantly upgrades realism — diverse lifelike faces and fine-grained reference consistency for characters, props, spaces, and style. The video editing capability introduced in Wan2.7 carries forward.

How does Wan3.0 handle video editing?

Wan3.0 supports modifying visuals, plot, and dialogue on generated videos, including flexible story re-creation, so you can iterate on content without regenerating from scratch.

Where can I access Wan3.0?

Wan3.0 is available through Alibaba Cloud Model Studio. Visit the Wan3.0 model page to view documentation and start generating.

What are Wan3.0's current limitations?

Audio texture and on-screen text rendering accuracy are still improving. Both are being actively refined and will continue to mature in subsequent versions.

Start Creating with Wan3.0

From a single prompt to a full 30-second brand film — and from a PPT to a product ad — Wan3.0 turns any input into video. Try Wan3.0 on Alibaba Cloud Model Studio today.

To sharpen your results, explore our Wan3.0 Model Release page for more demos.

0 0 0
Share on

Alibaba Cloud Community

1,501 posts | 510 followers

You may also like

Comments