All Products
Search
Document Center

:wan3.0-video

Last Updated:Aug 06, 2026

Wan3.0-Video is an All-in-One video generation model that supports multi-modal inputs including text, image, video, and audio, enabling text-to-video, image-to-video (first frame/first-last frame), and reference-based video generation within a single model. It supports 480P, 720P, and 1080P resolution output, with a maximum video duration of 30 seconds. This model is currently in invitational preview.

Inference service provider

The inference service for the wan3.0-video model is provided by Alibaba Cloud Model Studio.

Model capabilities

Capability

Support status

Capability

Support status

Input modalities

Text Image Video Audio File Link

Output modalities

Video

Model playground

Supported

Function Calling

Unsupported

Structured output

Unsupported

Web search

Unsupported

Prefix completion

Unsupported

Context caching

Unsupported

Batch inference

Unsupported

Model fine-tuning

Unsupported

Context limits

Parameter

Value

Parameter

Value

Maximum input length

Maximum output length

Context length

Pricing

This topic only shows the original price for model calls, excluding any limited-time promotions or other offers. Please visit the Model Studio console for promotional offers.

China (Beijing)

Billing item

Price (USD)

Unit

Video generation (480P)

0.041256

Per second

Video generation (720P)

0.082513

Per second

Video generation (1080P)

0.165025

Per second

Singapore

Deployment scope: International

Billing item

Price (USD)

Unit

Video generation (480P)

0.05

Per second

Video generation (720P)

0.1

Per second

Video generation (1080P)

0.2

Per second

Rate limits

China (Beijing)

Parameter

Value

RPM (Requests per minute)

Singapore

Deployment scope: International

Parameter

Value

RPM (Requests per minute)