All Products
Search
Document Center

Alibaba Cloud Model Studio:Training and deployment pricing

Last Updated:Jul 14, 2026

This topic describes the billing rules and pricing for model training and model deployment on Alibaba Cloud Model Studio.

Training billing

Text generation models – Qwen

Note

For the training workflow, see Fine-tune Qwen. After training completes, deploy the new model before evaluating or calling it.

Method

Billed by training tokens

Formula

Model training fee = (Total tokens in training data + Total tokens in mixed training data) × Number of epochs × Training unit price (Minimum billing unit: 1 token)

View the estimated training fee at the bottom of the model training console, and click Computing Details to view the total number of training tokens, number of epochs, and training unit price.

Unit price for training

The following table lists the unit prices for training pre-trained models. The unit price for training a custom model matches that of the corresponding pre-trained model.

Qwen

Service

Code

Price

Qwen3-32B

qwen3-32b

$0.008/1,000 tokens

Qwen3-14B

qwen3-14b

$0.0016/1,000 tokens

Qwen-VL

Service

Code

Price

Qwen3-VL-8B-Instruct

qwen3-vl-8b-instruct

$0.002/1,000 tokens

Qwen3-VL-8B-Thinking

qwen3-vl-8b-thinking

$0.002/1,000 tokens

Image generation models – Wan

Note

For the training workflow, see Fine-tune image generation models. After training completes, deploy the new model before calling it.

Method

Billed by training tokens

Formula

Model training fee = Total training tokens × Training unit price (Billing unit: per 1,000 tokens)

Formula for total training tokens

Where:

  • max_steps: A hyperparameter specified during training, representing the maximum number of training steps (configured when creating a fine-tuning job).

  • Lstep: The token consumption per step. The formula is:

    Lstep is approximately equal to Lmax. Lmax is determined by the max_token_length and generation_type, as shown below:

generation_type

max_token_length

Lmax

t2i (text-to-image)

1k

12,800

2k

23,220

i2i (image-to-image)

1k

23,220

2k

32,000

Note

The above formula provides an approximation. Actual billing is based on the usage field returned by the system.

Model

Code

Training price (per 1K tokens)

Wan image generation

wan2.7-image-pro

$0.015

Wan image generation

wan2.7-image

$0.015

Billing example

Suppose you fine-tune the wan2.7-image-pro model for t2i. The parameters are: max_steps = 200, max_token_length = "1k", and the training price is $0.015 per 1,000 tokens:

  • From the table: Lmax = 12,800 (generation_type=t2i, max_token_length=1k), Lstep ≈ Lmax = 12,800

  • Total training tokens ≈ 200 × 12800 = 2560000 = 2560 thousand tokens

  • Model training fee ≈ 2560× 0.015 = $38.4

Video generation models – Wan

Note

For the training workflow, see Fine-tuning video generation models. After training completes, deploy the new model before calling it.

Method

Billed by training tokens

Formula

Model training fee = Total training tokens × Training unit price (Billing unit: per 1,000 tokens)

Formula for total training tokens

Where:

  • N: Total number of videos in the training set.

  • max_pixels: A hyperparameter specified during training, representing the maximum number of pixels for a video (configured when creating a fine-tuning job).

  • n_epochs: A hyperparameter specified during training, representing the number of loops (configured when creating a fine-tuning job).

    • The conversion between n_epochs and steps is: steps = n_epochs × ⌈dataset_size / batch_size⌉, i.e., n_epochs = steps / ⌈dataset_size / batch_size⌉.

    • When the dataset contains only 1 sample and batch_size = 1, n_epochs = steps. We recommend a total of at least 800 steps.

  • Billing duration calculation rule for a single video: First, round the original video duration (in seconds) to the nearest integer, then determine the final value based on model limits.

    • wan2.7 model: Billing duration=min(10, rounded duration), meaning a single video is billed for a maximum of 10 seconds.

    • wan2.6 model: Billing duration=min(10, rounded duration), meaning a single video is billed for a maximum of 10 seconds.

    • wan2.5 model: Billing duration=min(10, rounded duration), meaning a single video is billed for a maximum of 10 seconds.

    • wan2.2 model: Billing duration=min(5, rounded duration), meaning a single video is billed for a maximum of 5 seconds.

Model

Code

Training price (per 1K tokens)

Image-to-video (first frame)

wan2.7-i2v

$0.3

wan2.6-i2v

$0.08

wan2.5-i2v-preview

$0.05

wan2.2-i2v-flash

$0.03

Image-to-video (first and last frames)

wan2.2-kf2v-flash

$0.03

Billing examples

  1. wan2.7-i2v cost estimation (single data)

Assume a training set contains 1 video with a duration of 10 seconds. With batch_size = 1 (recommended), n_epochs = steps / ⌈1(dataset_size) / 1(batch_size)⌉ = steps.

Training unit price = $0.3/thousand tokens. Taking max_pixels = 36864 and n_epochs = 800 as an example:

  • Total training tokens = 10 × (36864 / 1024) × 800 = 288,000 = 288 thousand tokens

  • Model training fee = 288 × $0.3 = $86.4

max_pixels

Common steps

n_epochs

Estimated tokens

Estimated cost (USD)

36864

800

800

288,000

$86.4

1,000

1,000

360,000

$108

2,000

2,000

720,000

$216

65536

800

800

512,000

$153.6

1,000

1,000

640,000

$192

2,000

2,000

1,280,000

$384

102400

800

800

800,000

$240

1,000

1,000

1,000,000

$300

2,000

2,000

2,000,000

$600

  1. wan2.7-i2v cost estimation (multiple data)

Assume a training set contains 2 videos with durations of 3.4 seconds and 11.5 seconds. Parameters: max_pixels = 36864, n_epochs = 800. Training unit price = $0.3/thousand tokens:

  • Duration calculation:

    • Video 1: 3.4 seconds is rounded to 3. Billable duration = min(10, 3) = 3.

    • Video 2: 11.5 seconds is rounded to 11. Billable duration = min(10, 11) = 10.

    • Total billable duration = 3 + 10 = 13 seconds.

  • Total training tokens = 13 × (36864/1024) × 800 = 374,400 = 374.4 thousand tokens.

  • Model training fee = 374.4 × 0.3 = $112.32.

  1. wan2.5-i2v-preview cost estimation (multiple data)

Suppose you fine-tune the wan2.5 model. The training set contains two videos: 3.4 seconds and 11.5 seconds. The parameters are max_pixels = 36864 and n_epochs = 400. The unit training price is $0.05 per 1,000 tokens.

  • Duration calculation:

    • Video 1: 3.4 seconds is rounded to 3. Billable duration: min(10, 3) = 3 seconds.

    • Video 2: 11.5 seconds is rounded to 11. Billable duration: min(10, 11) = 10 seconds.

    • Total billable duration: 3 + 10 = 13 seconds.

  • Total training tokens = 13 × (36864 / 1024) × 400 = 187,200 = 187.2 thousand tokens.

  • Model training fee = 187.2 × 0.05 = $9.36.

Deployment billing

Text generation models: Qwen

Time-based billing (Provisioned Throughput)

Cost = Usage Duration × (Input TPM Unit Price × Input TPM + Output TPM Unit Price × Output TPM)

For the pay-as-you-go method, usage is billed hourly, and the unit price is based on the hourly rates in the table below. For the subscription method, usage is billed daily, and the unit price is based on the daily rates in the table below.

  • Subscription orders take effect immediately after payment. An N-day subscription is valid until 23:59 on the Nth day. If an order is placed after 22:00, the expiration date is automatically extended by one day.

  • After a subscription order expires, the service is stopped after a 2-hour grace period. After the service is stopped, the resources are retained for 14 hours and then released.

  • Subscription orders cannot be terminated early.

  • For the pay-as-you-go method, if your account has an overdue payment, the deployed resources are retained and continue to be billed for 24 hours, during which the service remains available. After 24 hours, the system stops billing, and the model deployment enters an overdue state. The underlying resources are deleted, but the model deployment task is retained. After you pay the overdue amount, the system reallocates resources, restores the service, and resumes billing. To stop incurring charges, you must delete the model deployment task. Billing stops after the task is successfully deleted.

If the model input exceeds the maximum input tokens or the purchased TPM, the call automatically switches to the pay-as-you-go mode for the current model. In this case, inference performance may decrease and will be subject to the public traffic control of the current snapshot model in the workspace. Costs are charged based on the model invocation (pay-as-you-go) standard.

  • In this case, the API call returns a header that contains x-dashscope-ptu-overflow:true.

  • To view TPM statistics, go to Model Monitoring.

For the specific refund rules for scale-in scenarios (downgrades), see Refund rules for downgrades.

Singapore
Qwen

Model name

Model code

Max input tokens

Pay-as-you-go input

Per 10k TPM/hour

Pay-as-you-go output

Per 1k TPM/hour

Subscription input

Per 10k TPM/day

Subscription output

Per 1k TPM/day

Qwen3.7-Max-2026-05-20

qwen3.7-max-2026-05-20

256K

$6

$1.8

$72

$21.6

Qwen3.7-Plus-2026-05-26

qwen3.7-plus-2026-05-26

256K

$0.96

$0.384

$11.52

$4.608

Qwen3.6-Plus-2026-04-02

qwen3.6-plus-2026-04-02

128K

$1.2

$0.72

$14.4

$8.64

Qwen3.5-Plus-2026-04-20

qwen3.5-plus-2026-04-20

128K

$0.96

$0.576

$11.52

$6.912

DeepSeek

Model name

Model code

Max input tokens

Pay-as-you-go input

Per 10k TPM/hour

Pay-as-you-go output

Per 1k TPM/hour

Subscription input

Per 10k TPM/day

Subscription output

Per 1k TPM/day

DeepSeek-v4-Flash

deepseek-v4-flash

256K

$0.72

$0.144

$8.64

$1.728

DeepSeek-v4-Pro

deepseek-v4-pro

256K

$8.64

$1.728

$103.68

$20.736

DeepSeek-v3.2

deepseek-v3.2

64K

$2.05

$0.616

$24.62

$7.387

Qwen-VL

Model name

Model code

Max input tokens

Pay-as-you-go input

Per 10k TPM/hour

Pay-as-you-go output

Per 1k TPM/hour

Subscription input

Per 10k TPM/day

Subscription output

Per 1k TPM/day

Qwen3-VL-Plus-2025-09-23

qwen3-vl-plus-2025-09-23

128K

$0.48

$0.384

$5.76

$4.608

More models

Model name

Model code

Max input tokens

Pay-as-you-go input

Per 10k TPM/hour

Pay-as-you-go output

Per 1k TPM/hour

Subscription input

Per 10k TPM/day

Subscription output

Per 1k TPM/day

GLM-5.1

glm-5.1

64K

$5.04

$1.584

$64.8

$19.008

China (Beijing)
Qwen

Model name

Model code

Max input tokens

Pay-as-you-go input

Per 10k TPM/hour

Pay-as-you-go output

Per 1k TPM/hour

Subscription input

Per 10k TPM/day

Subscription output

Per 1k TPM/day

Qwen3.7-Max-2026-05-20

qwen3.7-max-2026-05-20

256K

$3.96

$1.188

$47.53

$14.258

Qwen3.7-Plus-2026-05-26

qwen3.7-plus-2026-05-26

256K

$0.66

$0.264

$7.92

$3.168

Qwen3.6-Plus-2026-04-02

qwen3.6-plus-2026-04-02

128K

$0.67

$0.397

$7.93

$4.753

Qwen3.5-Plus-2026-04-20

qwen3.5-plus-2026-04-20

128K

$0.26

$0.16

$3.17

$1.9

Qwen3-Max-2025-09-23

qwen3-max-2025-09-23

128K

$1.11

$0.45

$13.32

$5.4

Qwen-Flash-2025-07-28

qwen-flash-2025-07-28

128K

$0.06

$0.06

$0.72

$0.72

Qwen-Plus-2025-12-01

qwen-plus-2025-12-01

128K

$0.28

Non-thinking mode: $0.07

Thinking mode: $0.28

$3.36

Non-thinking mode: $0.84

Thinking mode: $3.36

DeepSeek

Model name

Model code

Max input tokens

Pay-as-you-go input

Per 10k TPM/hour

Pay-as-you-go output

Per 1k TPM/hour

Subscription input

Per 10k TPM/day

Subscription output

Per 1k TPM/day

DeepSeek-v4-Flash

deepseek-v4-flash

256K

$0.5

$0.099

$5.94

$1.188

DeepSeek-v4-Pro

deepseek-v4-pro

256K

$5.94

$1.188

$71.3

$14.26

DeepSeek-v3.2

deepseek-v3.2

64K

$1.04

$0.16

$12.48

$1.92

DeepSeek-v3

deepseek-v3

64K

$0.99

$0.396

$11.9

$4.75

Qwen-VL

Model name

Model code

Max input tokens

Pay-as-you-go input

Per 10k TPM/hour

Pay-as-you-go output

Per 1k TPM/hour

Subscription input

Per 10k TPM/day

Subscription output

Per 1k TPM/day

Qwen3-VL-Plus-2025-09-23

qwen3-vl-plus-2025-09-23

128K

$0.35

$0.35

$4.2

$4.2

More models

Model name

Model code

Max input tokens

Pay-as-you-go input

Per 10k TPM/hour

Pay-as-you-go output

Per 1k TPM/hour

Subscription input

Per 10k TPM/day

Subscription output

Per 1k TPM/day

GLM-5.1

glm-5.1

64K

$2.97

$1.19

$35.65

$14.26

Time-based billing (Model Unit)

Cost = Usage Duration (hours) × Number of Model Units × Model Unit Price

For the pay-as-you-go method, the "Model Unit Price" is the "Hourly Price" from the table below. For the monthly subscription method, the formula is: Number of Months × Number of Model Units × Monthly Price.

  • For subscriptions, if you unsubscribe within the first month, the daily unit price (≈ monthly unit price / 30) is charged at 1.2 times the standard rate. Usage for less than a day is billed as a full day.

Note

For the Model Unit pay-as-you-go method, computing power resources are allocated on a first-come, first-served basis. A full refund is issued if the purchase is unsuccessful.

Singapore
Text generation

Model name

Model code

Model unit specification

Hourly price ($)

Minimum billing unit: minute

Monthly price ($)

Minimum billing unit: day

Qwen3.6-Plus-2026-04-02

qwen3.6-plus-2026-04-02

MU1 x 8

$88

$41,832

Qwen3.5-39B-A17B

qwen3.5-397b-a17b

MU2 x 8

$112

$52,392

Qwen3.5-35B-A3B

qwen3.5-35b-a3b

MU2 x 8

$112

$52,392

Qwen3-32B

qwen3-32b

MU1 x 4

$44

$20,916

MU2 x 8

$112

$52,392

Qwen3-14B

qwen3-14b

MU1 x 4

$44

$20,916

GLM-5.1

glm-5.1

MU2 x 8

$112

$52,392

DeepSeek-V4-Flash

deepseek-v4-flash

MU1 x 8

$88

$41,832

Multimodal

Model name

Model code

Model unit specification

Hourly price ($)

Minimum billing unit: minute

Monthly price ($)

Minimum billing unit: day

Qwen3-VL-32B-Instruct

qwen3-vl-32b-instruct

MU2 x 8

$112

$52,392

Qwen3-VL-8B-Instruct

qwen3-vl-8b-instruct

MU1 x 2

$22

$10,458

Model types:

  • Instruct - The deployed model performs inference in non-thinking mode.

China (Beijing)
Text generation
Qwen

Model name

Model code

Model unit specification

Hourly price ($)

Minimum billing unit: minute

Monthly price ($)

Minimum billing unit: day

Qwen3.7-Plus-2026-05-26

qwen3.7-plus-2026-05-26

MU3 x 8

$150.72

$72,577.152

Qwen3.6-35B-A3B

qwen3.6-35b-a3b

MU8 x 1

$6.464

$3,080.477

MU9 x 1

$7.014

$3,383.024

Qwen3.6-27B

qwen3.6-27b

MU9 x 1

$7.014

$3,383.024

Qwen3.6-Flash-2026-04-16

qwen3.6-flash-2026-04-16

MU1 x 2

$14.852

$7,183.564

Qwen3.6-Plus-2026-04-02

qwen3.6-plus-2026-04-02

MU1 x 8

$59.408

$28,734.256

Qwen3.5-397B-A17B

qwen3.5-397b-a17b

MU2 x 8

$69.312

$33,044.72

MU3 x 8

$150.72

$72,577.152

MU6 x 16

$55.008

$26,599.92

Qwen3.5-122B-A10B

qwen3.5-122b-a10b

MU1 x 4

$29.704

$14,367.128

MU2 x 8

$69.312

$33,044.72

MU6 x 16

$55.008

$26,599.92

MU9 x 2

$14.028

$6,766.048

Qwen3.5-35B-A3B

qwen3.5-35b-a3b

MU1 x 2

$14.852

$7,183.564

MU2 x 8

$69.312

$33,044.72

MU8 x 1

$6.464

$3,080.477

MU9 x 1

$7.014

$3,383.024

Qwen3.5-27B

qwen3.5-27b

MU9 x 1

$7.014

$3,383.024

Qwen3.5-9B

qwen3.5-9b

MU8 x 1

$6.464

$3,080.477

MU9 x 1

$7.014

$3,383.024

Qwen3.5-Flash-2026-02-23

qwen3.5-flash-2026-02-23

MU1 x 2

$14.852

$7,183.564

Qwen3.5-Plus-2026-02-15

qwen3.5-plus-2026-02-15

MU1 x 8

$59.408

$28,734.256

MU3 x 8

$150.72

$72,577.152

Qwen3-235B-A22B-Instruct

qwen3-235b-a22b-instruct-2507

MU1 x 4

$29.704

$14,367.128

MU2 x 8

$69.312

$33,044.72

Qwen3-Next-80B-A3B-Instruct

qwen3-next-80b-a3b-instruct

MU1 x 2

$14.852

$7,183.564

Qwen3-32B

qwen3-32b

MU1 x 4

$29.704

$14,367.128

MU6 x 4

$13.752

$6,649.98

Qwen3-30B-A3B

qwen3-30b-a3b

MU9 x 2

$14.028

$6,766.048

Qwen3-30B-A3B-Instruct-2507

qwen3-30b-a3b-instruct-2507

MU1 x 4

$29.704

$14,367.128

MU2 x 8

$69.312

$33,044.72

Qwen3-8B

qwen3-8b

MU1 x 2

$14.852

$7,183.564

MU2 x 2

$17.328

$8,261.18

MU5 x 1

$2.888

$1,394.329

Qwen3-4B

qwen3-4b

MU1 x 2

$14.852

$7,183.564

MU5 x 1

$2.888

$1,394.329

Qwen3-1.7B

qwen3-1.7b

MU1 x 2

$14.852

$7,183.564

MU5 x 1

$2.888

$1,394.329

Qwen3-Max-2025-09-23

qwen3-max-2025-09-23

MU2 x 8

$69.312

$33,044.72

MU3 x 8

$150.72

$72,577.152

Qwen2.5-72B

qwen2.5-72b-instruct

MU1 x 4

$29.704

$14,367.128

Qwen2.5-32B

qwen2.5-32b-instruct

MU1 x 4

$29.704

$14,367.128

Qwen2.5-14B

qwen2.5-14b-instruct

MU1 x 2

$14.852

$7,183.564

Qwen2.5-7B

qwen2.5-7b-instruct

MU1 x 2

$14.852

$7,183.564

MU5 x 1

$2.888

$1,394.329

Qwen2.5-3B-Instruct

qwen2.5-3b-instruct

MU5 x 1

$2.888

$1,394.329

Qwen-Flash-2025-07-28

qwen-flash-2025-07-28

MU1 x 4

$29.704

$14,367.128

Qwen-Plus-2025-07-28

qwen-plus-2025-07-28

MU1 x 4

$29.704

$14,367.128

Qwen-Plus-2025-12-01

qwen-plus-2025-12-01

MU1 x 4

$29.704

$14,367.128

GLM

Model name

Model code

Model unit specification

Hourly price ($)

Minimum billing unit: minute

Monthly price ($)

Minimum billing unit: day

GLM-5.1

glm-5.1

MU2 x 8

$69.312

$33,044.72

MU3 x 8

$150.72

$72,577.152

MU6 x 16

$55.008

$26,599.92

GLM-5

glm-5

MU3 x 8

$150.72

$72,577.152

GLM-4.7

glm-4.7

MU6 x 16

$55.008

$26,599.92

DeepSeek

Model name

Model code

Model unit specification

Hourly price ($)

Minimum billing unit: minute

Monthly price ($)

Minimum billing unit: day

DeepSeek-V4-Flash

deepseek-v4-flash

MU1 x 8

$59.408

$28,734.256

DeepSeek-V3.2

deepseek-v3.2

MU2 x 8

$69.312

$33,044.72

Other models

Model name

Model code

Model unit specification

Hourly price ($)

Minimum billing unit: minute

Monthly price ($)

Minimum billing unit: day

MiniMax-M2.5

MiniMax-M2.5

MU1 x 8

$59.408

$28,734.256

Kimi-K2.5

kimi-k2.5

MU2 x 8

$69.312

$33,044.72

Multimodal
Qwen-VL

Model name

Model code

Model unit specification

Hourly price ($)

Minimum billing unit: minute

Monthly price ($)

Minimum billing unit: day

Qwen3-VL-235B-A22B-Instruct

qwen3-vl-235b-a22b-instruct

MU1 x 4

$29.704

$14,367.128

Qwen3-VL-235B-A22B-Thinking

qwen3-vl-235b-a22b-thinking

MU1 x 4

$29.704

$14,367.128

Qwen3-VL-32B-Instruct

qwen3-vl-32b-instruct

MU2 x 8

$69.312

$33,044.72

Qwen3-VL-8B-Instruct

qwen3-vl-8b-instruct

MU1 x 2

$14.852

$7,183.564

Qwen3-VL-Flash-2025-10-15

qwen3-vl-flash-2025-10-15

MU1 x 4

$29.704

$14,367.128

Qwen3-VL-Plus-2025-09-23

qwen3-vl-plus-2025-09-23

MU1 x 4

$29.704

$14,367.128

Qwen-VL-Max-2025-08-13

qwen-vl-max-2025-08-13

MU6 x 4

$13.752

$6,649.98

Qwen-VL-OCR-2025-11-20

qwen-vl-ocr-2025-11-20

MU6 x 4

$13.752

$6,649.98

Qwen Omni

Model name

Model code

Model unit specification

Hourly price ($)

Minimum billing unit: minute

Monthly price ($)

Minimum billing unit: day

Qwen3.5-Omni-Flash

qwen3.5-omni-flash

MU8 x 1

$6.464

$3,080.477

MU9 x 1

$7.014

$3,383.024

Qwen3.5-Omni-Plus

qwen3.5-omni-plus

MU9 x 8

$56.112

$27,064.192

Model types:

  • Instruct - The deployed model performs inference in non-thinking mode.

  • Thinking - The deployed model performs inference in thinking mode.

By model token usage

Cost = Number of Input Tokens × Input Unit Price + Number of Output Tokens × Output Unit Price (Minimum billing unit: 1 token)

  • Billing by model token usage is only supported after you have completed Supervised Fine-Tuning (SFT) for the following foundation models and you have obtained a custom model.

Singapore

Foundation model

Model code

Input

$/1k tokens

Output

$/1k tokens

Qwen3-14B

qwen3-14b

$0.00035

Non-thinking mode: $0.0014

Thinking mode: $0.0042

Image generation models – Wan

Deployment is free. Invocations are billed at the standard rate of the fine-tuned base model. For the training workflow, see Fine-tune image generation models.

Model ID

LoRA Deployment & Invocation Price

wan2.7-image-pro

$0.075/image

wan2.7-image

$0.03/image

FAQ

Q: When does billing for model deployment start?

A: Billing starts when the model status changes to Running. No charges apply during Deploying, Overdue Payment, or Deployment Failed.

Q: Am I charged if I cancel a training job?

A: Yes. If you cancel training manually, you are charged for all tokens processed before cancellation. Training jobs interrupted by system errors or other non-user causes are not charged.

Q: How do I view invocation statistics for a deployed model?

A: Visit the Model Monitoring (Singapore), Model Monitoring (Virginia), or Model Monitoring (Beijing) page.