All Products
Search
Document Center

Alibaba Cloud Model Studio:DTU Dedicated compute deployment

Last Updated:Sep 16, 2026

Dedicated compute deployment offers two plans: Dedicated Throughput Unit (DTU) and Model Unit (MU), providing dedicated GPU resources for specified models. DTU targets newly released models (billed by input/output TPM x duration); MU is used for existing models (billed by model unit count x duration). This document introduces both plans' features, billing, usage flow, and supported models.

Overview

Dedicated compute deployment provides dedicated GPU resources and performance guarantees for specified models, supporting both base and custom model deployment. Bailian offers two dedicated compute plans:

  • DTU (Dedicated Throughput Unit): Sold by Input/Output TPM, providing more direct throughput guarantees. Fully managed inference service with dedicated underlying GPU resources maintained by the platform.
  • Model Unit (MU): Configures compute power by usage duration and model unit count, with dedicated resources. Supports custom performance metrics and PD disaggregation mode. Billing granularity is model unit count x duration.

DTU and MU are differentiated by model applicability: newly released models use DTU (metered by input/output TPM), while existing models continue to use MU (metered by model unit count). Both provide dedicated compute, resource isolation, and fine-tuned model deployment. For new deployments, choose DTU if the model supports it.

Shared capabilities of both plans:

  • Dedicated GPU resources; the inference environment is physically isolated from other users.
  • Supports base models and custom models (Bailian fine-tuned models or user-uploaded models).
  • The platform maintains the underlying operations, no need to manage GPUs.
  • No proactive RPM/TPM limits; traffic is bound by actual capacity.

For the general model deployment workflow, see Dedicated Deployment Overview and Provisioned Throughput.

Solution Selection

Bailian offers multiple capacity and billing plans for inference calls. DTU suits scenarios requiring dedicated deployment, fine-tuned model deployment, low latency with high concurrency, or data isolation. The plans are compared in the table below.

For a complete comparison of all billing methods including Token pay-as-you-go and PTU, see the Dedicated Deployment Overview main table.

Comparison Dimension

DTU

PTU

Token Pay-as-you-go

PAI/Lingjun

Feature

Dedicated compute + fine-tuned models

Throughput guarantee + low latency

Elastic and zero-threshold

Custom runtime

Resource isolation

Physical isolation

Logical isolation

Shared pool

Physical isolation

Fine-tuned models

Full-parameter + LoRA

Not supported

LoRA

Supported

Custom framework

Not supported

Not supported

Not supported

Supported

Billing mode

TPM quota × duration

TPM quota × duration

Token usage

GPU × duration

Billing rules

DTU billing

DTU is billed separately by input TPM and output TPM, using a pre-paid (monthly) model. Each model has fixed input/output baseline TPM (see table below), and you must purchase in integer multiples of the baseline TPM. Purchase at least 1× baseline TPM for both input and output; for production, 2× or above is recommended for each.

Fees are calculated by the backend based on model, input/output TPM, service region, and purchase duration. The input/output unit prices and monthly price for each model are shown in the table below. Unit prices are billed per kTPM·month, and the monthly price is the total for 1× baseline TPM of both input and output. Fine-tuned model prices are the same as the corresponding base model, subject to the actual console display.

MU billing

Cost = Usage duration (hours) × Number of Model Units × Model Unit unit price

The "Model Unit unit price" takes the "Hourly unit price" column in the table below for postpaid scenarios; for monthly subscription prepaid billing, the formula becomes Number of monthly subscriptions × Number of Model Units × Monthly unit price.

  • For the first month of a prepaid purchase, if you cancel the subscription early within the first month, the daily unit price (≈ Monthly unit price / 30) will be billed at 1.2 times the rate (less than one day is billed as one day)

NoteThe compute resources for the Model Unit postpaid method are first-come, first-served. If the purchase fails, a full refund will be issued.

Supported models and pricing

DTU pricing and performance benchmarks

Singapore

Model

Baseline Input (TPM)

Baseline Output (TPM)

Input Unit Price (USD/kTPM·mo)

Output Unit Price (USD/kTPM·mo)

Monthly Price (USD)

qwen3.7-plus-2026-05-26

1,192,000

148,000

37

295

87,764

1,372,000

170,000

32

257

87,594

660,000

124,000

66

352

87,208

704,000

132,000

62

331

87,340

The performance reference data for each model under standard workload is shown in the table below, subject to the actual console display.

The following performance reference data was measured at a 0% cache hit rate. In actual use, as the cache hit rate increases, model performance improves accordingly.

Model

Input Length

Output Length

Cache Hit Rate

First Token Latency (ms)

Per-Token Latency (ms)

qwen3.7-plus-2026-05-26

16,000

2,000

0

888

10

16,000

2,000

0

2,418

15

2,600

500

0

819

14

2,600

500

0

970

13

Among these, qwen3.8-27b currently requires whitelist access, while the other models are on the Singapore official site.

MU pricing

Singapore

Text Generation

Model

Baseline Input (TPM)

Baseline Output (TPM)

Input Unit Price (USD/kTPM·Month)

Output Unit Price (USD/kTPM·Month)

Monthly Total Price (USD)

qwen3.7-plus-2026-05-26

1,192,000

148,000

37

295

87,764

1,372,000

170,000

32

257

87,594

660,000

124,000

66

352

87,208

704,000

132,000

62

331

87,340

Multimodal

Model

Input Length

Output Length

Cache Hit Rate

First Token Latency (ms)

Per-token Latency (ms)

qwen3.7-plus-2026-05-26

16,000

2,000

0

888

10

16,000

2,000

0

2,418

15

2,600

500

0

819

14

2,600

500

0

970

13

Model type:

  • Instruct - After the model is deployed, inference runs in non-thinking mode.

China (Beijing)

Text generation

Qwen

Model Name

Model Code

Model Unit Specification

Hourly Unit Price (USD)

Min Billing: Minute

Monthly Unit Price (USD)

Min Billing: Day

Qwen3.8-27B

qwen3.8-27b

MU9 x 4

USD28.056

USD13,532.096

Qwen3.7-Plus-2026-05-26

qwen3.7-plus-2026-05-26

MU2 x 8

USD69.312

USD33,044.72

MU3 x 8

USD150.72

USD72,577.152

Qwen3.6-35B-A3B

qwen3.6-35b-a3b

MU1 x 8

USD59.408

USD28,734.256

MU2 x 8

USD69.312

USD33,044.72

MU3 x 8

USD150.72

USD72,577.152

MU8 x 1

USD6.464

USD3,080.477

MU9 x 1

USD7.014

USD3,383.024

Qwen3.6-27B

qwen3.6-27b

MU9 x 1

USD7.014

USD3,383.024

Qwen3.6-Flash-2026-04-16

qwen3.6-flash-2026-04-16

MU1 x 2

USD14.852

USD7,183.564

MU3 x 8

USD150.72

USD72,577.152

Qwen3.6-Plus-2026-04-02

qwen3.6-plus-2026-04-02

MU1 x 8

MU1 x 16(PD Disaggregation Mode)

USD59.408

PD Disaggregation Mode: USD118.816

USD28,734.256

PD Disaggregation Mode: USD57,468.512

MU2 x 8

USD69.312

USD33,044.72

Qwen3.5-397B-A17B

qwen3.5-397b-a17b

MU3 x 8

MU3 x 16(PD Disaggregation Mode)

USD150.72

PD Disaggregation Mode: USD301.44

USD72,577.152

PD Disaggregation Mode: USD145,154.304

MU6 x 16

USD55.008

USD26,599.92

Qwen3.5-122B-A10B

qwen3.5-122b-a10b

MU1 x 4

USD29.704

USD14,367.128

MU6 x 16

USD55.008

USD26,599.92

Qwen3.5-35B-A3B

qwen3.5-35b-a3b

MU1 x 2

USD14.852

USD7,183.564

MU2 x 8

USD69.312

USD33,044.72

MU3 x 8

USD150.72

USD72,577.152

MU9 x 1

USD7.014

USD3,383.024

Qwen3.5-27B

qwen3.5-27b

MU2 x 8

USD69.312

USD33,044.72

MU3 x 8

USD150.72

USD72,577.152

MU8 x 1

USD6.464

USD3,080.477

MU9 x 1

USD7.014

USD3,383.024

Qwen3.5-9B

qwen3.5-9b

MU1 x 2

USD14.852

USD7,183.564

MU2 x 8

USD69.312

USD33,044.72

Qwen3.5-Flash-2026-02-23

qwen3.5-flash-2026-02-23

MU1 x 2

USD14.852

USD7,183.564

Qwen3.5-Plus-2026-02-15

qwen3.5-plus-2026-02-15

MU1 x 8

MU1 x 16(PD Disaggregation Mode)

USD59.408

PD Disaggregation Mode: USD118.816

USD28,734.256

PD Disaggregation Mode: USD57,468.512

MU2 x 8

USD69.312

USD33,044.72

MU3 x 8

MU3 x 16(PD Disaggregation Mode)

USD150.72

PD Disaggregation Mode: USD301.44

USD72,577.152

PD Disaggregation Mode: USD145,154.304

Qwen3-235B-A22B-Instruct-2507

qwen3-235b-a22b-instruct-2507

MU1 x 4

USD29.704

USD14,367.128

MU2 x 8

USD69.312

USD33,044.72

MU3 x 8

USD150.72

USD72,577.152

Qwen3-32B

qwen3-32b

MU6 x 16

USD55.008

USD26,599.92

Qwen3-30B-A3B-Thinking-2507

qwen3-30b-a3b-thinking-2507

MU1 x 2

USD14.852

USD7,183.564

Qwen3-4B

qwen3-4b

MU1 x 2

USD14.852

USD7,183.564

MU5 x 1

USD2.888

USD1,394.329

Qwen3-Embedding-0.6B

qwen3-embedding-0.6b

MU5 x 1

USD2.888

USD1,394.329

MU6 x 1

USD3.438

USD1,662.495

Qwen3-MoE-Rerank-0.6B

qwen3-moe-rerank-0.6b

MU5 x 1

USD2.888

USD1,394.329

Qwen3-Rerank-0.6B

qwen3-rerank-0.6b

MU5 x 1

USD2.888

USD1,394.329

MU6 x 1

USD3.438

USD1,662.495

Qwen3-Max-2025-09-23

qwen3-max-2025-09-23

MU2 x 8

USD69.312

USD33,044.72

MU3 x 8

USD150.72

USD72,577.152

Qwen3-Rerank

qwen3-rerank

MU5 x 1

USD2.888

USD1,394.329

Qwen2.5-Open Source-72B

qwen2.5-72b-instruct

MU1 x 8

USD59.408

USD28,734.256

Qwen2.5-Open Source-14B

qwen2.5-14b-instruct

MU1 x 2

USD14.852

USD7,183.564

Qwen2.5-Open Source-7B

qwen2.5-7b-instruct

MU1 x 2

USD14.852

USD7,183.564

MU5 x 1

USD2.888

USD1,394.329

Qwen-Plus-2025-07-28

qwen-plus-2025-07-28

MU1 x 4

MU1 x 16(PD Disaggregation Mode)

USD29.704

PD Disaggregation Mode: USD118.816

USD14,367.128

PD Disaggregation Mode: USD57,468.512

Qwen-Plus-2025-12-01

qwen-plus-2025-12-01

MU1 x 4

USD29.704

USD14,367.128

Qwen-Plus-Character-2025-11-06

qwen-plus-character-2025-11-06

MU1 x 4

USD29.704

USD14,367.128

GLM

Model Name

Model Code

Model Unit Specification

Hourly Unit Price (USD)

Min Billing: Minute

Monthly Unit Price (USD)

Min Billing: Day

GLM-5.1

glm-5.1

MU2 x 8

USD69.312

USD33,044.72

MU3 x 16(PD Disaggregation Mode)

PD Disaggregation Mode: USD301.44

PD Disaggregation Mode: USD145,154.304

MU6 x 16

USD55.008

USD26,599.92

GLM-5

glm-5

MU3 x 16(PD Disaggregation Mode)

PD Disaggregation Mode: USD301.44

PD Disaggregation Mode: USD145,154.304

GLM-4.7

glm-4.7

MU6 x 32(PD Disaggregation Mode)

PD Disaggregation Mode: USD110.016

PD Disaggregation Mode: USD53,199.84

DeepSeek

Model Name

Model Code

Model Unit Specification

Hourly Unit Price (USD)

Min Billing: Minute

Monthly Unit Price (USD)

Min Billing: Day

DeepSeek-v4-Flash

deepseek-v4-flash

MU1 x 8

USD59.408

USD28,734.256

MU3 x 8

USD150.72

USD72,577.152

DeepSeek-v3.2

deepseek-v3.2

MU2 x 16(PD Disaggregation Mode)

PD Disaggregation Mode: USD138.624

PD Disaggregation Mode: USD66,089.44

Other models

Model Name

Model Code

Model Unit Specification

Hourly Unit Price (USD)

Min Billing: Minute

Monthly Unit Price (USD)

Min Billing: Day

Kimi-K2.5

kimi-k2.5

MU2 x 8

USD69.312

USD33,044.72

Multimodal

Qwen VL

Model Name

Model Code

Model Unit Specification

Hourly Unit Price (USD)

Min Billing: Minute

Monthly Unit Price (USD)

Min Billing: Day

Qwen3-VL-235B-A22B-Thinking

qwen3-vl-235b-a22b-thinking

MU1 x 8

USD59.408

USD28,734.256

MU2 x 8

USD69.312

USD33,044.72

MU3 x 8

USD150.72

USD72,577.152

Qwen3-VL-32B-Instruct

qwen3-vl-32b-instruct

MU2 x 8

USD69.312

USD33,044.72

MU3 x 8

USD150.72

USD72,577.152

Qwen3-VL-8B-Instruct

qwen3-vl-8b-instruct

MU1 x 2

USD14.852

USD7,183.564

MU5 x 1

USD2.888

USD1,394.329

Qwen3-VL-4B-Instruct

qwen3-vl-4b-instruct

MU1 x 2

USD14.852

USD7,183.564

Qwen3-VL-2B-Instruct

qwen3-vl-2b-instruct

MU5 x 1

USD2.888

USD1,394.329

Qwen3-VL-Embedding-2B

qwen3-vl-embedding-2b

MU5 x 1

USD2.888

USD1,394.329

Qwen3-VL-Flash-2025-10-15

qwen3-vl-flash-2025-10-15

MU1 x 4

USD29.704

USD14,367.128

Qwen3-VL-Plus-2025-09-23

qwen3-vl-plus-2025-09-23

MU1 x 4

USD29.704

USD14,367.128

Qwen VL-Max-2025-08-13

qwen-vl-max-2025-08-13

MU6 x 4

USD13.752

USD6,649.98

Qwen Omni

Model Name

Model Code

Model Unit Specification

Hourly Unit Price (USD)

Min Billing: Minute

Monthly Unit Price (USD)

Min Billing: Day

Qwen3.5-Omni-Flash

qwen3.5-omni-flash

MU8 x 1

USD6.464

USD3,080.477

MU9 x 1

USD7.014

USD3,383.024

Model type:

  • Instruct - The model performs inference in non-thinking mode after deployment.
  • Thinking - The model performs inference in thinking mode after deployment.

Deployment creation

DTU deployment creation

Before using DTU, enable the DTU feature in the console and apply for a resource quota. Once enabled, select the target model on the Dedicated Deployment page of the Bailian console and choose DTU as the billing method to create a deployment.

WarningDTU deployment does not currently support creation and management via API. Please complete enablement, creation, scaling, and renewal in the Bailian console.

For the basic workflow of general model deployment, see Dedicated Deployment Overview.

The form fields for creating a deployment are shown in the table below.

Parameter

Description

Required

Value Description

Service Name

Name of the deployment service

Yes

Custom

Model

Target model to deploy

Yes

Dropdown selection

Deployment Template

Deployment architecture

No

Dropdown selection, default to the first

Payment Type

Billing method

Yes

Pre-paid (monthly)

Input Throughput Quota

Purchased input TPM capacity

Yes

Integer multiple of baseline input TPM (kTPM)

Output Throughput Quota

Purchased output TPM capacity

Yes

Integer multiple of baseline output TPM (kTPM)

Purchase Duration

Purchase period

Yes

1-12 (integer, months)

Auto Renewal

Auto-renew on expiry

No

On/Off

Single Renewal Duration

Required when auto-renewal is enabled

Yes

1-12 (integer, months)

MU deployment configuration

Model Unit

Configuration item

Configuration details

Service name

A custom name for the deployment service.

Model

Select the model to deploy, including platform preset models and fine-tuned models.

Model unit type

Select the deployment specification. Different specifications correspond to different computing power and performance.

Replica count

Set the initial number of deployment replicas, which affects the concurrent processing capability of the service.

Deployment template

Select a deployment template (for example, "single-node deployment"). Different templates correspond to different resource configuration schemes. Available only in the model unit billing mode.

Model inference mode

For some models, when deployed inModel Unit mode, you can configure the inference mode, maximum context, and more.

  • Instruct - The model performs inference in non-thinking mode after deployment.

  • Thinking - The model performs inference in thinking mode after deployment.

Maximum context

TheModel Unit deployment mode of some models supports this setting. The maximum context length depends on the model type.

Service throttling

TheModel Unit deployment mode of some models supports this setting, which can limit the RPM and TPM of model calls.

Capacity Planning

DTU is sold by input/output TPM; the purchased TPM defines the maximum tokens processable per minute. Capacity planning aims to derive the purchase multiple from your business's peak token demand, then verify through load testing that the actually sustainable concurrency and latency meet requirements. Baseline input/output TPM and standard-workload performance references for each model are in the tables above.

Use a load-testing tool such as evalscope to test against your real business scenario (input/output length, concurrency, latency requirements), then determine the purchase multiple against the pricing table.

Capacity Planning Method

  1. Profile peak business metrics: typical request input/output token length, target concurrency, and requirements for first-token latency (TTFT) and per-token generation latency (TPOT).
  2. Estimate peak throughput demand: peak input TPM ≈ concurrency × per-request input length ÷ per-request processing time (minutes); peak output TPM ≈ concurrency × per-request output length ÷ per-request generation time (minutes).
  3. Compute the purchase multiple: input multiple = ⌈peak input TPM ÷ baseline input TPM⌉, output multiple = ⌈peak output TPM ÷ baseline output TPM⌉. DTU requires input and output to scale together, so take the larger of the two as the final multiple.
  4. Verify with load testing: after purchasing the computed multiple, re-test with evalscope under the target workload to confirm actual throughput and latency meet business requirements. If latency is high, raise concurrency within the TPM headroom to improve effective throughput.

If latency requirements are not strict, you can raise concurrency — within the purchased TPM cap — to increase actual throughput.

Business estimation example:

For a model with baseline input 656K TPM and baseline output 82K TPM: business peak profiling shows peak input ~1,300K TPM and peak output ~160K TPM. Compute multiples: input ⌈1300÷656⌉=2, output ⌈160÷82⌉=2; take the larger (2×), purchasing input 1,312K TPM (656K×2) + output 164K TPM (82K×2). Then load-test with evalscope at the peak workload to confirm latency passes. If business volume doubles (peak input ~2,600K, output ~320K), multiples become input ⌈2600÷656⌉=4, output ⌈320÷82⌉=4; purchase 4×: input 2,624K TPM (656K×4) + output 328K TPM (82K×4), and re-test.

Estimation formula:

Purchase multiple = max(⌈peak input TPM ÷ baseline input TPM⌉, ⌈peak output TPM ÷ baseline output TPM⌉)

Scaling and renewal/unsubscription

DTU scaling and renewal/unsubscription

The deployment list shows the purchased input/output TPM quotas; the details page shows the TPM capacity details.

Scaling: Click "Scale" in the deployment list to modify the input/output TPM capacity. Scaling up shows the additional amount due; scaling down shows the estimated refund. Input and output must be increased or decreased together; one cannot increase while the other decreases. Only running deployments can be operated.

Renewal: Pre-paid monthly deployments can be renewed upon expiry. The renewal button on the details page enters the renewal flow, supporting an auto-renewal toggle and a single renewal duration.

Unsubscription: Prepaid monthly orders support early termination; the used portion is settled at a 1.2x coefficient for refund. For details, see Refund rules for configuration downgrades. Unsubscription is irreversible; after unsubscription, dedicated resources are released and the service stops.

MU scaling and renewal

  • Model Unit (billed by duration): Click the Scaling button to self-service, manually adjust the number of instances (replicas). You can also configure an auto-scaling policy (including scaling thresholds, minimum/maximum replica count, scheduled scaling, etc.) via the scaling configuration button in the operation column.
  • Renewal: Prepaid monthly services can be renewed to extend the service period; auto-renewal is supported.
  • Unsubscription: Early termination within the first month of a prepaid purchase is billed at 1.2x the daily unit price; for details, see Refund rules for configuration downgrades.

FAQ

What is the difference between DTU deployment and Token pay-as-you-go billing?

Token pay-as-you-go bills by token usage on a shared resource pool; DTU uses dedicated GPU resources and bills by input/output TPM, suitable for scenarios requiring stable throughput and dedicated deployment.

What billing methods does DTU deployment support?

Only pre-paid (monthly) billing is supported.

How do I view the usage of purchased TPM?

You can view the purchased input/output TPM quotas in the deployment list and details page of Dedicated Deployment in the Bailian console. See Dedicated Deployment Overview.