Dedicated compute deployment offers two plans: Dedicated Throughput Unit (DTU) and Model Unit (MU), providing dedicated GPU resources for specified models. DTU targets newly released models (billed by input/output TPM x duration); MU is used for existing models (billed by model unit count x duration). This document introduces both plans' features, billing, usage flow, and supported models.
Overview
Dedicated compute deployment provides dedicated GPU resources and performance guarantees for specified models, supporting both base and custom model deployment. Bailian offers two dedicated compute plans:
- DTU (Dedicated Throughput Unit): Sold by Input/Output TPM, providing more direct throughput guarantees. Fully managed inference service with dedicated underlying GPU resources maintained by the platform.
- Model Unit (MU): Configures compute power by usage duration and model unit count, with dedicated resources. Supports custom performance metrics and PD disaggregation mode. Billing granularity is model unit count x duration.
DTU and MU are differentiated by model applicability: newly released models use DTU (metered by input/output TPM), while existing models continue to use MU (metered by model unit count). Both provide dedicated compute, resource isolation, and fine-tuned model deployment. For new deployments, choose DTU if the model supports it.
Shared capabilities of both plans:
- Dedicated GPU resources; the inference environment is physically isolated from other users.
- Supports base models and custom models (Bailian fine-tuned models or user-uploaded models).
- The platform maintains the underlying operations, no need to manage GPUs.
- No proactive RPM/TPM limits; traffic is bound by actual capacity.
For the general model deployment workflow, see Dedicated Deployment Overview and Provisioned Throughput.
Solution Selection
Bailian offers multiple capacity and billing plans for inference calls. DTU suits scenarios requiring dedicated deployment, fine-tuned model deployment, low latency with high concurrency, or data isolation. The plans are compared in the table below.
For a complete comparison of all billing methods including Token pay-as-you-go and PTU, see the Dedicated Deployment Overview main table.
Comparison Dimension | DTU | PTU | Token Pay-as-you-go | PAI/Lingjun |
|---|---|---|---|---|
Feature | Dedicated compute + fine-tuned models | Throughput guarantee + low latency | Elastic and zero-threshold | Custom runtime |
Resource isolation | Physical isolation | Logical isolation | Shared pool | Physical isolation |
Fine-tuned models | Full-parameter + LoRA | Not supported | LoRA | Supported |
Custom framework | Not supported | Not supported | Not supported | Supported |
Billing mode | TPM quota × duration | TPM quota × duration | Token usage | GPU × duration |
Billing rules
DTU billing
DTU is billed separately by input TPM and output TPM, using a pre-paid (monthly) model. Each model has fixed input/output baseline TPM (see table below), and you must purchase in integer multiples of the baseline TPM. Purchase at least 1× baseline TPM for both input and output; for production, 2× or above is recommended for each.
Fees are calculated by the backend based on model, input/output TPM, service region, and purchase duration. The input/output unit prices and monthly price for each model are shown in the table below. Unit prices are billed per kTPM·month, and the monthly price is the total for 1× baseline TPM of both input and output. Fine-tuned model prices are the same as the corresponding base model, subject to the actual console display.
MU billing
Cost = Usage duration (hours) × Number of Model Units × Model Unit unit price
The "Model Unit unit price" takes the "Hourly unit price" column in the table below for postpaid scenarios; for monthly subscription prepaid billing, the formula becomes Number of monthly subscriptions × Number of Model Units × Monthly unit price.
- For the first month of a prepaid purchase, if you cancel the subscription early within the first month, the daily unit price (≈ Monthly unit price / 30) will be billed at 1.2 times the rate (less than one day is billed as one day)
NoteThe compute resources for the Model Unit postpaid method are first-come, first-served. If the purchase fails, a full refund will be issued.
Supported models and pricing
DTU pricing and performance benchmarks
Singapore
Model | Baseline Input (TPM) | Baseline Output (TPM) | Input Unit Price (USD/kTPM·mo) | Output Unit Price (USD/kTPM·mo) | Monthly Price (USD) |
|---|---|---|---|---|---|
qwen3.7-plus-2026-05-26 | 1,192,000 | 148,000 | 37 | 295 | 87,764 |
1,372,000 | 170,000 | 32 | 257 | 87,594 | |
660,000 | 124,000 | 66 | 352 | 87,208 | |
704,000 | 132,000 | 62 | 331 | 87,340 |
The performance reference data for each model under standard workload is shown in the table below, subject to the actual console display.
The following performance reference data was measured at a 0% cache hit rate. In actual use, as the cache hit rate increases, model performance improves accordingly.
Model | Input Length | Output Length | Cache Hit Rate | First Token Latency (ms) | Per-Token Latency (ms) |
|---|---|---|---|---|---|
qwen3.7-plus-2026-05-26 | 16,000 | 2,000 | 0 | 888 | 10 |
16,000 | 2,000 | 0 | 2,418 | 15 | |
2,600 | 500 | 0 | 819 | 14 | |
2,600 | 500 | 0 | 970 | 13 |
Among these, qwen3.8-27b currently requires whitelist access, while the other models are on the Singapore official site.
MU pricing
Singapore
Text Generation
Model | Baseline Input (TPM) | Baseline Output (TPM) | Input Unit Price (USD/kTPM·Month) | Output Unit Price (USD/kTPM·Month) | Monthly Total Price (USD) |
|---|---|---|---|---|---|
qwen3.7-plus-2026-05-26 | 1,192,000 | 148,000 | 37 | 295 | 87,764 |
1,372,000 | 170,000 | 32 | 257 | 87,594 | |
660,000 | 124,000 | 66 | 352 | 87,208 | |
704,000 | 132,000 | 62 | 331 | 87,340 |
Multimodal
Model | Input Length | Output Length | Cache Hit Rate | First Token Latency (ms) | Per-token Latency (ms) |
|---|---|---|---|---|---|
qwen3.7-plus-2026-05-26 | 16,000 | 2,000 | 0 | 888 | 10 |
16,000 | 2,000 | 0 | 2,418 | 15 | |
2,600 | 500 | 0 | 819 | 14 | |
2,600 | 500 | 0 | 970 | 13 |
Model type:
- Instruct - After the model is deployed, inference runs in non-thinking mode.
China (Beijing)
Text generation
Qwen
Model Name | Model Code | Model Unit Specification | Hourly Unit Price (USD) Min Billing: Minute | Monthly Unit Price (USD) Min Billing: Day |
|---|---|---|---|---|
Qwen3.8-27B | qwen3.8-27b | MU9 x 4 | USD28.056 | USD13,532.096 |
Qwen3.7-Plus-2026-05-26 | qwen3.7-plus-2026-05-26 | MU2 x 8 | USD69.312 | USD33,044.72 |
MU3 x 8 | USD150.72 | USD72,577.152 | ||
Qwen3.6-35B-A3B | qwen3.6-35b-a3b | MU1 x 8 | USD59.408 | USD28,734.256 |
MU2 x 8 | USD69.312 | USD33,044.72 | ||
MU3 x 8 | USD150.72 | USD72,577.152 | ||
MU8 x 1 | USD6.464 | USD3,080.477 | ||
MU9 x 1 | USD7.014 | USD3,383.024 | ||
Qwen3.6-27B | qwen3.6-27b | MU9 x 1 | USD7.014 | USD3,383.024 |
Qwen3.6-Flash-2026-04-16 | qwen3.6-flash-2026-04-16 | MU1 x 2 | USD14.852 | USD7,183.564 |
MU3 x 8 | USD150.72 | USD72,577.152 | ||
Qwen3.6-Plus-2026-04-02 | qwen3.6-plus-2026-04-02 | MU1 x 8 MU1 x 16(PD Disaggregation Mode) | USD59.408 PD Disaggregation Mode: USD118.816 | USD28,734.256 PD Disaggregation Mode: USD57,468.512 |
MU2 x 8 | USD69.312 | USD33,044.72 | ||
Qwen3.5-397B-A17B | qwen3.5-397b-a17b | MU3 x 8 MU3 x 16(PD Disaggregation Mode) | USD150.72 PD Disaggregation Mode: USD301.44 | USD72,577.152 PD Disaggregation Mode: USD145,154.304 |
MU6 x 16 | USD55.008 | USD26,599.92 | ||
Qwen3.5-122B-A10B | qwen3.5-122b-a10b | MU1 x 4 | USD29.704 | USD14,367.128 |
MU6 x 16 | USD55.008 | USD26,599.92 | ||
Qwen3.5-35B-A3B | qwen3.5-35b-a3b | MU1 x 2 | USD14.852 | USD7,183.564 |
MU2 x 8 | USD69.312 | USD33,044.72 | ||
MU3 x 8 | USD150.72 | USD72,577.152 | ||
MU9 x 1 | USD7.014 | USD3,383.024 | ||
Qwen3.5-27B | qwen3.5-27b | MU2 x 8 | USD69.312 | USD33,044.72 |
MU3 x 8 | USD150.72 | USD72,577.152 | ||
MU8 x 1 | USD6.464 | USD3,080.477 | ||
MU9 x 1 | USD7.014 | USD3,383.024 | ||
Qwen3.5-9B | qwen3.5-9b | MU1 x 2 | USD14.852 | USD7,183.564 |
MU2 x 8 | USD69.312 | USD33,044.72 | ||
Qwen3.5-Flash-2026-02-23 | qwen3.5-flash-2026-02-23 | MU1 x 2 | USD14.852 | USD7,183.564 |
Qwen3.5-Plus-2026-02-15 | qwen3.5-plus-2026-02-15 | MU1 x 8 MU1 x 16(PD Disaggregation Mode) | USD59.408 PD Disaggregation Mode: USD118.816 | USD28,734.256 PD Disaggregation Mode: USD57,468.512 |
MU2 x 8 | USD69.312 | USD33,044.72 | ||
MU3 x 8 MU3 x 16(PD Disaggregation Mode) | USD150.72 PD Disaggregation Mode: USD301.44 | USD72,577.152 PD Disaggregation Mode: USD145,154.304 | ||
Qwen3-235B-A22B-Instruct-2507 | qwen3-235b-a22b-instruct-2507 | MU1 x 4 | USD29.704 | USD14,367.128 |
MU2 x 8 | USD69.312 | USD33,044.72 | ||
MU3 x 8 | USD150.72 | USD72,577.152 | ||
Qwen3-32B | qwen3-32b | MU6 x 16 | USD55.008 | USD26,599.92 |
Qwen3-30B-A3B-Thinking-2507 | qwen3-30b-a3b-thinking-2507 | MU1 x 2 | USD14.852 | USD7,183.564 |
Qwen3-4B | qwen3-4b | MU1 x 2 | USD14.852 | USD7,183.564 |
MU5 x 1 | USD2.888 | USD1,394.329 | ||
Qwen3-Embedding-0.6B | qwen3-embedding-0.6b | MU5 x 1 | USD2.888 | USD1,394.329 |
MU6 x 1 | USD3.438 | USD1,662.495 | ||
Qwen3-MoE-Rerank-0.6B | qwen3-moe-rerank-0.6b | MU5 x 1 | USD2.888 | USD1,394.329 |
Qwen3-Rerank-0.6B | qwen3-rerank-0.6b | MU5 x 1 | USD2.888 | USD1,394.329 |
MU6 x 1 | USD3.438 | USD1,662.495 | ||
Qwen3-Max-2025-09-23 | qwen3-max-2025-09-23 | MU2 x 8 | USD69.312 | USD33,044.72 |
MU3 x 8 | USD150.72 | USD72,577.152 | ||
Qwen3-Rerank | qwen3-rerank | MU5 x 1 | USD2.888 | USD1,394.329 |
Qwen2.5-Open Source-72B | qwen2.5-72b-instruct | MU1 x 8 | USD59.408 | USD28,734.256 |
Qwen2.5-Open Source-14B | qwen2.5-14b-instruct | MU1 x 2 | USD14.852 | USD7,183.564 |
Qwen2.5-Open Source-7B | qwen2.5-7b-instruct | MU1 x 2 | USD14.852 | USD7,183.564 |
MU5 x 1 | USD2.888 | USD1,394.329 | ||
Qwen-Plus-2025-07-28 | qwen-plus-2025-07-28 | MU1 x 4 MU1 x 16(PD Disaggregation Mode) | USD29.704 PD Disaggregation Mode: USD118.816 | USD14,367.128 PD Disaggregation Mode: USD57,468.512 |
Qwen-Plus-2025-12-01 | qwen-plus-2025-12-01 | MU1 x 4 | USD29.704 | USD14,367.128 |
Qwen-Plus-Character-2025-11-06 | qwen-plus-character-2025-11-06 | MU1 x 4 | USD29.704 | USD14,367.128 |
GLM
Model Name | Model Code | Model Unit Specification | Hourly Unit Price (USD) Min Billing: Minute | Monthly Unit Price (USD) Min Billing: Day |
|---|---|---|---|---|
GLM-5.1 | glm-5.1 | MU2 x 8 | USD69.312 | USD33,044.72 |
MU3 x 16(PD Disaggregation Mode) | PD Disaggregation Mode: USD301.44 | PD Disaggregation Mode: USD145,154.304 | ||
MU6 x 16 | USD55.008 | USD26,599.92 | ||
GLM-5 | glm-5 | MU3 x 16(PD Disaggregation Mode) | PD Disaggregation Mode: USD301.44 | PD Disaggregation Mode: USD145,154.304 |
GLM-4.7 | glm-4.7 | MU6 x 32(PD Disaggregation Mode) | PD Disaggregation Mode: USD110.016 | PD Disaggregation Mode: USD53,199.84 |
DeepSeek
Model Name | Model Code | Model Unit Specification | Hourly Unit Price (USD) Min Billing: Minute | Monthly Unit Price (USD) Min Billing: Day |
|---|---|---|---|---|
DeepSeek-v4-Flash | deepseek-v4-flash | MU1 x 8 | USD59.408 | USD28,734.256 |
MU3 x 8 | USD150.72 | USD72,577.152 | ||
DeepSeek-v3.2 | deepseek-v3.2 | MU2 x 16(PD Disaggregation Mode) | PD Disaggregation Mode: USD138.624 | PD Disaggregation Mode: USD66,089.44 |
Other models
Model Name | Model Code | Model Unit Specification | Hourly Unit Price (USD) Min Billing: Minute | Monthly Unit Price (USD) Min Billing: Day |
|---|---|---|---|---|
Kimi-K2.5 | kimi-k2.5 | MU2 x 8 | USD69.312 | USD33,044.72 |
Multimodal
Qwen VL
Model Name | Model Code | Model Unit Specification | Hourly Unit Price (USD) Min Billing: Minute | Monthly Unit Price (USD) Min Billing: Day |
|---|---|---|---|---|
Qwen3-VL-235B-A22B-Thinking | qwen3-vl-235b-a22b-thinking | MU1 x 8 | USD59.408 | USD28,734.256 |
MU2 x 8 | USD69.312 | USD33,044.72 | ||
MU3 x 8 | USD150.72 | USD72,577.152 | ||
Qwen3-VL-32B-Instruct | qwen3-vl-32b-instruct | MU2 x 8 | USD69.312 | USD33,044.72 |
MU3 x 8 | USD150.72 | USD72,577.152 | ||
Qwen3-VL-8B-Instruct | qwen3-vl-8b-instruct | MU1 x 2 | USD14.852 | USD7,183.564 |
MU5 x 1 | USD2.888 | USD1,394.329 | ||
Qwen3-VL-4B-Instruct | qwen3-vl-4b-instruct | MU1 x 2 | USD14.852 | USD7,183.564 |
Qwen3-VL-2B-Instruct | qwen3-vl-2b-instruct | MU5 x 1 | USD2.888 | USD1,394.329 |
Qwen3-VL-Embedding-2B | qwen3-vl-embedding-2b | MU5 x 1 | USD2.888 | USD1,394.329 |
Qwen3-VL-Flash-2025-10-15 | qwen3-vl-flash-2025-10-15 | MU1 x 4 | USD29.704 | USD14,367.128 |
Qwen3-VL-Plus-2025-09-23 | qwen3-vl-plus-2025-09-23 | MU1 x 4 | USD29.704 | USD14,367.128 |
Qwen VL-Max-2025-08-13 | qwen-vl-max-2025-08-13 | MU6 x 4 | USD13.752 | USD6,649.98 |
Qwen Omni
Model Name | Model Code | Model Unit Specification | Hourly Unit Price (USD) Min Billing: Minute | Monthly Unit Price (USD) Min Billing: Day |
|---|---|---|---|---|
Qwen3.5-Omni-Flash | qwen3.5-omni-flash | MU8 x 1 | USD6.464 | USD3,080.477 |
MU9 x 1 | USD7.014 | USD3,383.024 |
Model type:
- Instruct - The model performs inference in non-thinking mode after deployment.
- Thinking - The model performs inference in thinking mode after deployment.
Deployment creation
DTU deployment creation
Before using DTU, enable the DTU feature in the console and apply for a resource quota. Once enabled, select the target model on the Dedicated Deployment page of the Bailian console and choose DTU as the billing method to create a deployment.
WarningDTU deployment does not currently support creation and management via API. Please complete enablement, creation, scaling, and renewal in the Bailian console.
For the basic workflow of general model deployment, see Dedicated Deployment Overview.
The form fields for creating a deployment are shown in the table below.
Parameter | Description | Required | Value Description |
|---|---|---|---|
Service Name | Name of the deployment service | Yes | Custom |
Model | Target model to deploy | Yes | Dropdown selection |
Deployment Template | Deployment architecture | No | Dropdown selection, default to the first |
Payment Type | Billing method | Yes | Pre-paid (monthly) |
Input Throughput Quota | Purchased input TPM capacity | Yes | Integer multiple of baseline input TPM (kTPM) |
Output Throughput Quota | Purchased output TPM capacity | Yes | Integer multiple of baseline output TPM (kTPM) |
Purchase Duration | Purchase period | Yes | 1-12 (integer, months) |
Auto Renewal | Auto-renew on expiry | No | On/Off |
Single Renewal Duration | Required when auto-renewal is enabled | Yes | 1-12 (integer, months) |
MU deployment configuration
Model Unit
Configuration item | Configuration details |
|---|---|
Service name | A custom name for the deployment service. |
Model | Select the model to deploy, including platform preset models and fine-tuned models. |
Model unit type | Select the deployment specification. Different specifications correspond to different computing power and performance. |
Replica count | Set the initial number of deployment replicas, which affects the concurrent processing capability of the service. |
Deployment template | Select a deployment template (for example, "single-node deployment"). Different templates correspond to different resource configuration schemes. Available only in the model unit billing mode. |
Model inference mode | For some models, when deployed inModel Unit mode, you can configure the inference mode, maximum context, and more.
|
Maximum context | TheModel Unit deployment mode of some models supports this setting. The maximum context length depends on the model type. |
Service throttling | TheModel Unit deployment mode of some models supports this setting, which can limit the RPM and TPM of model calls. |
Capacity Planning
DTU is sold by input/output TPM; the purchased TPM defines the maximum tokens processable per minute. Capacity planning aims to derive the purchase multiple from your business's peak token demand, then verify through load testing that the actually sustainable concurrency and latency meet requirements. Baseline input/output TPM and standard-workload performance references for each model are in the tables above.
Use a load-testing tool such as evalscope to test against your real business scenario (input/output length, concurrency, latency requirements), then determine the purchase multiple against the pricing table.
Capacity Planning Method
- Profile peak business metrics: typical request input/output token length, target concurrency, and requirements for first-token latency (TTFT) and per-token generation latency (TPOT).
- Estimate peak throughput demand: peak input TPM ≈ concurrency × per-request input length ÷ per-request processing time (minutes); peak output TPM ≈ concurrency × per-request output length ÷ per-request generation time (minutes).
- Compute the purchase multiple: input multiple = ⌈peak input TPM ÷ baseline input TPM⌉, output multiple = ⌈peak output TPM ÷ baseline output TPM⌉. DTU requires input and output to scale together, so take the larger of the two as the final multiple.
- Verify with load testing: after purchasing the computed multiple, re-test with evalscope under the target workload to confirm actual throughput and latency meet business requirements. If latency is high, raise concurrency within the TPM headroom to improve effective throughput.
If latency requirements are not strict, you can raise concurrency — within the purchased TPM cap — to increase actual throughput.
Business estimation example:
For a model with baseline input 656K TPM and baseline output 82K TPM: business peak profiling shows peak input ~1,300K TPM and peak output ~160K TPM. Compute multiples: input ⌈1300÷656⌉=2, output ⌈160÷82⌉=2; take the larger (2×), purchasing input 1,312K TPM (656K×2) + output 164K TPM (82K×2). Then load-test with evalscope at the peak workload to confirm latency passes. If business volume doubles (peak input ~2,600K, output ~320K), multiples become input ⌈2600÷656⌉=4, output ⌈320÷82⌉=4; purchase 4×: input 2,624K TPM (656K×4) + output 328K TPM (82K×4), and re-test.
Estimation formula:
Purchase multiple = max(⌈peak input TPM ÷ baseline input TPM⌉, ⌈peak output TPM ÷ baseline output TPM⌉)
Scaling and renewal/unsubscription
DTU scaling and renewal/unsubscription
The deployment list shows the purchased input/output TPM quotas; the details page shows the TPM capacity details.
Scaling: Click "Scale" in the deployment list to modify the input/output TPM capacity. Scaling up shows the additional amount due; scaling down shows the estimated refund. Input and output must be increased or decreased together; one cannot increase while the other decreases. Only running deployments can be operated.
Renewal: Pre-paid monthly deployments can be renewed upon expiry. The renewal button on the details page enters the renewal flow, supporting an auto-renewal toggle and a single renewal duration.
Unsubscription: Prepaid monthly orders support early termination; the used portion is settled at a 1.2x coefficient for refund. For details, see Refund rules for configuration downgrades. Unsubscription is irreversible; after unsubscription, dedicated resources are released and the service stops.
MU scaling and renewal
- Model Unit (billed by duration): Click the Scaling button to self-service, manually adjust the number of instances (replicas). You can also configure an auto-scaling policy (including scaling thresholds, minimum/maximum replica count, scheduled scaling, etc.) via the scaling configuration button in the operation column.
- Renewal: Prepaid monthly services can be renewed to extend the service period; auto-renewal is supported.
- Unsubscription: Early termination within the first month of a prepaid purchase is billed at 1.2x the daily unit price; for details, see Refund rules for configuration downgrades.
FAQ
What is the difference between DTU deployment and Token pay-as-you-go billing?Token pay-as-you-go bills by token usage on a shared resource pool; DTU uses dedicated GPU resources and bills by input/output TPM, suitable for scenarios requiring stable throughput and dedicated deployment.
What billing methods does DTU deployment support?Only pre-paid (monthly) billing is supported.
How do I view the usage of purchased TPM?You can view the purchased input/output TPM quotas in the deployment list and details page of Dedicated Deployment in the Bailian console. See Dedicated Deployment Overview.