Create a model deployment task.
Prerequisites
- Read Introduction to model deployment and Deploy a model by using the API to understand how model deployment works and the basic workflow on Alibaba Cloud Model Studio.
- Configure your API key for Model Studio. For details, see Get an API key.
Model deployment
Endpoint
POST https://dashscope-intl.aliyuncs.com/api/v1/deployments
Request examples
Billing by provisioned throughput units (PTU)
NoteAfter running the deployment command below, billing starts immediately upon successful deployment—even if you have not yet called the model. Confirm the billing rules before deploying.
The PTU billing mode charges based on the duration of provisioned throughput usage. It suits scenarios requiring stable throughput guarantees, high concurrency, low latency, and predictable traffic. In this mode, throughput/concurrency and generation speed are preset by the platform and cannot be adjusted.
curl "https://dashscope-intl.aliyuncs.com/api/v1/deployments" \
--header "Authorization: Bearer $DASHSCOPE_API_KEY" \
--header 'Content-Type: application/json' \
--data '{
"name": "my_qwen_flash",
"model_name": "qwen-flash-2025-07-28",
"plan": "ptu",
"ptu_capacity": {
"input_tpm": 10000,
"output_tpm": 1000
}
}'
Billing by model unit usage duration
Note
- After running the deployment command below, billing starts immediately upon successful deployment—even if you have not yet called the model. Confirm the billing rules before deploying.
- Model unit pay-as-you-go computing resources are allocated on a first-come, first-served basis. If purchase fails, you receive a full refund.
Select the model unit billing method. This mode charges based on model unit usage duration and suits large-scale inference workloads after model fine-tuning. Resources are dedicated, and performance and cost are flexible. Throughput/concurrency and generation speed are customizable.
curl "https://dashscope-intl.aliyuncs.com/api/v1/deployments" \
--header "Authorization: Bearer $DASHSCOPE_API_KEY" \
--header 'Content-Type: application/json' \
--data '{
"name": "my_qwen_plus",
"model_name": "qwen-plus-2025-12-01",
"plan": "mu",
"deploy_spec": "MU1",
"enable_thinking": true,
"capacity": 4,
"max_context_length": 10000,
"rpm_limit": 500,
"tpm_limit": 1000
}'
The model unit deployment mode supports additional settings:
Configuration item | Configuration details |
|---|---|
Service name | Custom name for the deployment service. |
Select model | Select the model to deploy, including platform built-in models and fine-tuned models. |
Model unit type | Select the deployment specification. Different specifications correspond to different compute power and performance. |
Deployment replica count | Set the initial number of deployment replicas, which affects the concurrency capacity of the service. |
Deployment template | Select a deployment template (such as "single-node deployment"). Different templates correspond to different resource configuration schemes. Available only under the model unit billing mode. |
Configure model inference mode | For some models deployed in Model Unit mode, you can configure the inference mode, maximum context length, and more.
|
Maximum context length | The Model Unit deployment mode of some models supports this setting. The maximum context length depends on the model type. |
Service rate limiting | The Model Unit deployment mode of some models supports this setting. You can limit the RPM and TPM of model invocations. |
For details on setting these options via API, see Create a model deployment task using the API.
Billing by token usage
Select the token-based billing method. This mode charges based on token usage and suits cost-sensitive scenarios with low requirements for concurrency and latency. It offers the highest price advantage. Throughput/concurrency and generation speed are preset by the platform and cannot be adjusted.
curl "https://dashscope-intl.aliyuncs.com/api/v1/deployments" \
--header "Authorization: Bearer $DASHSCOPE_API_KEY" \
--header 'Content-Type: application/json' \
--data '{
"model_name": "qwen3-8b-ft-202511132025-0260",
"plan": "lora",
"capacity": 1,
"name": "qwen3-8b-ft"
}'
The capacity parameter has no effect but must be included. To scale up or down, go to the Model Studio Dedicated Deployment console and submit a form request.
Request parameters
Parameter | Type | Location | Required | Description | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
model_name | String | body | Yes | The name of the model to deploy. This corresponds to the model ID in My Models. You can also get this ID from the output of the Create Training Job or Create Import Job operations. | |||||||||||
plan | String | body | Yes | The deployment plan. The following billing methods are supported:
You can quickly find the supported deployment plans for a fine-tuned model in My Models. Fine-tuned CosyVoice models currently only support | |||||||||||
name | String | body | Yes | The display name of the model in the console. | |||||||||||
capacity | Integer | body | No | Required only when CosyVoice models currently provide the following two deployment templates with corresponding
| |||||||||||
billing_method | String | body | No | Required only when | |||||||||||
deploy_spec | String | body | No | This setting is applicable only when For details about feature support, see Feature support for model unit deployment. | This parameter is required when You can get this value from the | ||||||||||
enable_thinking | Boolean | body | No | Supported by some models. You can set this to | |||||||||||
max_context_length | Number | body | No | Supported by some models. Example: | |||||||||||
rpm_limit | Number | body | No | Supported by some models. Specifies the maximum number of requests per minute (RPM). | |||||||||||
tpm_limit | Number | body | No | Supported by some models. Specifies the maximum number of tokens per minute (TPM). | |||||||||||
ptu_capacity | Object | body | No | This setting is applicable only when For details about feature support, see Feature Support for PTU Deployment. If you do not specify this parameter, the system defaults to | Example: Example: | ||||||||||
ptu_capacity.input_tpm | Number | body | No | Supported by all models. Specifies the maximum number of input tokens per minute (TPM). | |||||||||||
ptu_capacity.output_tpm | Number | body | No | Supported by all models. Specifies the maximum number of output tokens per minute (TPM). | |||||||||||
ptu_capacity.thinking_output_tpm | Number | body | No | Supported by some models. Specifies the maximum number of provisioned thinking output tokens per minute (TPM). | |||||||||||
suffix | String | body | No | After a model is deployed, a new model name is generated. The suffix parameter specifies the suffix for this new name. It must be globally unique and have a maximum length of 8 characters. You can omit the suffix for the first deployment of a model. If you deploy the same model multiple times, you must specify a unique suffix for each deployment. See the deployed_model output parameter for more information. | |||||||||||
Supported models
View supported features and billing.
Usage duration billing (Provisioned Throughput)
Cost = Usage duration × (Input TPM unit price × Input TPM + Output TPM unit price × Output TPM)
Postpaid is calculated by hour: the usage duration unit is hours, and the unit price takes the "1 hour continuous" column in the table below; prepaid is calculated by day: the usage duration unit is days, and the unit price takes the "1 day continuous" column in the table below.
- Prepaid orders take effect immediately after payment, valid for N days until 23:59 on day N. If ordered after 22:00, the expiration date is automatically extended by 1 day.
- After a prepaid order expires, the service will be stopped with a 2-hour delay, and resources will be retained for 14 hours after stopping before being released.
- Prepaid orders cannot be terminated early.
- For post-paid billing, if your account is in arrears, the deployed resources will be retained and billed for 24 hours, during which the service can still be used normally. After 24 hours, the system stops billing, the model deployment enters an arrears state, and the underlying resources will be deleted, but the model deployment task will be retained. After you pay off the arrears, the system will reallocate resources and resume usage (fees will continue to accrue after resumption). If you do not want to continue incurring charges, you can delete the model deployment task; once deleted successfully, billing will stop.
When the model input exceeds the maximum input Token, the relevant call will automatically switch to the pay-as-you-go mode of the current model; when the purchased TPM is exceeded, it is handled according to the overflow policy selected at creation ("Auto-overflow" switches to pay-as-you-go, "PTU capacity only" returns 429). At this time, inference performance may degrade and will be governed by the public traffic of the current snapshot model in the workspace, and the fee is charged according to the model invocation (pay-as-you-go) standard.
- At this time (only under the "Auto-overflow" policy), the call API response Header will include:
x-dashscope-ptu-overflow:true. - For TPM statistics, go to: Model Monitoring.
For specific fee-reduction and refund rules for scale-in (downgrade) scenarios, please refer to: Refund rules for configuration downgrades.
NotePTU deployment supports stepped capacity coefficients and cache discounts for long inputs. For details, see Provisioned Throughput long input and cache.
Singapore
Qwen
Model name | Model code | Max input tokens | Postpaid input Per 10K TPM/hour | Postpaid output Per 1K TPM/hour | Prepaid input Per 10K TPM/day | Prepaid output Per 1K TPM/day |
|---|---|---|---|---|---|---|
Qwen3.8-Max | qwen3.8-max | 1M | $4.8 | $1.44 | $57.6 | $17.28 |
Qwen3.7-Max-2026-05-20 | qwen3.7-max-2026-05-20 | 256K | $1.92 | $1.8 | $72 | $21.6 |
Qwen3.7-Plus-2026-05-26 | qwen3.7-plus-2026-05-26 | 256K | $0.96 | $0.384 | $11.52 | $4.608 |
Qwen3.6-Plus-2026-04-02 | qwen3.6-plus-2026-04-02 | 128K | $1.2 | $0.72 | $14.4 | $8.64 |
Qwen3.5-Plus-2026-04-20 | qwen3.5-plus-2026-04-20 | 128K | $0.96 | $0.576 | $11.52 | $6.912 |
DeepSeek
Model name | Model code | Max input tokens | Postpaid input Per 10K TPM/hour | Postpaid output Per 1K TPM/hour | Prepaid input Per 10K TPM/day | Prepaid output Per 1K TPM/day |
|---|---|---|---|---|---|---|
DeepSeek-v4-Flash | deepseek-v4-flash | 256K | $0.72 | $0.144 | $8.64 | $1.728 |
DeepSeek-v4-Pro | deepseek-v4-pro | 256K | $0.96 | $1.728 | $103.68 | $20.736 |
DeepSeek-v3.2 | deepseek-v3.2 | 64K | $2.05 | $0.616 | $24.62 | $7.387 |
GLM
Model name | Model code | Max input tokens | Postpaid input Per 10K TPM/hour | Postpaid output Per 1K TPM/hour | Prepaid input Per 10K TPM/day | Prepaid output Per 1K TPM/day |
|---|---|---|---|---|---|---|
GLM-5.2 | glm-5.2 | 1M | $5.04 | $1.584 | $60.48 | $19.008 |
GLM-5.1 | glm-5.1 | 64K | $5.04 | $1.584 | $64.8 | $19.008 |
Qwen-VL
Model name | Model code | Max input tokens | Postpaid input Per 10K TPM/hour | Postpaid output Per 1K TPM/hour | Prepaid input Per 10K TPM/day | Prepaid output Per 1K TPM/day |
|---|---|---|---|---|---|---|
Qwen3-VL-Plus-2025-09-23 | qwen3-vl-plus-2025-09-23 | 128K | $0.48 | $0.384 | $5.76 | $4.608 |
North China 2 (Beijing)
Qwen
Model name | Model code | Max input tokens | Postpaid input Per 10K TPM/hour | Postpaid output Per 1K TPM/hour | Subscription input Per 10K TPM/day | Subscription output Per 1K TPM/day |
|---|---|---|---|---|---|---|
Qwen3.8-Max | qwen3.8-max | 1M | $3.96 | $1.188 | $47.53 | $14.258 |
Qwen3.7-Max-2026-05-20 | qwen3.7-max-2026-05-20 | 256K | $3.96 | $1.188 | $47.53 | $14.258 |
Qwen3.7-Plus-2026-05-26 | qwen3.7-plus-2026-05-26 | 256K | $0.66 | $0.264 | $7.92 | $3.168 |
Qwen3.6-Plus-2026-04-02 | qwen3.6-plus-2026-04-02 | 128K | $0.67 | $0.066 | $7.93 | $4.753 |
Qwen3.5-Plus-2026-04-20 | qwen3.5-plus-2026-04-20 | 128K | $0.26 | $0.16 | $3.17 | $1.9 |
Qwen3-Max-2025-09-23 | qwen3-max-2025-09-23 | 128K | $1.11 | $0.45 | $13.32 | $5.4 |
Qwen-Flash-2025-07-28 | qwen-flash-2025-07-28 | 128K | $0.06 | $0.06 | $0.72 | $0.72 |
Qwen-Plus-2025-12-01 | qwen-plus-2025-12-01 | 128K | $0.28 | Non-thinking: $0.07 Thinking: $0.28 | $3.36 | Non-thinking: $0.84 Thinking: $3.36 |
DeepSeek
Model name | Model code | Max input tokens | Postpaid input Per 10K TPM/hour | Postpaid output Per 1K TPM/hour | Subscription input Per 10K TPM/day | Subscription output Per 1K TPM/day |
|---|---|---|---|---|---|---|
DeepSeek-v4-Flash | deepseek-v4-flash | 256K | $0.5 | $0.099 | $5.94 | $1.188 |
DeepSeek-v4-Pro | deepseek-v4-pro | 256K | $5.94 | $1.188 | $71.3 | $14.26 |
DeepSeek-v3.2 | deepseek-v3.2 | 64K | $1.04 | $0.16 | $12.48 | $1.92 |
DeepSeek-v3 | deepseek-v3 | 64K | $0.99 | $0.396 | $11.9 | $4.75 |
GLM
Model name | Model code | Max input tokens | Postpaid input Per 10K TPM/hour | Postpaid output Per 1K TPM/hour | Subscription input Per 10K TPM/day | Subscription output Per 1K TPM/day |
|---|---|---|---|---|---|---|
GLM-5.2 | glm-5.2 | 1M | $3.96 | $1.386 | $47.53 | $16.635 |
GLM-5.1 | glm-5.1 | 64K | $2.97 | $1.19 | $35.65 | $14.26 |
Qwen-VL
Model name | Model code | Max input tokens | Postpaid input Per 10K TPM/hour | Postpaid output Per 1K TPM/hour | Subscription input Per 10K TPM/day | Subscription output Per 1K TPM/day |
|---|---|---|---|---|---|---|
Qwen3-VL-Plus-2025-09-23 | qwen3-vl-plus-2025-09-23 | 128K | $0.35 | $0.35 | $4.2 | $4.2 |
Billing by usage duration (model unit)
Cost = Usage duration (hours) × Number of model units × Model unit price
"Model unit price" takes the "Hourly unit price" column in the table below for pay-as-you-go scenarios; for monthly prepaid billing, the formula becomes Number of months × Number of model units × Monthly unit price.
- For the first month of a prepaid purchase, if you cancel early within the first month, the daily unit price (≈ Monthly unit price / 30) is billed at 1.2× (less than one day is billed as one day)
NoteCompute resources under the model unit pay-as-you-go method are first-come, first-served. If the purchase fails, a full refund is issued.
Singapore
Text generation
Model name | Model code | Model unit specification | Hourly unit price (USD) Minimum billing: minute | Monthly unit price (USD) Minimum billing: day |
|---|---|---|---|---|
Qwen3.6-Plus-2026-04-02 | qwen3.6-plus-2026-04-02 | MU1 x 8 | $88 | $41,832 |
Qwen3.5-397B-A17B | qwen3.5-397b-a17b | MU2 x 8 | $112 | $52,392 |
Qwen3.5-122B-A10B | qwen3.5-122b-a10b | MU2 x 8 | $112 | $52,392 |
Qwen3.5-27B | qwen3.5-27b | MU1 x 2 | $22 | $10,458 |
Qwen3.5-9B | qwen3.5-9b | MU1 x 2 | $22 | $10,458 |
GLM-5.1 | glm-5.1 | MU2 x 8 | $112 | $52,392 |
MU3 x 8 | $216 | $102,696 | ||
DeepSeek-v4-Flash | deepseek-v4-flash | MU2 x 8 | $112 | $52,392 |
Qwen-Plus-Character-2025-11-06 | qwen-plus-character-2025-11-06 | MU1 x 4 | $44 | $20,916 |
Multimodal
Model name | Model code | Model unit specification | Hourly unit price (USD) Minimum billing: minute | Monthly unit price (USD) Minimum billing: day |
|---|---|---|---|---|
Qwen3-VL-32B-Instruct | qwen3-vl-32b-instruct | MU2 x 8 | $112 | $52,392 |
Model type:
- Instruct - After model deployment, inference is performed in non-thinking mode.
North China 2 (Beijing)
Text generation
Tab
Qwen
Model name | Model code | Model unit specification | Hourly unit price (USD) Minimum billing: minute | Monthly unit price (USD) Minimum billing: day |
|---|---|---|---|---|
Qwen3.6-35B-A3B | qwen3.6-35b-a3b | MU1 x 8 | $59.408 | $28,734.256 |
MU2 x 8 | $69.312 | $33,044.72 | ||
MU3 x 8 | $150.72 | $72,577.152 | ||
MU9 x 1 | $7.014 | $3,383.024 | ||
Qwen3.6-27B | qwen3.6-27b | MU9 x 1 | $7.014 | $3,383.024 |
Qwen3.6-Flash-2026-04-16 | qwen3.6-flash-2026-04-16 | MU1 x 2 | $14.852 | $7,183.564 |
MU3 x 8 | $150.72 | $72,577.152 | ||
Qwen3.6-Plus-2026-04-02 | qwen3.6-plus-2026-04-02 | MU1 x 8 MU1 x 16(PD separation mode) | $59.408 PD separation mode:$118.816 | $28,734.256 PD separation mode:$57,468.512 |
Qwen3.5-397B-A17B | qwen3.5-397b-a17b | MU3 x 8 MU3 x 16(PD separation mode) | $150.72 PD separation mode:$301.44 | $72,577.152 PD separation mode:$145,154.304 |
MU6 x 16 | $55.008 | $26,599.92 | ||
Qwen3.5-122B-A10B | qwen3.5-122b-a10b | MU1 x 4 | $29.704 | $14,367.128 |
MU6 x 16 | $55.008 | $26,599.92 | ||
Qwen3.5-35B-A3B | qwen3.5-35b-a3b | MU1 x 2 | $14.852 | $7,183.564 |
MU2 x 8 | $69.312 | $33,044.72 | ||
MU3 x 8 | $150.72 | $72,577.152 | ||
MU9 x 1 | $7.014 | $3,383.024 | ||
Qwen3.5-27B | qwen3.5-27b | MU2 x 8 | $69.312 | $33,044.72 |
MU3 x 8 | $150.72 | $72,577.152 | ||
MU8 x 1 | $6.464 | $3,080.477 | ||
MU9 x 1 | $7.014 | $3,383.024 | ||
Qwen3.5-9B | qwen3.5-9b | MU2 x 2 | $17.328 | $8,261.18 |
Qwen3.5-Flash-2026-02-23 | qwen3.5-flash-2026-02-23 | MU1 x 2 | $14.852 | $7,183.564 |
Qwen3.5-Plus-2026-02-15 | qwen3.5-plus-2026-02-15 | MU1 x 8 MU1 x 16(PD separation mode) | $59.408 PD separation mode:$118.816 | $28,734.256 PD separation mode:$57,468.512 |
MU2 x 8 | $69.312 | $33,044.72 | ||
MU3 x 8 MU3 x 16(PD separation mode) | $150.72 PD separation mode:$301.44 | $72,577.152 PD separation mode:$145,154.304 | ||
Qwen3-235B-A22B-Instruct-2507 | qwen3-235b-a22b-instruct-2507 | MU1 x 4 | $29.704 | $14,367.128 |
MU2 x 8 | $69.312 | $33,044.72 | ||
MU3 x 8 | $150.72 | $72,577.152 | ||
Qwen3-32B | qwen3-32b | MU6 x 16 | $55.008 | $26,599.92 |
Qwen3-30B-A3B-Thinking-2507 | qwen3-30b-a3b-thinking-2507 | MU1 x 2 | $14.852 | $7,183.564 |
Qwen3-8B | qwen3-8b | MU1 x 2 | $14.852 | $7,183.564 |
MU2 x 2 | $17.328 | $8,261.18 | ||
Qwen3-4B | qwen3-4b | MU1 x 2 | $14.852 | $7,183.564 |
MU5 x 1 | $2.888 | $1,394.329 | ||
Qwen3-Embedding-0.6B | qwen3-embedding-0.6b | MU5 x 1 | $2.888 | $1,394.329 |
MU6 x 1 | $3.438 | $1,662.495 | ||
Qwen3-MoE-Rerank-0.6B | qwen3-moe-rerank-0.6b | MU5 x 1 | $2.888 | $1,394.329 |
Qwen3-Rerank-0.6B | qwen3-rerank-0.6b | MU5 x 1 | $2.888 | $1,394.329 |
MU6 x 1 | $3.438 | $1,662.495 | ||
Qwen3-Max-2025-09-23 | qwen3-max-2025-09-23 | MU2 x 8 | $69.312 | $33,044.72 |
MU3 x 8 | $150.72 | $72,577.152 | ||
Qwen3-Rerank | qwen3-rerank | MU5 x 1 | $2.888 | $1,394.329 |
Qwen2.5-Open-Source-72B | qwen2.5-72b-instruct | MU1 x 8 | $59.408 | $28,734.256 |
Qwen2.5-Open-Source-14B | qwen2.5-14b-instruct | MU1 x 2 | $14.852 | $7,183.564 |
Qwen2.5-Open-Source-7B | qwen2.5-7b-instruct | MU1 x 2 | $14.852 | $7,183.564 |
MU5 x 1 | $2.888 | $1,394.329 | ||
Qwen-Plus-2025-07-28 | qwen-plus-2025-07-28 | MU1 x 4 MU1 x 16(PD separation mode) | $29.704 PD separation mode:$118.816 | $14,367.128 PD separation mode:$57,468.512 |
Qwen-Plus-2025-12-01 | qwen-plus-2025-12-01 | MU1 x 4 | $29.704 | $14,367.128 |
Qwen-Plus-Character-2025-11-06 | qwen-plus-character-2025-11-06 | MU1 x 4 | $29.704 | $14,367.128 |
Tab
GLM
Model name | Model code | Model unit specification | Hourly unit price (USD) Minimum billing: minute | Monthly unit price (USD) Minimum billing: day |
|---|---|---|---|---|
GLM-5.1 | glm-5.1 | MU2 x 8 | $69.312 | $33,044.72 |
MU3 x 16(PD separation mode) | PD separation mode:$301.44 | PD separation mode:$145,154.304 | ||
MU6 x 16 | $55.008 | $26,599.92 | ||
GLM-5 | glm-5 | MU3 x 16(PD separation mode) | PD separation mode:$301.44 | PD separation mode:$145,154.304 |
GLM-4.7 | glm-4.7 | MU6 x 32(PD separation mode) | PD separation mode:$110.016 | PD separation mode:$53,199.84 |
Tab
DeepSeek
Model name | Model code | Model unit specification | Hourly unit price (USD) Minimum billing: minute | Monthly unit price (USD) Minimum billing: day |
|---|---|---|---|---|
DeepSeek-v4-Flash | deepseek-v4-flash | MU1 x 8 | $59.408 | $28,734.256 |
MU3 x 8 | $150.72 | $72,577.152 | ||
DeepSeek-v3.2 | deepseek-v3.2 | MU2 x 16 (PD separation mode) | PD separation mode: $138.624 | PD separation mode: $66,089.44 |
Tab
Other models
Model name | Model code | Model unit specification | Hourly unit price ($) Minimum billing: minute | Monthly unit price ($) Minimum billing: day |
|---|---|---|---|---|
Kimi-K2.5 | kimi-k2.5 | MU2 x 8 | $69.312 | $33,044.72 |
Multimodal
Tab
Qwen-VL
Model name | Model code | Model unit specification | Hourly unit price ($) Minimum billing: minute | Monthly unit price ($) Minimum billing: day |
|---|---|---|---|---|
Qwen3-VL-235B-A22B-Thinking | qwen3-vl-235b-a22b-thinking | MU1 x 8 | $59.408 | $28,734.256 |
MU2 x 8 | $69.312 | $33,044.72 | ||
MU3 x 8 | $150.72 | $72,577.152 | ||
Qwen3-VL-32B-Instruct | qwen3-vl-32b-instruct | MU2 x 8 | $69.312 | $33,044.72 |
MU3 x 8 | $150.72 | $72,577.152 | ||
Qwen3-VL-8B-Instruct | qwen3-vl-8b-instruct | MU1 x 2 | $14.852 | $7,183.564 |
MU5 x 1 | $2.888 | $1,394.329 | ||
Qwen3-VL-4B-Instruct | qwen3-vl-4b-instruct | MU1 x 2 | $14.852 | $7,183.564 |
Qwen3-VL-2B-Instruct | qwen3-vl-2b-instruct | MU5 x 1 | $2.888 | $1,394.329 |
Qwen3-VL-Embedding-2B | qwen3-vl-embedding-2b | MU5 x 1 | $2.888 | $1,394.329 |
Qwen3-VL-Flash-2025-10-15 | qwen3-vl-flash-2025-10-15 | MU1 x 4 | $29.704 | $14,367.128 |
Qwen3-VL-Plus-2025-09-23 | qwen3-vl-plus-2025-09-23 | MU1 x 4 | $29.704 | $14,367.128 |
Qwen-VL-Max-2025-08-13 | qwen-vl-max-2025-08-13 | MU6 x 4 | $13.752 | $6,649.98 |
Tab
Qwen-Omni
Model name | Model code | Model unit specification | Hourly unit price ($) Minimum billing: minute | Monthly unit price ($) Minimum billing: day |
|---|---|---|---|---|
Qwen3.5-Omni-Flash | qwen3.5-omni-flash | MU9 x 1 | $7.014 | $3,383.024 |
Model types:
- Instruct - The model performs inference in non-thinking mode after deployment.
- Thinking - The model performs inference in thinking mode after deployment.
By model Token usage
Fee = Model input token count × Model input unit price + Model output token count × Model output unit price (minimum billing unit: 1 token)
- Billing by model Token usage is supported only after you complete SFT efficient training on the following base models and obtain a custom model.
Singapore
Base model | Model code | Input $/million tokens | Output $/million tokens |
|---|---|---|---|
Qwen3-14B | qwen3-14b | Non-thinking mode:$0.35 Thinking mode:$0.35 | Non-thinking mode:$1.4 Thinking mode:$4.2 |
Beijing
Base model | Model code | Input $/million tokens | Output $/million tokens |
|---|---|---|---|
Qwen3.5-27B | qwen3.5-27b | <128K $0.086 128K-256K $0.258 | <128K $0.688 128K-256K $2.064 |
Qwen3-32B | qwen3-32b | Non-thinking mode:$0.287 Thinking mode:$0.287 | Non-thinking mode:$1.147 Thinking mode:$2.868 |
Qwen3-14B | qwen3-14b | Non-thinking mode:$0.144 Thinking mode:$0.144 | Non-thinking mode:$0.574 Thinking mode:$1.434 |
Qwen3-8B | qwen3-8b | Non-thinking mode:$0.072 Thinking mode:$0.072 | Non-thinking mode:$0.287 Thinking mode:$0.717 |
Qwen3-VL-8B-Instruct | qwen3-vl-8b-instruct | $0.072 | $0.287 |
Response example
The command returns the following:
{
"request_id": "f2ae64f7-83cc-410c-bc0b-840443f7eb86",
"output": {
"deployed_model": "emo-35b3f106-sample01",
"gmt_create": "2025-06-17T11:00:38.68",
"gmt_modified": "2025-06-17T11:00:38.68",
"status": "PENDING",
"model_name": "emo",
"base_model": "emo",
"base_capacity": 1,
"capacity": 1,
"ready_capacity": 0,
"workspace_id": "llm-v71tlv3d***",
"charge_type": "post_paid",
"creator": "175805416***",
"modifier": "175805416***"
}
}
Response parameters
Parameter | Type | Description |
|---|---|---|
request_id | String | The ID of the request. |
output | Object | Details of the deployment task. |
deployed_model | String | A unique identifier for the deployed model. This ID is used for API operations, such as querying deployment details, modifying deployment rate limiting, deployment scaling, and deleting deployments, and is also passed as an SDK parameter when you invoke the model. |
gmt_create | String | The creation time of the deployment task. |
gmt_modified | String | The last modification time of the deployment task. |
status | String | The status of the deployment task.
|
model_name | String | The name of the model used in the deployment task. |
base_model | String | The ID of the base model used in the deployment task. |
base_capacity | Number | The minimum number of resource units required to run the base model. |
capacity | Number | The number of resource units used by the deployment task. |
ready_capacity | Number | The number of resource units that are ready to process requests immediately. Resource initialization speed or hardware status can limit this value. |
workspace_id | String | The ID of the deployment task's workspace. |
charge_type | String | The billing method for the deployment task. |
creator | String | The UID of the user who created the deployment task. |
modifier | String | The UID of the user who last modified the deployment task. |
plan | String | The billing model for the deployment task. This parameter is not returned for some billing models. |
Returned only for Model Unit deployments. | ||
model_unit_spec | String | The model unit specification. |
enable_thinking | Boolean | Specifies if Thinking mode is enabled. This feature is only available for certain models. |
max_context_length | Number | The maximum context length. |
rpm_limit | String | The maximum number of requests per minute (RPM). |
tpm_limit | Number | The maximum number of tokens per minute (TPM). |
Returned only for provisioned throughput (PTU) deployments. | ||
ptu_capacity | Object | This parameter takes effect only when Example: |
ptu_capacity.input_tpm | Number | The maximum number of input tokens per minute (TPM) for the deployed model. This feature is supported by all models. |
ptu_capacity.output_tpm | Number | The maximum number of output tokens per minute (TPM) for the deployed model. This feature is supported by all models. |
ptu_capacity.thinking_output_tpm | Number | The maximum number of thinking output tokens per minute (TPM) for the deployed model. This feature is only available for certain models. |
Error response
Response example
{
"request_id": "ca218d57-b91b-46b2-bd35-c41c6287bcf4",
"message": "Model: qwen-plus-20230703-cx7f not found!",
"code": "NotFound"
}
Response parameters
Parameter | Type | Description |
|---|---|---|
request_id | String | The unique ID of the request. |
code | String | The error code. |
message | String | The error message. |
The following errors can occur when a request fails:
Error code | Error message | Reason |
|---|---|---|
NotFound | Model: xxx not found! |
|
Conflict | Deployed model xxx already exists, please specify a suffix. | You are creating a deployment task with a suffix that is already in use. |
InvalidParameter | Invalid capacity (xx), capacity must be larger than or equal to 0 and multiples of 1 and less than 1000! | You are creating or updating a deployment task with an invalid number of capacity units. |