All Products
Search
Document Center

Alibaba Cloud Model Studio:Create deployment

Last Updated:Jul 06, 2026

Create a model deployment task.

Prerequisites

Model deployment

Endpoint

POST https://dashscope-intl.aliyuncs.com/api/v1/deployments

Request examples

Billing by provisioned throughput units (PTU)

Note

After running the deployment command below, billing starts immediately upon successful deployment—even if you have not yet called the model. Confirm the billing rules before deploying.

The PTU billing mode charges based on the duration of provisioned throughput usage. It suits scenarios requiring stable throughput guarantees, high concurrency, low latency, and predictable traffic. In this mode, throughput/concurrency and generation speed are preset by the platform and cannot be adjusted.

curl "https://dashscope-intl.aliyuncs.com/api/v1/deployments" \
--header "Authorization: Bearer $DASHSCOPE_API_KEY" \
--header 'Content-Type: application/json' \
--data '{
    "name": "my_qwen_flash",
    "model_name": "qwen-flash-2025-07-28",
    "plan": "ptu",
    "ptu_capacity": {
        "input_tpm": 10000,
	"output_tpm": 1000
    }
}'

Billing by model unit usage duration

Note
  • After running the deployment command below, billing starts immediately upon successful deployment—even if you have not yet called the model. Confirm the billing rules before deploying.

  • Model unit pay-as-you-go computing resources are allocated on a first-come, first-served basis. If purchase fails, you receive a full refund.

Select the model unit billing method. This mode charges based on model unit usage duration and suits large-scale inference workloads after model fine-tuning. Resources are dedicated, and performance and cost are flexible. Throughput/concurrency and generation speed are customizable.

curl "https://dashscope-intl.aliyuncs.com/api/v1/deployments" \
--header "Authorization: Bearer $DASHSCOPE_API_KEY" \
--header 'Content-Type: application/json' \
--data '{
    "name": "my_qwen_plus",
    "model_name": "qwen-plus-2025-12-01",
    "plan": "mu",
    "deploy_spec": "MU1",
    "enable_thinking": true,
    "capacity": 4,
    "max_context_length": 10000,
    "rpm_limit": 500,
    "tpm_limit": 1000
}'

The model unit deployment mode supports additional settings:

Configuration

Details

Service Name

The custom name of the deployment service.

Select Model

Select the model to deploy, including pre-trained platform models and fine-tuned models.

Model Unit Type

Select the deployment specification. Different specifications correspond to different computing power and performance.

Number of Replicas

Set the initial number of deployment replicas, which affects the service's concurrent processing capability.

Deployment Template

Select a deployment template, such as "Single-Machine Deployment". Different templates correspond to different resource configuration plans. This is only available in the model unit billing mode.

Configure Model Inference Mode

When some models are deployed using the Model Unit method, you can configure the inference mode, maximum context length, and other settings.

  • Instruct - The deployed model performs inference in non-thinking mode.

  • Thinking - The deployed model performs inference in thinking mode.

Maximum Context

This setting is supported for the Model Unit deployment mode of some models. The maximum context length is based on the model type.

Service Throttling

This setting is supported for the Model Unit deployment mode of some models. You can limit the RPM and TPM of model calls.

For details on setting these options via API, see Create a model deployment task using the API.

Billing by token usage

Select the token-based billing method. This mode charges based on token usage and suits cost-sensitive scenarios with low requirements for concurrency and latency. It offers the highest price advantage. Throughput/concurrency and generation speed are preset by the platform and cannot be adjusted.

curl "https://dashscope-intl.aliyuncs.com/api/v1/deployments" \
--header "Authorization: Bearer $DASHSCOPE_API_KEY" \
--header 'Content-Type: application/json' \
--data '{
    "model_name": "qwen3-8b-ft-202511132025-0260",
    "plan": "lora",
    "capacity": 1,
    "name": "qwen3-8b-ft"
}'
The capacity parameter has no effect but must be included. To scale up or down, go to the Model Studio model deployment console and submit a form request.

Request parameters

Parameter

Type

Location

Required

Description

model_name

String

body

Yes

The name of the model to deploy. This corresponds to the model ID in My Models. You can also get this ID from the output of the Create Training Job or Create Import Job operations.

plan

String

body

Yes

The deployment plan. The following billing methods are supported:

Billing method

Plan setting

Billing by model unit

"plan": "mu"

Billing by compute unit

"plan": "cu"

Provisioned throughput

"plan": "ptu"

LoRA shared deployment (billed by token usage)

"plan": "lora"

You can quickly find the supported deployment plans for a fine-tuned model in My Models.

Note

Fine-tuned CosyVoice models currently only support "plan": "mu".

name

String

body

Yes

The display name of the model in the console.

capacity

Integer

body

No

Required only when "plan": "mu" is specified. Specifies the number of resource units for the deployment. The value must be an integer multiple of base_capacity. The constraints vary based on the deploy_spec value. For example, for MU2, the value must be a multiple of 8, while for MU5, it can be 1. Example: "capacity": 1.

Note

CosyVoice models currently provide the following two deployment templates with corresponding capacity constraints:

  • single-node deployment: capacity must be an integer multiple of 1, such as 1, 2, 3, 4, or 5.

  • single-node deployment - flagship complex inference edition: capacity must be an integer multiple of 8, such as 8, 16, 24, or 32.

billing_method

String

body

No

Required only when "plan": "mu" is specified. Currently, only "POST_PAY" (Post-paid) is supported. Example: "billing_method": "POST_PAY".

deploy_spec

String

body

No

This setting is applicable only when "plan": "mu" is specified.

For details about feature support, see Feature support for model unit deployment.

This parameter is required when "plan": "mu" is specified. Example: "deploy_spec": "MU1".

Note

You can get this value from the template_id field returned by the Get Deployable Model List operation.

enable_thinking

Boolean

body

No

Supported by some models. You can set this to true or false.

max_context_length

Number

body

No

Supported by some models. Example: "max_context_length": 131072.

rpm_limit

Number

body

No

Supported by some models. Specifies the maximum number of requests per minute (RPM).

tpm_limit

Number

body

No

Supported by some models. Specifies the maximum number of tokens per minute (TPM).

ptu_capacity

Object

body

No

This setting is applicable only when "plan": "ptu" is specified.

For details about feature support, see Feature Support for PTU Deployment.

If you do not specify this parameter, the system defaults to 10,000 input_tpm and 1,000 output_tpm.

Example: "ptu_capacity": { "input_tpm": 10000, "output_tpm": 1000 }.

Example: "ptu_capacity": { "input_tpm": 10000, "output_tpm": 1000 }.

ptu_capacity.input_tpm

Number

body

No

Supported by all models. Specifies the maximum number of input tokens per minute (TPM).

ptu_capacity.output_tpm

Number

body

No

Supported by all models. Specifies the maximum number of output tokens per minute (TPM).

ptu_capacity.thinking_output_tpm

Number

body

No

Supported by some models. Specifies the maximum number of provisioned thinking output tokens per minute (TPM).

suffix

String

body

No

After a model is deployed, a new model name is generated. The suffix parameter specifies the suffix for this new name. It must be globally unique and have a maximum length of 8 characters. You can omit the suffix for the first deployment of a model. If you deploy the same model multiple times, you must specify a unique suffix for each deployment.

See the deployed_model output parameter for more information.

Supported models

View supported features and billing.

Time-based billing (Provisioned Throughput)

Cost = Usage Duration × (Input TPM Unit Price × Input TPM + Output TPM Unit Price × Output TPM)

For the pay-as-you-go method, usage is billed hourly, and the unit price is based on the hourly rates in the table below. For the subscription method, usage is billed daily, and the unit price is based on the daily rates in the table below.

  • Subscription orders take effect immediately after payment. An N-day subscription is valid until 23:59 on the Nth day. If an order is placed after 22:00, the expiration date is automatically extended by one day.

  • After a subscription order expires, the service is stopped after a 2-hour grace period. After the service is stopped, the resources are retained for 14 hours and then released.

  • Subscription orders cannot be terminated early.

  • For the pay-as-you-go method, if your account has an overdue payment, the deployed resources are retained and continue to be billed for 24 hours, during which the service remains available. After 24 hours, the system stops billing, and the model deployment enters an overdue state. The underlying resources are deleted, but the model deployment task is retained. After you pay the overdue amount, the system reallocates resources, restores the service, and resumes billing. To stop incurring charges, you must delete the model deployment task. Billing stops after the task is successfully deleted.

If the model input exceeds the maximum input tokens or the purchased TPM, the call automatically switches to the pay-as-you-go mode for the current model. In this case, inference performance may decrease and will be subject to the public traffic control of the current snapshot model in the workspace. Costs are charged based on the model invocation (pay-as-you-go) standard.

  • In this case, the API call returns a header that contains x-dashscope-ptu-overflow:true.

  • To view TPM statistics, go to Model Monitoring.

For the specific refund rules for scale-in scenarios (downgrades), see Refund rules for downgrades.

Singapore
Qwen

Model name

Model code

Max input tokens

Pay-as-you-go input

Per 10k TPM/hour

Pay-as-you-go output

Per 1k TPM/hour

Subscription input

Per 10k TPM/day

Subscription output

Per 1k TPM/day

Qwen3.7-Max-2026-05-20

qwen3.7-max-2026-05-20

256K

$6

$1.8

$72

$21.6

Qwen3.7-Plus-2026-05-26

qwen3.7-plus-2026-05-26

256K

$0.96

$0.384

$11.52

$4.608

Qwen3.6-Plus-2026-04-02

qwen3.6-plus-2026-04-02

128K

$1.2

$0.72

$14.4

$8.64

Qwen3.5-Plus-2026-04-20

qwen3.5-plus-2026-04-20

128K

$0.96

$0.576

$11.52

$6.912

DeepSeek

Model name

Model code

Max input tokens

Pay-as-you-go input

Per 10k TPM/hour

Pay-as-you-go output

Per 1k TPM/hour

Subscription input

Per 10k TPM/day

Subscription output

Per 1k TPM/day

DeepSeek-v4-Flash

deepseek-v4-flash

256K

$0.72

$0.144

$8.64

$1.728

DeepSeek-v4-Pro

deepseek-v4-pro

256K

$8.64

$1.728

$103.68

$20.736

DeepSeek-v3.2

deepseek-v3.2

64K

$2.05

$0.616

$24.62

$7.387

Qwen-VL

Model name

Model code

Max input tokens

Pay-as-you-go input

Per 10k TPM/hour

Pay-as-you-go output

Per 1k TPM/hour

Subscription input

Per 10k TPM/day

Subscription output

Per 1k TPM/day

Qwen3-VL-Plus-2025-09-23

qwen3-vl-plus-2025-09-23

128K

$0.48

$0.384

$5.76

$4.608

More models

Model name

Model code

Max input tokens

Pay-as-you-go input

Per 10k TPM/hour

Pay-as-you-go output

Per 1k TPM/hour

Subscription input

Per 10k TPM/day

Subscription output

Per 1k TPM/day

GLM-5.1

glm-5.1

64K

$5.04

$1.584

$64.8

$19.008

China (Beijing)
Qwen

Model name

Model code

Max input tokens

Pay-as-you-go input

Per 10k TPM/hour

Pay-as-you-go output

Per 1k TPM/hour

Subscription input

Per 10k TPM/day

Subscription output

Per 1k TPM/day

Qwen3.7-Max-2026-05-20

qwen3.7-max-2026-05-20

256K

$3.96

$1.188

$47.53

$14.258

Qwen3.7-Plus-2026-05-26

qwen3.7-plus-2026-05-26

256K

$0.66

$0.264

$7.92

$3.168

Qwen3.6-Plus-2026-04-02

qwen3.6-plus-2026-04-02

128K

$0.67

$0.397

$7.93

$4.753

Qwen3.5-Plus-2026-04-20

qwen3.5-plus-2026-04-20

128K

$0.26

$0.16

$3.17

$1.9

Qwen3-Max-2025-09-23

qwen3-max-2025-09-23

128K

$1.11

$0.45

$13.32

$5.4

Qwen-Flash-2025-07-28

qwen-flash-2025-07-28

128K

$0.06

$0.06

$0.72

$0.72

Qwen-Plus-2025-12-01

qwen-plus-2025-12-01

128K

$0.28

Non-thinking mode: $0.07

Thinking mode: $0.28

$3.36

Non-thinking mode: $0.84

Thinking mode: $3.36

DeepSeek

Model name

Model code

Max input tokens

Pay-as-you-go input

Per 10k TPM/hour

Pay-as-you-go output

Per 1k TPM/hour

Subscription input

Per 10k TPM/day

Subscription output

Per 1k TPM/day

DeepSeek-v4-Flash

deepseek-v4-flash

256K

$0.5

$0.099

$5.94

$1.188

DeepSeek-v4-Pro

deepseek-v4-pro

256K

$5.94

$1.188

$71.3

$14.26

DeepSeek-v3.2

deepseek-v3.2

64K

$1.04

$0.16

$12.48

$1.92

DeepSeek-v3

deepseek-v3

64K

$0.99

$0.396

$11.9

$4.75

Qwen-VL

Model name

Model code

Max input tokens

Pay-as-you-go input

Per 10k TPM/hour

Pay-as-you-go output

Per 1k TPM/hour

Subscription input

Per 10k TPM/day

Subscription output

Per 1k TPM/day

Qwen3-VL-Plus-2025-09-23

qwen3-vl-plus-2025-09-23

128K

$0.35

$0.35

$4.2

$4.2

More models

Model name

Model code

Max input tokens

Pay-as-you-go input

Per 10k TPM/hour

Pay-as-you-go output

Per 1k TPM/hour

Subscription input

Per 10k TPM/day

Subscription output

Per 1k TPM/day

GLM-5.1

glm-5.1

64K

$2.97

$1.19

$35.65

$14.26

Time-based billing (Model Unit)

Cost = Usage Duration (hours) × Number of Model Units × Model Unit Price

For the pay-as-you-go method, the "Model Unit Price" is the "Hourly Price" from the table below. For the monthly subscription method, the formula is: Number of Months × Number of Model Units × Monthly Price.

  • For subscriptions, if you unsubscribe within the first month, the daily unit price (≈ monthly unit price / 30) is charged at 1.2 times the standard rate. Usage for less than a day is billed as a full day.

Note

For the Model Unit pay-as-you-go method, computing power resources are allocated on a first-come, first-served basis. A full refund is issued if the purchase is unsuccessful.

Singapore
Text generation

Model name

Model code

Model unit specification

Hourly price ($)

Minimum billing unit: minute

Monthly price ($)

Minimum billing unit: day

Qwen3.6-Plus-2026-04-02

qwen3.6-plus-2026-04-02

MU1 x 8

$88

$41,832

Qwen3.5-39B-A17B

qwen3.5-397b-a17b

MU2 x 8

$112

$52,392

Qwen3.5-35B-A3B

qwen3.5-35b-a3b

MU2 x 8

$112

$52,392

Qwen3-32B

qwen3-32b

MU1 x 4

$44

$20,916

MU2 x 8

$112

$52,392

Qwen3-14B

qwen3-14b

MU1 x 4

$44

$20,916

GLM-5.1

glm-5.1

MU2 x 8

$112

$52,392

DeepSeek-V4-Flash

deepseek-v4-flash

MU1 x 8

$88

$41,832

Multimodal

Model name

Model code

Model unit specification

Hourly price ($)

Minimum billing unit: minute

Monthly price ($)

Minimum billing unit: day

Qwen3-VL-32B-Instruct

qwen3-vl-32b-instruct

MU2 x 8

$112

$52,392

Qwen3-VL-8B-Instruct

qwen3-vl-8b-instruct

MU1 x 2

$22

$10,458

Model types:

  • Instruct - The deployed model performs inference in non-thinking mode.

China (Beijing)
Text generation
Qwen

Model name

Model code

Model unit specification

Hourly price ($)

Minimum billing unit: minute

Monthly price ($)

Minimum billing unit: day

Qwen3.7-Plus-2026-05-26

qwen3.7-plus-2026-05-26

MU3 x 8

$150.72

$72,577.152

Qwen3.6-35B-A3B

qwen3.6-35b-a3b

MU8 x 1

$6.464

$3,080.477

MU9 x 1

$7.014

$3,383.024

Qwen3.6-27B

qwen3.6-27b

MU9 x 1

$7.014

$3,383.024

Qwen3.6-Flash-2026-04-16

qwen3.6-flash-2026-04-16

MU1 x 2

$14.852

$7,183.564

Qwen3.6-Plus-2026-04-02

qwen3.6-plus-2026-04-02

MU1 x 8

$59.408

$28,734.256

Qwen3.5-397B-A17B

qwen3.5-397b-a17b

MU2 x 8

$69.312

$33,044.72

MU3 x 8

$150.72

$72,577.152

MU6 x 16

$55.008

$26,599.92

Qwen3.5-122B-A10B

qwen3.5-122b-a10b

MU1 x 4

$29.704

$14,367.128

MU2 x 8

$69.312

$33,044.72

MU6 x 16

$55.008

$26,599.92

MU9 x 2

$14.028

$6,766.048

Qwen3.5-35B-A3B

qwen3.5-35b-a3b

MU1 x 2

$14.852

$7,183.564

MU2 x 8

$69.312

$33,044.72

MU8 x 1

$6.464

$3,080.477

MU9 x 1

$7.014

$3,383.024

Qwen3.5-27B

qwen3.5-27b

MU9 x 1

$7.014

$3,383.024

Qwen3.5-9B

qwen3.5-9b

MU8 x 1

$6.464

$3,080.477

MU9 x 1

$7.014

$3,383.024

Qwen3.5-Flash-2026-02-23

qwen3.5-flash-2026-02-23

MU1 x 2

$14.852

$7,183.564

Qwen3.5-Plus-2026-02-15

qwen3.5-plus-2026-02-15

MU1 x 8

$59.408

$28,734.256

MU3 x 8

$150.72

$72,577.152

Qwen3-235B-A22B-Instruct

qwen3-235b-a22b-instruct-2507

MU1 x 4

$29.704

$14,367.128

MU2 x 8

$69.312

$33,044.72

Qwen3-Next-80B-A3B-Instruct

qwen3-next-80b-a3b-instruct

MU1 x 2

$14.852

$7,183.564

Qwen3-32B

qwen3-32b

MU1 x 4

$29.704

$14,367.128

MU6 x 4

$13.752

$6,649.98

Qwen3-30B-A3B

qwen3-30b-a3b

MU9 x 2

$14.028

$6,766.048

Qwen3-30B-A3B-Instruct-2507

qwen3-30b-a3b-instruct-2507

MU1 x 4

$29.704

$14,367.128

MU2 x 8

$69.312

$33,044.72

Qwen3-8B

qwen3-8b

MU1 x 2

$14.852

$7,183.564

MU2 x 2

$17.328

$8,261.18

MU5 x 1

$2.888

$1,394.329

Qwen3-4B

qwen3-4b

MU1 x 2

$14.852

$7,183.564

MU5 x 1

$2.888

$1,394.329

Qwen3-1.7B

qwen3-1.7b

MU1 x 2

$14.852

$7,183.564

MU5 x 1

$2.888

$1,394.329

Qwen3-Max-2025-09-23

qwen3-max-2025-09-23

MU2 x 8

$69.312

$33,044.72

MU3 x 8

$150.72

$72,577.152

Qwen2.5-72B

qwen2.5-72b-instruct

MU1 x 4

$29.704

$14,367.128

Qwen2.5-32B

qwen2.5-32b-instruct

MU1 x 4

$29.704

$14,367.128

Qwen2.5-14B

qwen2.5-14b-instruct

MU1 x 2

$14.852

$7,183.564

Qwen2.5-7B

qwen2.5-7b-instruct

MU1 x 2

$14.852

$7,183.564

MU5 x 1

$2.888

$1,394.329

Qwen2.5-3B-Instruct

qwen2.5-3b-instruct

MU5 x 1

$2.888

$1,394.329

Qwen-Flash-2025-07-28

qwen-flash-2025-07-28

MU1 x 4

$29.704

$14,367.128

Qwen-Plus-2025-07-28

qwen-plus-2025-07-28

MU1 x 4

$29.704

$14,367.128

Qwen-Plus-2025-12-01

qwen-plus-2025-12-01

MU1 x 4

$29.704

$14,367.128

GLM

Model name

Model code

Model unit specification

Hourly price ($)

Minimum billing unit: minute

Monthly price ($)

Minimum billing unit: day

GLM-5.1

glm-5.1

MU2 x 8

$69.312

$33,044.72

MU3 x 8

$150.72

$72,577.152

MU6 x 16

$55.008

$26,599.92

GLM-5

glm-5

MU3 x 8

$150.72

$72,577.152

GLM-4.7

glm-4.7

MU6 x 16

$55.008

$26,599.92

DeepSeek

Model name

Model code

Model unit specification

Hourly price ($)

Minimum billing unit: minute

Monthly price ($)

Minimum billing unit: day

DeepSeek-V4-Flash

deepseek-v4-flash

MU1 x 8

$59.408

$28,734.256

DeepSeek-V3.2

deepseek-v3.2

MU2 x 8

$69.312

$33,044.72

Other models

Model name

Model code

Model unit specification

Hourly price ($)

Minimum billing unit: minute

Monthly price ($)

Minimum billing unit: day

MiniMax-M2.5

MiniMax-M2.5

MU1 x 8

$59.408

$28,734.256

Kimi-K2.5

kimi-k2.5

MU2 x 8

$69.312

$33,044.72

Multimodal
Qwen-VL

Model name

Model code

Model unit specification

Hourly price ($)

Minimum billing unit: minute

Monthly price ($)

Minimum billing unit: day

Qwen3-VL-235B-A22B-Instruct

qwen3-vl-235b-a22b-instruct

MU1 x 4

$29.704

$14,367.128

Qwen3-VL-235B-A22B-Thinking

qwen3-vl-235b-a22b-thinking

MU1 x 4

$29.704

$14,367.128

Qwen3-VL-32B-Instruct

qwen3-vl-32b-instruct

MU2 x 8

$69.312

$33,044.72

Qwen3-VL-8B-Instruct

qwen3-vl-8b-instruct

MU1 x 2

$14.852

$7,183.564

Qwen3-VL-Flash-2025-10-15

qwen3-vl-flash-2025-10-15

MU1 x 4

$29.704

$14,367.128

Qwen3-VL-Plus-2025-09-23

qwen3-vl-plus-2025-09-23

MU1 x 4

$29.704

$14,367.128

Qwen-VL-Max-2025-08-13

qwen-vl-max-2025-08-13

MU6 x 4

$13.752

$6,649.98

Qwen-VL-OCR-2025-11-20

qwen-vl-ocr-2025-11-20

MU6 x 4

$13.752

$6,649.98

Qwen Omni

Model name

Model code

Model unit specification

Hourly price ($)

Minimum billing unit: minute

Monthly price ($)

Minimum billing unit: day

Qwen3.5-Omni-Flash

qwen3.5-omni-flash

MU8 x 1

$6.464

$3,080.477

MU9 x 1

$7.014

$3,383.024

Qwen3.5-Omni-Plus

qwen3.5-omni-plus

MU9 x 8

$56.112

$27,064.192

Model types:

  • Instruct - The deployed model performs inference in non-thinking mode.

  • Thinking - The deployed model performs inference in thinking mode.

By model token usage

Cost = Number of Input Tokens × Input Unit Price + Number of Output Tokens × Output Unit Price (Minimum billing unit: 1 token)

  • Billing by model token usage is only supported after you have completed Supervised Fine-Tuning (SFT) for the following foundation models and you have obtained a custom model.

Singapore

Foundation model

Model code

Input

$/1k tokens

Output

$/1k tokens

Qwen3-14B

qwen3-14b

$0.00035

Non-thinking mode: $0.0014

Thinking mode: $0.0042

Response example

The command returns the following:

{
  "request_id": "f2ae64f7-83cc-410c-bc0b-840443f7eb86",
  "output": {
    "deployed_model": "emo-35b3f106-sample01",
    "gmt_create": "2025-06-17T11:00:38.68",
    "gmt_modified": "2025-06-17T11:00:38.68",
    "status": "PENDING",
    "model_name": "emo",
    "base_model": "emo",
    "base_capacity": 1,
    "capacity": 1,
    "ready_capacity": 0,
    "workspace_id": "llm-v71tlv3d***",
    "charge_type": "post_paid",
    "creator": "175805416***",
    "modifier": "175805416***"
  }
}

Response parameters

Parameter

Type

Description

request_id

String

The ID of the request.

output

Object

Details of the deployment task.

deployed_model

String

A unique identifier for the deployed model. This ID is used for API operations, such as querying deployment details, modifying deployment rate limiting, deployment scaling, and deleting deployments, and is also passed as an SDK parameter when you invoke the model.

gmt_create

String

The creation time of the deployment task.

gmt_modified

String

The last modification time of the deployment task.

status

String

The status of the deployment task.

  • PENDING: The task is being created.

  • UPDATING: The task is being updated.

  • RUNNING: The deployment task is running, and the deployed model can process requests.

  • STOPPED: The deployment task is stopped and is not billed.

  • DELETING: The task is being deleted.

  • FAILED: The creation or update of the task failed.

model_name

String

The name of the model used in the deployment task.

base_model

String

The ID of the base model used in the deployment task.

base_capacity

Number

The minimum number of resource units required to run the base model.

capacity

Number

The number of resource units used by the deployment task.

ready_capacity

Number

The number of resource units that are ready to process requests immediately. Resource initialization speed or hardware status can limit this value.

workspace_id

String

The ID of the deployment task's workspace.

charge_type

String

The billing method for the deployment task.

post_paid: Post-paid.

creator

String

The UID of the user who created the deployment task.

modifier

String

The UID of the user who last modified the deployment task.

plan

String

The billing model for the deployment task. This parameter is not returned for some billing models.

Returned only for Model Unit deployments.

model_unit_spec

String

The model unit specification.

enable_thinking

Boolean

Specifies if Thinking mode is enabled. This feature is only available for certain models.

max_context_length

Number

The maximum context length.

rpm_limit

String

The maximum number of requests per minute (RPM).

tpm_limit

Number

The maximum number of tokens per minute (TPM).

Returned only for provisioned throughput (PTU) deployments.

ptu_capacity

Object

This parameter takes effect only when "plan": "ptu" is set.

Example: "ptu_capacity": { "input_tpm": 10000, "output_tpm": 1000 }.

ptu_capacity.input_tpm

Number

The maximum number of input tokens per minute (TPM) for the deployed model. This feature is supported by all models.

ptu_capacity.output_tpm

Number

The maximum number of output tokens per minute (TPM) for the deployed model. This feature is supported by all models.

ptu_capacity.thinking_output_tpm

Number

The maximum number of thinking output tokens per minute (TPM) for the deployed model. This feature is only available for certain models.

Error response

Response example

{
    "request_id": "ca218d57-b91b-46b2-bd35-c41c6287bcf4",
    "message": "Model: qwen-plus-20230703-cx7f not found!",
    "code": "NotFound"
}

Response parameters

Parameter

Type

Description

request_id

String

The unique ID of the request.

code

String

The error code.

message

String

The error message.

The following errors can occur when a request fails:

Error code

Error message

Reason

NotFound

Model: xxx not found!

  • You are creating a deployment task with a model that does not exist.

  • You are querying, updating, or deleting a deployment task with a model that does not exist.

Conflict

Deployed model xxx already exists, please specify a suffix.

You are creating a deployment task with a suffix that is already in use.

InvalidParameter

Invalid capacity (xx), capacity must be larger than or equal to 0 and multiples of 1 and less than 1000!

You are creating or updating a deployment task with an invalid number of capacity units.