All Products
Search
Document Center

Alibaba Cloud Model Studio:PTU provisioned throughput deployment

Last Updated:Sep 10, 2026

This topic describes PTU (Provisioned Throughput) deployment billing rules, supported models and pricing, long input and prefix caching, Provisioned Throughput Quota Calculator usage, and API response fields.

Overview

PTU (Provisioned Throughput Unit) reserves platform resources to guarantee a specific TPM throughput capacity for model deployment; no rate limiting within the guaranteed quota. It is suitable for high-throughput, high-performance, high-load production environments, providing stable throughput capacity, lower latency, and stronger resource certainty. Compared with Token-based billing, TPS (tokens generated per second) typically increases by approximately 1.5 to 2.0 times.

Use cases:

  • Intelligent customer service for banking apps (stable traffic, requires guaranteed concurrent experience).
  • Real-time content moderation for social platforms (requires stable processing of predictable pipeline tasks).
  • Public cloud translation API (provides baseline service guarantees for standard package users).

It is also commonly used in long-document analysis (contracts, research report summaries) and multi-turn conversations (customer service, coding assistants) where input exceeds 32K tokens. When creating a PTU, you can choose an overflow strategy: auto-overflow (default, excess automatically converts to pay-as-you-go) or use-only-PTU-capacity (excess returns 429). See Billing rules for details.

PTU also supports long-input tiered capacity coefficients and cache discounts. See Long input and prefix caching for details. For how prefix caching works, see Context Cache.

Billing rules

Fee = Usage duration × (Input TPM unit price × Input TPM + Output TPM unit price × Output TPM)

Post-paid is calculated hourly: the usage duration unit is hours, and the unit price is taken from the "Continuous 1 hour" column in the table below; prepaid is calculated daily: the usage duration unit is days, and the unit price is taken from the "Continuous 1 day" column in the table below.

  • Prepaid orders take effect in real time after payment, with a validity period of N days ending at 23:59 on day N. If the order is placed after 22:00, the expiration date will be automatically extended by 1 day.
  • After a prepaid order expires, the service will be stopped with a 2-hour delay, and resources will be retained for 14 hours after the stop and then released.
  • Prepaid orders support early termination; the used portion is settled at a 1.2x coefficient for refund. For details, see Refund rules for configuration downgrades.
  • For post-paid billing, if the account is in arrears, the deployed resources will continue to be retained and billed for 24 hours, during which the service can still be used normally. After 24 hours, the system stops billing, the model deployment enters an arrears state, and the underlying resources will be deleted, but the model deployment task will be retained. After the arrears are paid, the system will reallocate resources and restore usage (fees will continue to accrue after restoration). If you do not want to continue incurring fees, you can delete the model deployment task, and billing will stop after successful deletion.

When the model input exceeds the maximum input Token, the relevant call will automatically switch to the pay-as-you-go mode of the current model; when the purchased TPM is exceeded, it is handled according to the overflow strategy selected at creation ("auto-overflow" switches to pay-as-you-go, "use-only-PTU-capacity" returns 429). At this time, inference performance may degrade and will be subject to the public traffic control of the current snapshot model in the business space, and fees will be charged according to the model invocation (pay-as-you-go) standard.

  • In this case (only under the "auto-overflow" strategy), the API response Header will include: x-dashscope-ptu-overflow:true.
  • For TPM statistics, go to: Model Monitoring.

For the specific fee reduction and refund rules in scale-down (downgrade) scenarios, please refer to: Refund rules for configuration downgrades.

NotePTU deployment supports long-input tiered capacity coefficients and cache discounts; see Long input and prefix caching for details.

Supported models and pricing

Singapore

Qwen

Model Name

Model Code

Max Input Tokens

Postpaid Input

Per 10K TPM/Hour

Postpaid Output

Per 1K TPM/Hour

Prepaid Input

Per 10K TPM/Day

Prepaid Output

Per 1K TPM/Day

Qwen3.8-Max

qwen3.8-max

1M

USD4.8

USD1.44

USD57.6

USD17.28

Qwen3.7-Flash-2026-07-15 Contact your account manager to enable

qwen3.7-flash-2026-07-15

128K

USD0.072

USD0.031

USD0.864

USD0.374

Qwen3.7-Max-2026-05-20

qwen3.7-max-2026-05-20

256K

USD1.92

USD1.8

USD72

USD21.6

Qwen3.7-Plus-2026-05-26

qwen3.7-plus-2026-05-26

256K

USD0.96

USD0.384

USD11.52

USD4.608

Qwen3.6-Plus-2026-04-02

qwen3.6-plus-2026-04-02

128K

USD1.2

USD0.72

USD14.4

USD8.64

Qwen3.5-Plus-2026-04-20

qwen3.5-plus-2026-04-20

128K

USD0.96

USD0.576

USD11.52

USD6.912

DeepSeek

Model Name

Model Code

Max Input Tokens

Postpaid Input

Per 10K TPM/Hour

Postpaid Output

Per 1K TPM/Hour

Prepaid Input

Per 10K TPM/Day

Prepaid Output

Per 1K TPM/Day

DeepSeek-v4-Flash

deepseek-v4-flash

256K

USD0.72

USD0.144

USD8.64

USD1.728

DeepSeek-v4-Flash-0731

deepseek-v4-flash-0731

64K

USD1.44

USD0.288

USD17.28

USD3.456

DeepSeek-v4-Pro

deepseek-v4-pro

256K

USD0.96

USD1.728

USD103.68

USD20.736

Qwen VL

Model Name

Model Code

Max Input Tokens

Postpaid Input

Per 10K TPM/Hour

Postpaid Output

Per 1K TPM/Hour

Prepaid Input

Per 10K TPM/Day

Prepaid Output

Per 1K TPM/Day

Qwen3-VL-Plus-2025-09-23

qwen3-vl-plus-2025-09-23

128K

USD0.48

USD0.384

USD5.76

USD4.608

GLM

Model Name

Model Code

Max Input Tokens

Postpaid Input

Per 10K TPM/Hour

Postpaid Output

Per 1K TPM/Hour

Prepaid Input

Per 10K TPM/Day

Prepaid Output

Per 1K TPM/Day

GLM-5.2

glm-5.2

1M

USD5.04

USD1.584

USD60.48

USD19.008

China North 2 (Beijing)

Qwen

Model Name

Model Code

Max Input Tokens

Postpaid Input

Per 10K TPM/Hour

Postpaid Output

Per 1K TPM/Hour

Prepaid Input

Per 10K TPM/Day

Prepaid Output

Per 1K TPM/Day

Qwen3.8-Max

qwen3.8-max

1M

USD3.96

USD1.188

USD47.53

USD14.258

Qwen3.7-Flash-2026-07-15 Contact your account manager to enable

qwen3.7-flash-2026-07-15

128K

USD0.066

USD0.026

USD0.792

USD0.317

Qwen3.7-Max-2026-05-20

qwen3.7-max-2026-05-20

256K

USD3.96

USD1.188

USD47.53

USD14.258

Qwen3.7-Plus-2026-05-26

qwen3.7-plus-2026-05-26

256K

USD0.66

USD0.264

USD7.92

USD3.168

Qwen3.6-Plus-2026-04-02

qwen3.6-plus-2026-04-02

128K

USD0.67

USD0.066

USD7.93

USD4.753

Qwen3.5-Plus-2026-04-20

qwen3.5-plus-2026-04-20

128K

USD0.26

USD0.16

USD3.17

USD1.9

Qwen3-Max-2025-09-23

qwen3-max-2025-09-23

128K

USD1.11

USD0.45

USD13.32

USD5.4

Qwen-Flash-2025-07-28

qwen-flash-2025-07-28

128K

USD0.06

USD0.06

USD0.72

USD0.72

Qwen-Plus-2025-12-01

qwen-plus-2025-12-01

128K

USD0.28

Non-thinking: USD0.07

Thinking: USD0.28

USD3.36

Non-thinking: USD0.84

Thinking: USD3.36

DeepSeek

Model Name

Model Code

Max Input Tokens

Postpaid Input

Per 10K TPM/Hour

Postpaid Output

Per 1K TPM/Hour

Prepaid Input

Per 10K TPM/Day

Prepaid Output

Per 1K TPM/Day

DeepSeek-v4-Flash

deepseek-v4-flash

256K

USD0.5

USD0.099

USD5.94

USD1.188

DeepSeek-v4-Flash-0731

deepseek-v4-flash-0731

64K

USD0.99

USD0.198

USD11.88

USD2.376

DeepSeek-v4-Pro

deepseek-v4-pro

256K

USD5.94

USD1.188

USD71.3

USD14.26

DeepSeek-v3

deepseek-v3

64K

USD0.99

USD0.396

USD11.9

USD4.75

Qwen VL

Model Name

Model Code

Max Input Tokens

Postpaid Input

Per 10K TPM/Hour

Postpaid Output

Per 1K TPM/Hour

Prepaid Input

Per 10K TPM/Day

Prepaid Output

Per 1K TPM/Day

Qwen3-VL-Plus-2025-09-23

qwen3-vl-plus-2025-09-23

128K

USD0.35

USD0.35

USD4.2

USD4.2

GLM

Model Name

Model Code

Max Input Tokens

Postpaid Input

Per 10K TPM/Hour

Postpaid Output

Per 1K TPM/Hour

Prepaid Input

Per 10K TPM/Day

Prepaid Output

Per 1K TPM/Day

GLM-5.2

glm-5.2

1M

USD3.96

USD1.386

USD47.53

USD16.635

Long input and prefix caching

Long-input tiered coefficients and cache discounts vary by model. The following table lists the parameters for currently supported models:

Model

Maximum input length

Cache discount

Long-input tiered coefficients

qwen3.8-Max

1 Million

0.125 (cache-hit tokens consume 12.5% capacity)

No tiers (1.0)

qwen3.7-plus-2026-05-26

1 Million

0.2 (cache-hit tokens consume 20% capacity)

No tiers (1.0)

qwen3.7-flash-2026-07-15

1 Million

0.2 (cache-hit tokens consume 20% capacity)

[0, 32K): input 1.0 / output 1.0
(32K, 256K]: input 3.0 / output 3.0

(256K, 1Million]: input 6.0 / output 6.0

glm-5.2

1 Million

0.25 (cache-hit tokens consume 25% capacity)

No tiers (1.0)

glm-5.1

200K

0.2 (cache-hit tokens consume 20% capacity)

[0, 32K): input 1.0 / output 1.0
(32K, 200K]: input 1.33 / output 1.17

deepseek-v4-pro

256K

0.08 (cache-hit tokens consume 8% capacity)

No tiers (1.0)

deepseek-v4-flash-0731

1 Million

0.2 (cache-hit tokens consume 20% capacity)

No tiers (1.0)

Other models

Refer to the console

Refer to the console

No tiers (1.0)

Calculation examples (glm-5.1)

Scenario 1: Short input (10K tokens, no cache)
  Input consumption: 10K × 1.0 = 10 KTPM

Scenario 2: Long input (50K tokens, no cache)
  Input consumption: 32K × 1.0 + 18K × 1.33 = 55.94 KTPM
  Output consumption (assuming 1K tokens): 1K × 1.17 = 1.17 KTPM

Scenario 3: Long input + cache hit (50K tokens, first 30K hit cache)
  Cached input (first 30K, within [0,32K) tier):
    30K × 1.0 × 0.2 = 6 KTPM
  Non-cached input (last 20K):
    2K × 1.0 + 18K × 1.33 = 25.94 KTPM
  Total input = 31.940 KTPM (43% savings compared to no cache)

Provisioned Throughput Quota Calculator

NoteUse the calculator before creating or scaling a deployment to evaluate quota requirements for long-input scenarios. This helps avoid requests falling back to pay-as-you-go billing due to insufficient quota. Refer to the console for actual purchase limits.

Prerequisites: You have activated Model Studio and have PTU deployment permissions. Log on to the Model Studio console, go to the Model Deployment > Create Deployment page (or click Scale Up on an existing deployment's product page), select a deployable PTU (Provisioned Throughput) model, and expand the Provisioned Throughput Quota Calculator.

image

The Provisioned Throughput Quota Calculator automatically recommends the KTPM quota based on your workload. Fill in the following parameters, and the calculator will output the recommended input KTPM and output KTPM.

Parameter

Description

Impact on results

Requests per minute (RPM)

Number of requests per minute during peak business hours.

The larger the RPM, the more the recommended input and output KTPM increases proportionally.

Average input length (tokens)

Average number of input tokens per request.

The longer the input, the higher the tier and coefficient, and the higher the recommended input KTPM. Tier boundaries vary by model; refer to the console for actual values.

Average output length (tokens)

Average number of output tokens per request.

The longer the output, the higher the possible coefficient, and the higher the recommended output KTPM.

Cache hit rate (%)

The percentage of request prefix tokens that hit the cache. Actual hit rate depends on request content repetition; refer to actual runtime results.

A higher hit rate slows input quota consumption, reducing the recommended input KTPM. Affects input KTPM only, not output KTPM.

API response fields

API responses for PTU deployments include the following quota-related fields that identify the billing method and quota consumption:

Field

Type

Description

service_tier

String

A top-level response field (consistent across all API formats). A value of ptu-standard indicates PTU quota is used. A value of default or no return indicates pay-as-you-go billing.

provisioned_tokens

Integer

The actual PTU quota tokens consumed after conversion (includes tiered coefficients and cache discounts).

cached_tokens

Integer

The number of tokens that hit the prefix cache. For details, see Context Cache.

The JSON paths for these fields differ by API format:

OpenAI Chat compatible

Field

JSON path

Description

cached_tokens

usage.prompt_tokens_details.cached_tokens

Input-side cached token count

provisioned_tokens

usage.prompt_tokens_details.provisioned_tokens

Input-side PTU quota consumption

provisioned_tokens

usage.completion_tokens_details.provisioned_tokens

Output-side PTU quota consumption

OpenAI Responses

Field

JSON path

Description

cached_tokens

usage.input_tokens_details.cached_tokens

Input-side cached token count

provisioned_tokens

usage.input_tokens_details.provisioned_tokens

Input-side PTU quota consumption

provisioned_tokens

usage.output_tokens_details.provisioned_tokens

Output-side PTU quota consumption

Anthropic compatible

Field

JSON path

Description

provisioned_tokens

usage.prompt_tokens_details.provisioned_tokens

Input-side PTU quota consumption

provisioned_tokens

usage.output_tokens_details.provisioned_tokens

Output-side PTU quota consumption

NoteThe Anthropic-compatible format does not currently return the cached_tokens field. You can use provisioned_tokens to indirectly evaluate cache effectiveness.

DashScope

Field

JSON path

Description

cached_tokens

usage.prompt_tokens_details.cached_tokens

Input-side cached token count

provisioned_tokens

usage.prompt_tokens_details.provisioned_tokens

Input-side PTU quota consumption

provisioned_tokens

usage.completion_tokens_details.provisioned_tokens

Output-side PTU quota consumption

For complete field definitions and value ranges, see API reference.

Monitoring and verification

You can monitor PTU deployments through the model monitoring feature on the Model Studio platform. The following metrics are relevant to long-input and caching:

  • PTU utilization: Three independent curves for input, output, and thinking mode output. In long-input scenarios, tiered coefficients can cause utilization to exceed 100%. This is expected behavior.
  • Token usage and cache hits: Includes the cached_tokens data series to show the proportion of cache hits relative to total input.
  • In-quota and out-of-quota call counts: Shows the proportion of requests that fall back to pay-as-you-go billing after exceeding PTU quota.

For more monitoring metrics and instructions, see Model monitoring.

Scaling

Click the Scaling button to self-service, manually adjust the throughput (number of instances). For detailed fee reduction and refund rules, please refer to: Refund rules for configuration downgrades.

In addition, you can configure an auto-scaling policy (including scaling thresholds, minimum/maximum replica count, scheduled scaling, etc.) via the scaling configuration button in the operation column.

FAQ

Q: What happens when PTU quota is exceeded?

The request automatically falls back to pay-as-you-go billing. In the API response, the service_tier field is either absent or returns default, and the response header includes x-dashscope-ptu-overflow:true. Service is not interrupted.

Q: What happens when a single input exceeds the model limit?

Qwen-series models have a 128K token input limit, and DeepSeek-series models have a 64K token limit. Requests exceeding the limit also automatically fall back to pay-as-you-go billing.

Q: How do I verify that caching is working?

Check the cached_tokens field in the API response. A value greater than 0 indicates a prefix cache hit. Cache-hit tokens consume quota at the model's discount coefficient (for specific discount rates, see Quota consumption rules). You can also view trends in the token usage chart on the console monitoring page.

Q: What if cached_tokens is always 0 and caching is not working?

Common causes: The input prefix is inconsistent between requests (for example, the System Message changes), the interval between two requests exceeds the cache validity period, or the input token count is too low to trigger caching. For troubleshooting methods and cache limits, see Context Cache.

Q: Why does utilization exceed 100%?

Some models (such as glm-5.1) have long-input tiered coefficients that cause actual quota consumption to exceed the raw token count. Utilization = converted consumption ÷ purchased quota. Exceeding 100% means the consumption rate exceeds the purchased quota. The excess automatically falls back to pay-as-you-go billing without affecting service availability.