This topic describes PTU (Provisioned Throughput) deployment billing rules, supported models and pricing, long input and prefix caching, Provisioned Throughput Quota Calculator usage, and API response fields.
Overview
PTU (Provisioned Throughput Unit) reserves platform resources to guarantee a specific TPM throughput capacity for model deployment; no rate limiting within the guaranteed quota. It is suitable for high-throughput, high-performance, high-load production environments, providing stable throughput capacity, lower latency, and stronger resource certainty. Compared with Token-based billing, TPS (tokens generated per second) typically increases by approximately 1.5 to 2.0 times.
Use cases:
- Intelligent customer service for banking apps (stable traffic, requires guaranteed concurrent experience).
- Real-time content moderation for social platforms (requires stable processing of predictable pipeline tasks).
- Public cloud translation API (provides baseline service guarantees for standard package users).
It is also commonly used in long-document analysis (contracts, research report summaries) and multi-turn conversations (customer service, coding assistants) where input exceeds 32K tokens. When creating a PTU, you can choose an overflow strategy: auto-overflow (default, excess automatically converts to pay-as-you-go) or use-only-PTU-capacity (excess returns 429). See Billing rules for details.
PTU also supports long-input tiered capacity coefficients and cache discounts. See Long input and prefix caching for details. For how prefix caching works, see Context Cache.
Billing rules
Fee = Usage duration × (Input TPM unit price × Input TPM + Output TPM unit price × Output TPM)
Post-paid is calculated hourly: the usage duration unit is hours, and the unit price is taken from the "Continuous 1 hour" column in the table below; prepaid is calculated daily: the usage duration unit is days, and the unit price is taken from the "Continuous 1 day" column in the table below.
- Prepaid orders take effect in real time after payment, with a validity period of N days ending at 23:59 on day N. If the order is placed after 22:00, the expiration date will be automatically extended by 1 day.
- After a prepaid order expires, the service will be stopped with a 2-hour delay, and resources will be retained for 14 hours after the stop and then released.
- Prepaid orders support early termination; the used portion is settled at a 1.2x coefficient for refund. For details, see Refund rules for configuration downgrades.
- For post-paid billing, if the account is in arrears, the deployed resources will continue to be retained and billed for 24 hours, during which the service can still be used normally. After 24 hours, the system stops billing, the model deployment enters an arrears state, and the underlying resources will be deleted, but the model deployment task will be retained. After the arrears are paid, the system will reallocate resources and restore usage (fees will continue to accrue after restoration). If you do not want to continue incurring fees, you can delete the model deployment task, and billing will stop after successful deletion.
When the model input exceeds the maximum input Token, the relevant call will automatically switch to the pay-as-you-go mode of the current model; when the purchased TPM is exceeded, it is handled according to the overflow strategy selected at creation ("auto-overflow" switches to pay-as-you-go, "use-only-PTU-capacity" returns 429). At this time, inference performance may degrade and will be subject to the public traffic control of the current snapshot model in the business space, and fees will be charged according to the model invocation (pay-as-you-go) standard.
- In this case (only under the "auto-overflow" strategy), the API response Header will include:
x-dashscope-ptu-overflow:true. - For TPM statistics, go to: Model Monitoring.
For the specific fee reduction and refund rules in scale-down (downgrade) scenarios, please refer to: Refund rules for configuration downgrades.
NotePTU deployment supports long-input tiered capacity coefficients and cache discounts; see Long input and prefix caching for details.
Supported models and pricing
Singapore
Qwen
Model Name | Model Code | Max Input Tokens | Postpaid Input Per 10K TPM/Hour | Postpaid Output Per 1K TPM/Hour | Prepaid Input Per 10K TPM/Day | Prepaid Output Per 1K TPM/Day |
|---|---|---|---|---|---|---|
Qwen3.8-Max | qwen3.8-max | 1M | USD4.8 | USD1.44 | USD57.6 | USD17.28 |
Qwen3.7-Flash-2026-07-15 Contact your account manager to enable | qwen3.7-flash-2026-07-15 | 128K | USD0.072 | USD0.031 | USD0.864 | USD0.374 |
Qwen3.7-Max-2026-05-20 | qwen3.7-max-2026-05-20 | 256K | USD1.92 | USD1.8 | USD72 | USD21.6 |
Qwen3.7-Plus-2026-05-26 | qwen3.7-plus-2026-05-26 | 256K | USD0.96 | USD0.384 | USD11.52 | USD4.608 |
Qwen3.6-Plus-2026-04-02 | qwen3.6-plus-2026-04-02 | 128K | USD1.2 | USD0.72 | USD14.4 | USD8.64 |
Qwen3.5-Plus-2026-04-20 | qwen3.5-plus-2026-04-20 | 128K | USD0.96 | USD0.576 | USD11.52 | USD6.912 |
DeepSeek
Model Name | Model Code | Max Input Tokens | Postpaid Input Per 10K TPM/Hour | Postpaid Output Per 1K TPM/Hour | Prepaid Input Per 10K TPM/Day | Prepaid Output Per 1K TPM/Day |
|---|---|---|---|---|---|---|
DeepSeek-v4-Flash | deepseek-v4-flash | 256K | USD0.72 | USD0.144 | USD8.64 | USD1.728 |
DeepSeek-v4-Flash-0731 | deepseek-v4-flash-0731 | 64K | USD1.44 | USD0.288 | USD17.28 | USD3.456 |
DeepSeek-v4-Pro | deepseek-v4-pro | 256K | USD0.96 | USD1.728 | USD103.68 | USD20.736 |
Qwen VL
Model Name | Model Code | Max Input Tokens | Postpaid Input Per 10K TPM/Hour | Postpaid Output Per 1K TPM/Hour | Prepaid Input Per 10K TPM/Day | Prepaid Output Per 1K TPM/Day |
|---|---|---|---|---|---|---|
Qwen3-VL-Plus-2025-09-23 | qwen3-vl-plus-2025-09-23 | 128K | USD0.48 | USD0.384 | USD5.76 | USD4.608 |
GLM
Model Name | Model Code | Max Input Tokens | Postpaid Input Per 10K TPM/Hour | Postpaid Output Per 1K TPM/Hour | Prepaid Input Per 10K TPM/Day | Prepaid Output Per 1K TPM/Day |
|---|---|---|---|---|---|---|
GLM-5.2 | glm-5.2 | 1M | USD5.04 | USD1.584 | USD60.48 | USD19.008 |
China North 2 (Beijing)
Qwen
Model Name | Model Code | Max Input Tokens | Postpaid Input Per 10K TPM/Hour | Postpaid Output Per 1K TPM/Hour | Prepaid Input Per 10K TPM/Day | Prepaid Output Per 1K TPM/Day |
|---|---|---|---|---|---|---|
Qwen3.8-Max | qwen3.8-max | 1M | USD3.96 | USD1.188 | USD47.53 | USD14.258 |
Qwen3.7-Flash-2026-07-15 Contact your account manager to enable | qwen3.7-flash-2026-07-15 | 128K | USD0.066 | USD0.026 | USD0.792 | USD0.317 |
Qwen3.7-Max-2026-05-20 | qwen3.7-max-2026-05-20 | 256K | USD3.96 | USD1.188 | USD47.53 | USD14.258 |
Qwen3.7-Plus-2026-05-26 | qwen3.7-plus-2026-05-26 | 256K | USD0.66 | USD0.264 | USD7.92 | USD3.168 |
Qwen3.6-Plus-2026-04-02 | qwen3.6-plus-2026-04-02 | 128K | USD0.67 | USD0.066 | USD7.93 | USD4.753 |
Qwen3.5-Plus-2026-04-20 | qwen3.5-plus-2026-04-20 | 128K | USD0.26 | USD0.16 | USD3.17 | USD1.9 |
Qwen3-Max-2025-09-23 | qwen3-max-2025-09-23 | 128K | USD1.11 | USD0.45 | USD13.32 | USD5.4 |
Qwen-Flash-2025-07-28 | qwen-flash-2025-07-28 | 128K | USD0.06 | USD0.06 | USD0.72 | USD0.72 |
Qwen-Plus-2025-12-01 | qwen-plus-2025-12-01 | 128K | USD0.28 | Non-thinking: USD0.07 Thinking: USD0.28 | USD3.36 | Non-thinking: USD0.84 Thinking: USD3.36 |
DeepSeek
Model Name | Model Code | Max Input Tokens | Postpaid Input Per 10K TPM/Hour | Postpaid Output Per 1K TPM/Hour | Prepaid Input Per 10K TPM/Day | Prepaid Output Per 1K TPM/Day |
|---|---|---|---|---|---|---|
DeepSeek-v4-Flash | deepseek-v4-flash | 256K | USD0.5 | USD0.099 | USD5.94 | USD1.188 |
DeepSeek-v4-Flash-0731 | deepseek-v4-flash-0731 | 64K | USD0.99 | USD0.198 | USD11.88 | USD2.376 |
DeepSeek-v4-Pro | deepseek-v4-pro | 256K | USD5.94 | USD1.188 | USD71.3 | USD14.26 |
DeepSeek-v3 | deepseek-v3 | 64K | USD0.99 | USD0.396 | USD11.9 | USD4.75 |
Qwen VL
Model Name | Model Code | Max Input Tokens | Postpaid Input Per 10K TPM/Hour | Postpaid Output Per 1K TPM/Hour | Prepaid Input Per 10K TPM/Day | Prepaid Output Per 1K TPM/Day |
|---|---|---|---|---|---|---|
Qwen3-VL-Plus-2025-09-23 | qwen3-vl-plus-2025-09-23 | 128K | USD0.35 | USD0.35 | USD4.2 | USD4.2 |
GLM
Model Name | Model Code | Max Input Tokens | Postpaid Input Per 10K TPM/Hour | Postpaid Output Per 1K TPM/Hour | Prepaid Input Per 10K TPM/Day | Prepaid Output Per 1K TPM/Day |
|---|---|---|---|---|---|---|
GLM-5.2 | glm-5.2 | 1M | USD3.96 | USD1.386 | USD47.53 | USD16.635 |
Long input and prefix caching
Long-input tiered coefficients and cache discounts vary by model. The following table lists the parameters for currently supported models:
Model | Maximum input length | Cache discount | Long-input tiered coefficients |
|---|---|---|---|
qwen3.8-Max | 1 Million | 0.125 (cache-hit tokens consume 12.5% capacity) | No tiers (1.0) |
qwen3.7-plus-2026-05-26 | 1 Million | 0.2 (cache-hit tokens consume 20% capacity) | No tiers (1.0) |
qwen3.7-flash-2026-07-15 | 1 Million | 0.2 (cache-hit tokens consume 20% capacity) | [0, 32K): input 1.0 / output 1.0 (256K, 1Million]: input 6.0 / output 6.0 |
glm-5.2 | 1 Million | 0.25 (cache-hit tokens consume 25% capacity) | No tiers (1.0) |
glm-5.1 | 200K | 0.2 (cache-hit tokens consume 20% capacity) | [0, 32K): input 1.0 / output 1.0 |
deepseek-v4-pro | 256K | 0.08 (cache-hit tokens consume 8% capacity) | No tiers (1.0) |
deepseek-v4-flash-0731 | 1 Million | 0.2 (cache-hit tokens consume 20% capacity) | No tiers (1.0) |
Other models | Refer to the console | Refer to the console | No tiers (1.0) |
Calculation examples (glm-5.1)
Scenario 1: Short input (10K tokens, no cache)
Input consumption: 10K × 1.0 = 10 KTPM
Scenario 2: Long input (50K tokens, no cache)
Input consumption: 32K × 1.0 + 18K × 1.33 = 55.94 KTPM
Output consumption (assuming 1K tokens): 1K × 1.17 = 1.17 KTPM
Scenario 3: Long input + cache hit (50K tokens, first 30K hit cache)
Cached input (first 30K, within [0,32K) tier):
30K × 1.0 × 0.2 = 6 KTPM
Non-cached input (last 20K):
2K × 1.0 + 18K × 1.33 = 25.94 KTPM
Total input = 31.940 KTPM (43% savings compared to no cache)
Provisioned Throughput Quota Calculator
NoteUse the calculator before creating or scaling a deployment to evaluate quota requirements for long-input scenarios. This helps avoid requests falling back to pay-as-you-go billing due to insufficient quota. Refer to the console for actual purchase limits.
Prerequisites: You have activated Model Studio and have PTU deployment permissions. Log on to the Model Studio console, go to the Model Deployment > Create Deployment page (or click Scale Up on an existing deployment's product page), select a deployable PTU (Provisioned Throughput) model, and expand the Provisioned Throughput Quota Calculator.
The Provisioned Throughput Quota Calculator automatically recommends the KTPM quota based on your workload. Fill in the following parameters, and the calculator will output the recommended input KTPM and output KTPM.
Parameter | Description | Impact on results |
|---|---|---|
Requests per minute (RPM) | Number of requests per minute during peak business hours. | The larger the RPM, the more the recommended input and output KTPM increases proportionally. |
Average input length (tokens) | Average number of input tokens per request. | The longer the input, the higher the tier and coefficient, and the higher the recommended input KTPM. Tier boundaries vary by model; refer to the console for actual values. |
Average output length (tokens) | Average number of output tokens per request. | The longer the output, the higher the possible coefficient, and the higher the recommended output KTPM. |
Cache hit rate (%) | The percentage of request prefix tokens that hit the cache. Actual hit rate depends on request content repetition; refer to actual runtime results. | A higher hit rate slows input quota consumption, reducing the recommended input KTPM. Affects input KTPM only, not output KTPM. |
API response fields
API responses for PTU deployments include the following quota-related fields that identify the billing method and quota consumption:
Field | Type | Description |
|---|---|---|
| String | A top-level response field (consistent across all API formats). A value of |
| Integer | The actual PTU quota tokens consumed after conversion (includes tiered coefficients and cache discounts). |
| Integer | The number of tokens that hit the prefix cache. For details, see Context Cache. |
The JSON paths for these fields differ by API format:
OpenAI Chat compatible
Field | JSON path | Description |
|---|---|---|
|
| Input-side cached token count |
|
| Input-side PTU quota consumption |
|
| Output-side PTU quota consumption |
OpenAI Responses
Field | JSON path | Description |
|---|---|---|
|
| Input-side cached token count |
|
| Input-side PTU quota consumption |
|
| Output-side PTU quota consumption |
Anthropic compatible
Field | JSON path | Description |
|---|---|---|
|
| Input-side PTU quota consumption |
|
| Output-side PTU quota consumption |
NoteThe Anthropic-compatible format does not currently return the cached_tokens field. You can use provisioned_tokens to indirectly evaluate cache effectiveness.
DashScope
Field | JSON path | Description |
|---|---|---|
|
| Input-side cached token count |
|
| Input-side PTU quota consumption |
|
| Output-side PTU quota consumption |
For complete field definitions and value ranges, see API reference.
Monitoring and verification
You can monitor PTU deployments through the model monitoring feature on the Model Studio platform. The following metrics are relevant to long-input and caching:
- PTU utilization: Three independent curves for input, output, and thinking mode output. In long-input scenarios, tiered coefficients can cause utilization to exceed 100%. This is expected behavior.
- Token usage and cache hits: Includes the
cached_tokensdata series to show the proportion of cache hits relative to total input. - In-quota and out-of-quota call counts: Shows the proportion of requests that fall back to pay-as-you-go billing after exceeding PTU quota.
For more monitoring metrics and instructions, see Model monitoring.
Scaling
Click the Scaling button to self-service, manually adjust the throughput (number of instances). For detailed fee reduction and refund rules, please refer to: Refund rules for configuration downgrades.
In addition, you can configure an auto-scaling policy (including scaling thresholds, minimum/maximum replica count, scheduled scaling, etc.) via the scaling configuration button in the operation column.
FAQ
Q: What happens when PTU quota is exceeded?
The request automatically falls back to pay-as-you-go billing. In the API response, the service_tier field is either absent or returns default, and the response header includes x-dashscope-ptu-overflow:true. Service is not interrupted.
Q: What happens when a single input exceeds the model limit?
Qwen-series models have a 128K token input limit, and DeepSeek-series models have a 64K token limit. Requests exceeding the limit also automatically fall back to pay-as-you-go billing.
Q: How do I verify that caching is working?
Check the cached_tokens field in the API response. A value greater than 0 indicates a prefix cache hit. Cache-hit tokens consume quota at the model's discount coefficient (for specific discount rates, see Quota consumption rules). You can also view trends in the token usage chart on the console monitoring page.
Q: What if cached_tokens is always 0 and caching is not working?
Common causes: The input prefix is inconsistent between requests (for example, the System Message changes), the interval between two requests exceeds the cache validity period, or the input token count is too low to trigger caching. For troubleshooting methods and cache limits, see Context Cache.
Q: Why does utilization exceed 100%?
Some models (such as glm-5.1) have long-input tiered coefficients that cause actual quota consumption to exceed the raw token count. Utilization = converted consumption ÷ purchased quota. Exceeding 100% means the consumption rate exceeds the purchased quota. The excess automatically falls back to pay-as-you-go billing without affecting service availability.