Others

Model Studio PTU (Provisioned Throughput) Capability Upgrade Announcement

Affected Time

2026-06-18 23:00:00 (UTC+08) [Subject to the actual change time]

I. Upgrade Overview

To better support enterprise customers in handling long-context and multi-turn conversation

workloads, Alibaba Cloud Model Studio PTU has undergone five key capability upgrades. These

upgrades cover critical scenarios including long-input support, implicit caching, capacity

calculation, and API observability. Following these upgrades, while maintaining the same unit

pricing and purchasing models, PTU can support a broader range of workloads and deliver

higher effective output per unit of capacity.

Existing PTU customers do not need to take any action. The new rules will be automatically

applied upon completion of the upgrade, further expanding the range of workloads that your

existing PTU capacity can handle.

II. Upgrade Details

1. PTU Supports Longer Inputs, Aligned with API Limits

Previously, PTU enforced a hard limit on input length, and requests exceeding this limit were

automatically routed to the shared resource pool for pay-as-you-go billing. Following this

upgrade, the maximum input length for the following models within PTU has been significantly

increased:

图一英文

The scope of supported models will continue to expand in alignment with the model iteration

schedule.

2. PTU Supports Implicit Caching

PTU now supports the automatic reuse of KV Cache for request input prefixes. Tokens that hit

the cache consume capacity at a reduced cache coefficient, requiring no code modifications.

The models supporting implicit caching and their corresponding cache coefficients in this release

are as follows:

图二英文

Typical use cases that benefit from this include RAG knowledge base Q&A, multi-turn

conversations, Agent/tool calling, and batch processing. For workloads featuring fixed prefixes

(such as System Prompts, retrieved chunks, and tool definitions), the request throughput per unit

of PTU capacity will significantly increase.

3. Introduction of Capacity Coefficient for Long Inputs in Selected Models

To more accurately reflect the inference cost differences between long and short inputs, PTU

introduces a capacity coefficient mechanism: Actual Capacity Consumption = Original Token

Count × Capacity Coefficient. Long-input requests are processed within the same PTU,

eliminating the need to purchase dedicated long-input capacity.

In this release, this applies only to GLM-5.1:

图三英文

The unit pricing and purchasing process for PTU remain unchanged.

4. New Capacity Calculator in the Console

A new "Capacity Calculator" has been integrated into the PTU creation workflow. Customers can

input their business-specific metrics, including RPM, average input length, average output

length, and cache hit rate. The system will then automatically calculate and recommend the

required TPM based on the capacity coefficients, significantly reducing the effort and cost of

capacity evaluation prior to purchase.

5. New PTU Identifier and Capacity Fields in API Responses

To facilitate the monitoring of PTU hit rates and the breakdown of actual capacity consumption,

the following fields have been added to the API responses:

图四英文

For response details, please refer to the official documentation.

III. Impact on Existing Customers

Upon completion of the upgrade:

Existing PTUs will automatically adopt the new rules without the need for recreation or migration;

Requests that previously fell back to the shared resource pool due to exceeding input length limits will continue to be processed within the PTU and enjoy SLA guarantees, provided that the capacity consumption after applying the coefficients remains within the PTU limits;

The implicit caching feature will take effect simultaneously for existing PTUs, automatically reducing the capacity consumption for requests with identical prefixes;

The scope of workloads that can be supported by customers' existing PTUs will be further expanded.

For details on billing and capacity rules, please refer to the "PTU Billing Guide" in the official

Model Studio documentation. For technical support, please submit a support ticket.