Model Studio PTU (Provisioned Throughput) Capability Upgrade Announcement
Jun 22, 2026
Alibaba Cloud Model StudioAffected Time
I. Upgrade Overview
To better support enterprise customers in handling long-context and multi-turn conversation
workloads, Alibaba Cloud Model Studio PTU has undergone five key capability upgrades. These
upgrades cover critical scenarios including long-input support, implicit caching, capacity
calculation, and API observability. Following these upgrades, while maintaining the same unit
pricing and purchasing models, PTU can support a broader range of workloads and deliver
higher effective output per unit of capacity.
Existing PTU customers do not need to take any action. The new rules will be automatically
applied upon completion of the upgrade, further expanding the range of workloads that your
existing PTU capacity can handle.
II. Upgrade Details
1. PTU Supports Longer Inputs, Aligned with API Limits
Previously, PTU enforced a hard limit on input length, and requests exceeding this limit were
automatically routed to the shared resource pool for pay-as-you-go billing. Following this
upgrade, the maximum input length for the following models within PTU has been significantly
increased:

The scope of supported models will continue to expand in alignment with the model iteration
schedule.
2. PTU Supports Implicit Caching
PTU now supports the automatic reuse of KV Cache for request input prefixes. Tokens that hit
the cache consume capacity at a reduced cache coefficient, requiring no code modifications.
The models supporting implicit caching and their corresponding cache coefficients in this release
are as follows:

Typical use cases that benefit from this include RAG knowledge base Q&A, multi-turn
conversations, Agent/tool calling, and batch processing. For workloads featuring fixed prefixes
(such as System Prompts, retrieved chunks, and tool definitions), the request throughput per unit
of PTU capacity will significantly increase.
3. Introduction of Capacity Coefficient for Long Inputs in Selected Models
To more accurately reflect the inference cost differences between long and short inputs, PTU
introduces a capacity coefficient mechanism: Actual Capacity Consumption = Original Token
Count × Capacity Coefficient. Long-input requests are processed within the same PTU,
eliminating the need to purchase dedicated long-input capacity.
In this release, this applies only to GLM-5.1:

The unit pricing and purchasing process for PTU remain unchanged.
4. New Capacity Calculator in the Console
A new "Capacity Calculator" has been integrated into the PTU creation workflow. Customers can
input their business-specific metrics, including RPM, average input length, average output
length, and cache hit rate. The system will then automatically calculate and recommend the
required TPM based on the capacity coefficients, significantly reducing the effort and cost of
capacity evaluation prior to purchase.
5. New PTU Identifier and Capacity Fields in API Responses
To facilitate the monitoring of PTU hit rates and the breakdown of actual capacity consumption,
the following fields have been added to the API responses:

For response details, please refer to the official documentation.
III. Impact on Existing Customers
Upon completion of the upgrade:
Existing PTUs will automatically adopt the new rules without the need for recreation or migration;
Requests that previously fell back to the shared resource pool due to exceeding input length limits will continue to be processed within the PTU and enjoy SLA guarantees, provided that the capacity consumption after applying the coefficients remains within the PTU limits;
The implicit caching feature will take effect simultaneously for existing PTUs, automatically reducing the capacity consumption for requests with identical prefixes;
The scope of workloads that can be supported by customers' existing PTUs will be further expanded.
For details on billing and capacity rules, please refer to the "PTU Billing Guide" in the official
Model Studio documentation. For technical support, please submit a support ticket.