All Products
Search
Document Center

Alibaba Cloud Model Studio:Throughput Reservation Billing

Last Updated:Sep 22, 2026

This document describes the billing rules, pricing, refunds, and capacity conversion rules of Throughput Reservation (formerly TPM Reservation) on the Alibaba Cloud Bailian platform.

Billing Rules

Throughput Reservation purchases dedicated inference throughput via Subscription. Calls within the reserved capacity incur no additional charges, and the exceeding portion is handled per the overflow strategy. The dedicated Model Code (invocation identifier, used to replace the model parameter of the API) generated after creating a Throughput Reservation is used for billing and invocation statistics.

Performance Mode

Standard mode: same TPS as the standard API.

High-Speed mode: 1.5~2 Times TPS improvement over the standard API (i.e., PTU model deployment).

Billing Method

Subscription (By Day): supported by both Standard mode and High-Speed mode, billed by natural day. Input is priced per 10k TPM and output per 1k TPM.

Subscription (By 8-Hour Block): only Standard mode, fixed 8 hours, effective immediately. For more information, see 8-Hour Block Reservation Billing.

Pay-as-you-go (By Hour): only High-Speed mode, billed by actual duration, no need to purchase a Duration.

Billing Formula

Subscription: Reserved Fee = Duration ( Days) × (Input Unit price × Input kTPM ÷ 10 + Output Unit price × Output kTPM)

Pay-as-you-go (hour granularity, minute precision): Pay-as-you-go Fee = (Usage Minutes ÷ 60) × (Pay-as-you-go Input Unit price × Input 10k TPM + Pay-as-you-go Output Unit price × Output 1k TPM)

Billed on Activation

Billing starts as soon as creation succeeds; calls within the reserved capacity incur no additional charges; capacity fees are billed by the purchased quantity and are not reduced based on actual usage.

New Purchase Expiration Time

UTC+8: if purchased Before 22:00 on the current day, the expiration time is 23:59:59 of the current day.

If purchased after 22:00, the expiration time is 23:59:59 of the next day.

Input Capacity

Starting at 200 kTPM, with a step size of 10 kTPM.

Output Capacity

Starting at 20 kTPM, with a step size of 1 kTPM.

Duration

Supports 1~30, 60, 90, 120, 365 Days.

Renew

Fee = Number of days × New purchase price.

Auto-Renew on Expiration

Auto-deducts and renews 1 day Before expiration, with up to 4 attempts between 8:20–20:00 on that day. Enabled by default, can be turned off.

Standard mode

Take Qwen3.8-Max (China (Beijing)) as an example (200 Input + 100 Output kTPM, 1 Days):

  1. Assume Standard mode Input Unit price is 16.63/10kTPMday,OutputUnitpriceis16.63/10k TPM·day, Output Unit price is 4.99/1k TPM·day.
  2. Input Fee = 16.63 × 200 ÷ 10 = 332.6 USD.
  3. Output Fee = 4.99 × 100 = 499 USD.
  4. Reserved Fee = 1 × (332.6 + 499) = 831.6 USD.

High-Speed mode

Take Qwen3.8-Max (China (Beijing)) as an example (200 Input + 100 Output kTPM, 1 Days):

  1. Assume High-Speed mode Subscription Input Unit price is 47.53/10kTPMday,OutputUnitpriceis47.53/10k TPM·day, Output Unit price is 14.258/1k TPM·day.
  2. Input Fee = 47.53 × 200 ÷ 10 = 950.6 USD.
  3. Output Fee = 14.258 × 100 = 1425.8 USD.
  4. Reserved Fee = 1 × (950.6 + 1425.8) = 2376.4 USD.

Unit Description

  • kTPM = 1,000 Tokens/minute, the unit of purchase quantity.
  • 10k TPM = 10,000 Tokens/minute = 10 kTPM. The Input Unit price is quoted per 10k TPM; in the formula, kTPM ÷ 10 converts to 10k TPM before multiplying by the unit price. The Output Unit price is quoted per 1k TPM (1 1k TPM = 1 kTPM) and is multiplied directly.

Supported Models and Pricing

The Subscription pricing for Standard mode (same TPS as pay-as-you-go calls) and High-Speed mode (1.5~2 Times TPS) is as follows

China (Beijing)

Standard mode

Model

Subscription

Input (10k TPM·day)

Output (1k TPM·day)

Qwen3.8-Max

$16.63

$4.99

Qwen3.7-Flash-2026-07-15

$0.28

$0.11

Qwen3.7-Max-2026-05-20

$16.63

$4.99

Qwen3.7-Plus-2026-05-26

$2.77

$1.11

Qwen3.6-Flash-2026-04-16

$1.66

$1.00

GLM-5.3

$11.40

$3.99

GLM-5.2

$11.09

$3.88

GLM-5.1

$8.32

$3.33

DeepSeek-v4-Flash

$1.39

$0.28

DeepSeek-v4-Flash-0731

$4.27

$1.28

DeepSeek-v4-Pro

$16.63

$3.33

DeepSeek-v4-Pro-0813

$12.82

$3.85

Kimi-K2.6

$9.01

$3.74

High-Speed mode

Model

Pay-as-you-go

Subscription

Input (10k TPM·hour)

Output (1k TPM·hour)

Input (10k TPM·day)

Output (1k TPM·day)

Qwen3.8-Max

$3.96

$1.188

$47.53

$14.258

Qwen3.7-Flash-2026-07-15

$0.066

$0.026

$0.792

$0.317

Qwen3.7-Max-2026-05-20

$3.96

$1.188

$47.53

$14.258

Qwen3.7-Plus-2026-05-26

$0.66

$0.264

$7.92

$3.168

Qwen3.6-Flash-2026-04-16

$0.4

$0.238

$4.75

$2.852

GLM-5.3

GLM-5.2

$3.96

$1.386

$47.53

$16.635

GLM-5.1

DeepSeek-v4-Flash

$0.5

$0.099

$5.94

$1.188

DeepSeek-v4-Flash-0731

$0.99

$0.198

$11.88

$2.376

DeepSeek-v4-Pro

$5.94

$1.188

$71.3

$14.26

DeepSeek-v4-Pro-0813

Kimi-K2.6

Singapore

Standard mode

Model

Subscription

Input (10k TPM·day)

Output (1k TPM·day)

Qwen3.8-Max

$20.16

$6.05

Qwen3.7-Flash-2026-07-15

$0.30

$0.13

Qwen3.7-Max-2026-05-20

$25.20

$7.56

Qwen3.7-Plus-2026-05-26

$4.03

$1.61

Qwen3.6-Flash-2026-04-16

$2.52

$1.51

GLM-5.3

$14.11

$4.43

GLM-5.2

$14.11

$4.43

GLM-5.1

$14.11

$4.43

DeepSeek-v4-Flash

$2.02

$0.40

DeepSeek-v4-Flash-0731

$4.44

$1.33

DeepSeek-v4-Pro

$24.19

$4.84

DeepSeek-v4-Pro-0813

$13.31

$3.99

High-Speed mode

Model

Pay-as-you-go

Subscription

Input (10k TPM·hour)

Output (1k TPM·hour)

Input (10k TPM·day)

Output (1k TPM·day)

Qwen3.8-Max

$4.8

$1.44

$57.6

$17.28

Qwen3.7-Flash-2026-07-15

$0.072

$0.031

$0.864

$0.374

Qwen3.7-Max-2026-05-20

$1.92

$1.8

$72

$21.6

Qwen3.7-Plus-2026-05-26

$0.96

$0.384

$11.52

$4.608

Qwen3.6-Flash-2026-04-16

GLM-5.3

GLM-5.2

$5.04

$1.584

$60.48

$19.008

GLM-5.1

DeepSeek-v4-Flash

$0.72

$0.144

$8.64

$1.728

DeepSeek-v4-Flash-0731

$1.44

$0.288

$17.28

$3.456

DeepSeek-v4-Pro

$0.96

$1.728

$103.68

$20.736

DeepSeek-v4-Pro-0813

NoteThe cache hit rate only affects Input kTPM, not Output.

Capacity Conversion Rules

Long input is converted by a tiered coefficient, and the cache-hit portion is converted by a cache conversion coefficient; the converted amount is deducted from the purchased quota. The parameters of each model are as follows.

Model

Max Input Token

Cache Conversion Coefficient

Long Input Tier Coefficient

qwen3.8-max

1M

0.125

No tier (1.0)

qwen3.7-flash-2026-07-15

1M

0.2

Same for Input and Output
(0,32K] 1x
(32K,256K] 3x
(256,1m] 6x

qwen3.7-max-2026-05-20

256K

0.1

No tier (1.0)

qwen3.7-plus-2026-05-26

256K

0.2

No tier (1.0)

qwen3.6-flash-2026-04-16

256K

1

No tier (1.0)

glm-5.3

1M

0.25

No tier (1.0)

glm-5.2

1M

0.25

No tier (1.0)

glm-5.1

200K

0.2

(0,32K] Input 1x / Output 1x
(32K,200K] Input 1.33x / Output 1.17x

deepseek-v4-flash

256K

0.1

No tier (1.0)

deepseek-v4-flash-0731

1M

0.1

No tier (1.0)

deepseek-v4-pro

256K

0.08

No tier (1.0)

deepseek-v4-pro-0813

1M

0.1

No tier (1.0)

kimi-k2.6

256K

0.2

No tier (1.0)

Capacity conversion examples are as follows.

With Tier·Qwen3.7-Flash

Long input of 50K Tokens, no cache hit.

  1. The first 32K portion (coefficient 1x): 32K × 1 = 32 kTPM.
  2. The portion exceeding 32K (50K − 32K = 18K, coefficient 3x): 18K × 3 = 54 kTPM.
  3. Total consumption: 32 + 54 = 86 kTPM.

For the same 50K Tokens input, where the first 30K hits the cache (cache conversion coefficient 0.2):

  1. Cache-hit portion (within the first 32K tier, coefficient 1x, then multiplied by conversion coefficient 0.2): 30K × 1 × 0.2 = 6 kTPM.
  2. Non-hit portion (remaining 2K within the first 32K tier = 32K − 30K, coefficient 1x): 2K × 1 = 2 kTPM.
  3. The portion exceeding 32K (50K − 32K = 18K, coefficient 3x): 18K × 3 = 54 kTPM.
  4. Total consumption: 6 + 2 + 54 = 62 kTPM (about 28% less than without cache).

No Tier·GLM-5.2

Long input of 50K Tokens, no cache hit.

  1. No tier, all converted by coefficient 1x: 50K × 1 = 50 kTPM.

For the same 50K Tokens input, where the first 30K hits the cache (cache conversion coefficient 0.25):

  1. Cache-hit portion (coefficient 1x, then multiplied by conversion coefficient 0.25): 30K × 0.25 = 7.5 kTPM.
  2. Non-hit portion (remaining 20K = 50K − 30K, coefficient 1x): 20K × 1 = 20 kTPM.
  3. Total consumption: 7.5 + 20 = 27.5 kTPM (about 45% less than without cache).

8-Hour Block Reservation Billing

8-Hour Block Reservation (8h Block) provides deterministic TPM capacity for 8 consecutive hours for scenarios where the load is concentrated in a specific time period.

Purchase and Billing Rules

  • Billing Method: Subscription (By 8-Hour Block), only Standard mode; 10% off, subject to the actual display in the Bailian Console.
  • Orderable time period: daily 20:00–04:00 the next day (the order time must fall within this period).
  • Effective: effective immediately after purchase, with a fixed duration of 8 hours, and can cross natural days; the effective starting point is rounded down to the current hour based on the purchase time and cannot be customized; less than 1 hour is counted as 1 hour.
  • Two purchase forms:
    • Create a new independent reservation: generates a dedicated Model Code.
    • Add to an existing reservation as an Add-on Capacity Package: reuses the base reservation's dedicated Model Code and increases capacity within that 8-hour window. For more information, see Add-on Capacity Package Billing.

Purchase 8-Hour Block Reservation

Take Qwen3.8-Max (China (Beijing)) Standard mode as an example (200 Input + 100 Output kTPM, 8-hour block, fee multiplier 0.9; assume Input Unit price 16.63/10kTPMday,OutputUnitprice16.63/10k TPM·day, Output Unit price 4.99/1k TPM·day):

  1. By-day fee = 16.63 × 200 ÷ 10 + 4.99 × 100 = 332.6 + 499 = 831.6 USD (1 day).
  2. 8-hour block fee = 831.6 × (8 ÷ 24) × 0.9 ≈ 249.48 USD (10% off).

Add 8-Hour Add-on Capacity Package

Add an 8-hour add-on capacity package on top of an existing by-day reservation. Take Qwen3.8-Max (China (Beijing)) Standard mode as an example (add 200 Input + 100 Output kTPM of 8-hour capacity, fee multiplier 0.9; assume Input Unit price 16.63/10kTPMday,OutputUnitprice16.63/10k TPM·day, Output Unit price 4.99/1k TPM·day):

  1. Added capacity by-day fee = 16.63 × 200 ÷ 10 + 4.99 × 100 = 831.6 USD (1-day basis).
  2. 8-hour add-on package fee = 831.6 × (8 ÷ 24) × 0.9 ≈ 249.48 USD.
  3. After adding, the capacity of the same dedicated Model Code increases within that 8-hour window, and the add-on capacity is invalid outside the window.

Limitations

  • Scale Out/Scale In and Renew/Unsubscribe are not supported.
  • Capacity automatically becomes invalid after the block expires.
  • Orders cannot be placed between 04:00–20:00 on the current day.

Expiration and Lifecycle

After the service expires, it stops immediately, and resources are released 2 hours after the stop.

Phase (After Expiration)

Instance Status

Can Call/Renew

0~2 Hours

Stopped

Cannot be called, but can still be renewed

After 2 Hours

Released

Cannot be recovered

It is recommended to enable Auto-Renew on Expiration in advance to avoid service interruption.

Scale Out

Subscription Instances

  • The Subscription purchase granularity is by day, and the billing precision is by hour.
  • The newly added capacity from Scale Out is billed by the remaining validity period, converted to days: remaining hours ÷ 24 rounded up.
  • Remaining less than 1 day (24 hours) is still counted as 1 day (e.g., 5 hours remaining is counted as 1 day).

Scale Out Fee = Rounded up(Remaining hours ÷ 24) × (Input Capacity difference(kTPM) ÷ 10 × Input Unit price + Output Capacity difference(kTPM) × Output Unit price)

Remaining 5 Hours·Scale Out 100 Input+50 Output kTPM

Take Qwen3.8-Max (China (Beijing)) Standard mode as an example (assume daily Unit price $16.63/10k TPM·day, Output Unit price $4.99/1k TPM·day), original 100 Input + 50 Output kTPM, Scale Out 100 Input + 50 Output kTPM, with 5 hours remaining:

  1. Remaining Days = Rounded up(5 ÷ 24) = 1 Days (5 hours is less than 1 day, rounded up).
  2. Input Scale Out Fee = 1 × 100 ÷ 10 × 16.63 = 166.3 USD.
  3. Output Scale Out Fee = 1 × 50 × 4.99 = 249.5 USD.
  4. Scale Out Fee = 166.3 + 249.5 = 415.8 USD.

Remaining 500 Hours·Scale Out 100 Input+50 Output kTPM

Take Qwen3.8-Max (China (Beijing)) Standard mode as an example (assume daily Unit price $16.63/10k TPM·day, Output Unit price $4.99/1k TPM·day), original 100 Input + 50 Output kTPM, Scale Out 100 Input + 50 Output kTPM, with 500 hours remaining:

  1. Remaining Days = Rounded up(500 ÷ 24) = 21 Days.
  2. Input Scale Out Fee = 21 × 100 ÷ 10 × 16.63 = 3492.3 USD.
  3. Output Scale Out Fee = 21 × 50 × 4.99 = 5239.5 USD.
  4. Scale Out Fee = 3492.3 + 5239.5 = 8731.8 USD.

Pay-as-you-go Instances

  • The Pay-as-you-go purchase granularity is by hour, and the billing precision is by minute.
  • Scale Out is billed by actual usage duration; less than 1 hour (60 minutes) is billed by actual minutes.
  • Fee = Usage Minutes ÷ 60 × Unit price × Capacity (e.g., 30 minutes billed as 0.5 hours).

Scale Out Fee = (Usage Minutes ÷ 60) × (Pay-as-you-go Input Unit price × Input Capacity difference(kTPM) ÷ 10 + Pay-as-you-go Output Unit price × Output Capacity difference(kTPM))

Used 30 Minutes·Less than 1 Hour

Take Qwen3.8-Max (China (Beijing)) High-Speed mode as an example (assume Pay-as-you-go Input Unit price $3.96/10k TPM·hour, Output Unit price $1.188/1k TPM·hour), Scale Out 100 Input + 50 Output kTPM, used 30 minutes:

  1. Input Scale Out Fee = (30 ÷ 60) × 3.96 × (100 ÷ 10) = 0.5 × 3.96 × 10 = 19.8 USD.
  2. Output Scale Out Fee = (30 ÷ 60) × 1.188 × 50 = 0.5 × 1.188 × 50 = 29.7 USD.
  3. Scale Out Fee = 19.8 + 29.7 = 49.5 USD.

Used 10 Hours

Take Qwen3.8-Max (China (Beijing)) High-Speed mode as an example (assume Pay-as-you-go Input Unit price $3.96/10k TPM·hour, Output Unit price $1.188/1k TPM·hour), Scale Out 100 Input + 50 Output kTPM, used 10 hours:

  1. Input Scale Out Fee = (600 ÷ 60) × 3.96 × (100 ÷ 10) = 10 × 3.96 × 10 = 396 USD.
  2. Output Scale Out Fee = (600 ÷ 60) × 1.188 × 50 = 10 × 1.188 × 50 = 594 USD.
  3. Scale Out Fee = 396 + 594 = 990 USD.

Scale In

Subscription Instances

  • The Subscription purchase granularity is by day, and the billing precision is by minute (the settled amount is rounded up by hour).
  • Settled Amount = Rounded up used hours × Penalty coefficient.
  • New specification fee is precise to the minute; ≤30 Days of usage is multiplied by 1.2 Times, >30 Days by 1.0 Times.

Refund = max(Original specification Remaining fee - New specification purchase fee, 0)

  • Original specification Remaining fee = Amount actually paid - Settled Amount
  • Settled Amount = (Daily Unit price ÷ 24) × Rounded up(used hours) × Penalty coefficient × Original Capacity(kTPM) ÷ 10
  • New specification purchase fee = Daily Unit price × (Remaining hours ÷ 24) × New Capacity(kTPM) ÷ 10

NoteSetting capacity to zero (capacity adjusted to 0) no longer incurs capacity fees; the dedicated Model Code is retained, and this is treated as a reduction settled by the penalty coefficient; remaining hours are precise to the minute (e.g., 27 hours 37 minutes ≈ 27.62 hours).

Purchased 30 Days·Used 20 Days (≤30 Days, 1.2 Times)

Take Qwen3.8-Max (China (Beijing)) Standard mode as an example (assume daily Unit price $16.63/10k TPM·day, Output Unit price $4.99/1k TPM·day), original 100 kTPM purchased 30 Days, scaled in to 50 kTPM. Amount actually paid 4989 USD = 100 ÷ 10 × 30 × 16.63, used 20 Days (480 hours), 240 hours remaining:

  1. Settled Amount = (16.63 ÷ 24) × 480 × 1.2 × 100 ÷ 10 = 3991.2 USD.
  2. Original specification Remaining fee = 4989 - 3991.2 = 997.8 USD.
  3. New specification purchase fee = 16.63 × (240 ÷ 24) × 50 ÷ 10 = 831.5 USD.
  4. Refund = max(997.8 - 831.5, 0) = 166.3 USD.

Purchased 60 Days·Used 40 Days (>30 Days, 1.0 Times)

Take Qwen3.8-Max (China (Beijing)) Standard mode as an example (assume daily Unit price $16.63/10k TPM·day, Output Unit price $4.99/1k TPM·day), original 100 kTPM purchased 60 Days, scaled in to 50 kTPM. Amount actually paid 9978 USD = 100 ÷ 10 × 60 × 16.63, used 40 Days (960 hours), 480 hours remaining:

  1. Settled Amount = (16.63 ÷ 24) × 960 × 1.0 × 100 ÷ 10 = 6652 USD.
  2. Original specification Remaining fee = 9978 - 6652 = 3326 USD.
  3. New specification purchase fee = 16.63 × (480 ÷ 24) × 50 ÷ 10 = 1663 USD.
  4. Refund = max(3326 - 1663, 0) = 1663 USD.

Set to Zero (Capacity Adjusted to 0)

Take Qwen3.8-Max (China (Beijing)) Standard mode as an example (assume daily Unit price $16.63/10k TPM·day, original 100 kTPM, purchased 30 Days, used 20 Days):

  1. Settled Amount = (16.63 ÷ 24) × 480 × 1.2 × 100 ÷ 10 = 3991.2 USD (used 20 Days ≤30 Days, 1.2 Times penalty coefficient).
  2. Original specification Remaining fee = 4989 - 3991.2 = 997.8 USD.
  3. New specification purchase fee = 0 (capacity adjusted to 0).
  4. Refund = max(997.8 - 0, 0) = 997.8 USD.

After being set to zero, no capacity fees are incurred anymore, and the dedicated Model Code is retained. Setting to zero is a reduction, and the used portion is settled by the penalty coefficient (same as the Scale In formula, New Capacity = 0, Refund = Original specification Remaining fee).

Pay-as-you-go Instances

  • The Pay-as-you-go purchase granularity is by hour, and the billing precision is by minute.
  • The used amount is settled by Usage Minutes ÷ 60 and does not generate a refund.
  • After Scale In, the new capacity is billed by the new capacity from the effective time; deletion or setting to zero stops billing.

Used Amount = Pay-as-you-go Input Unit price × Input Capacity(kTPM) ÷ 10 × Used Minutes ÷ 60 + Pay-as-you-go Output Unit price × Output Capacity(kTPM) × Used Minutes ÷ 60

Used 30 Minutes·Scaled In to Half

Take Qwen3.8-Max (China (Beijing)) High-Speed mode as an example (assume Pay-as-you-go Input Unit price $3.96/10k TPM·hour, Output Unit price $1.188/1k TPM·hour), original 100 Input + 50 Output kTPM, scaled in to 50 Input + 25 Output kTPM, used 30 minutes:

  1. Used Amount = 3.96 × (100 ÷ 10) × 30 ÷ 60 + 1.188 × 50 × 30 ÷ 60 = 3.96 × 10 × 0.5 + 1.188 × 50 × 0.5 = 19.8 + 29.7 = 49.5 USD (actual consumption, non-refundable).
  2. After Scale In, the new capacity 50 Input + 25 Output kTPM is billed by the new capacity from the effective time.

Used 2 Hours·Scaled In to Half

Take Qwen3.8-Max (China (Beijing)) High-Speed mode as an example (assume Pay-as-you-go Input Unit price $3.96/10k TPM·hour, Output Unit price $1.188/1k TPM·hour), original 100 Input + 50 Output kTPM, scaled in to 50 Input + 25 Output kTPM, used 2 hours:

  1. Used Amount = 3.96 × (100 ÷ 10) × 120 ÷ 60 + 1.188 × 50 × 120 ÷ 60 = 3.96 × 10 × 2 + 1.188 × 50 × 2 = 79.2 + 118.8 = 198 USD (actual consumption, non-refundable).
  2. After Scale In, the new capacity is billed by the new capacity from the effective time.

Unsubscribe

Subscription Instances

  • The Subscription purchase granularity is by day, and the billing precision is by minute (the settled amount is rounded up by hour).
  • Unsubscribe is a special case of Scale In (New specification = 0), and the settled amount is settled by rounded up used hours × penalty coefficient.
  • ≤30 Days by 1.2 Times, >30 Days by 1.0 Times; Unsubscribe is irreversible, and the dedicated Model Code becomes invalid immediately.

Refund = max(Original specification Remaining fee, 0) (special case of Scale In where New specification purchase fee = 0)

The Unsubscribe process is handled in the Expenses Center.

Purchased 30 Days·Used 20 Days (≤30 Days, 1.2 Times)

Take Qwen3.8-Max (China (Beijing)) Standard mode as an example (assume daily Unit price $16.63/10k TPM·day, Output Unit price $4.99/1k TPM·day), original 100 kTPM purchased 30 Days, Unsubscribe:

  1. Settled Amount = (16.63 ÷ 24) × 480 × 1.2 × 100 ÷ 10 = 3991.2 USD.
  2. Original specification Remaining fee = 4989 - 3991.2 = 997.8 USD.
  3. Refund = max(997.8, 0) = 997.8 USD.

Purchased 60 Days·Used 40 Days (>30 Days, 1.0 Times)

Take Qwen3.8-Max (China (Beijing)) Standard mode as an example (assume daily Unit price $16.63/10k TPM·day, Output Unit price $4.99/1k TPM·day), original 100 kTPM purchased 60 Days, Unsubscribe:

  1. Settled Amount = (16.63 ÷ 24) × 960 × 1.0 × 100 ÷ 10 = 6652 USD.
  2. Original specification Remaining fee = 9978 - 6652 = 3326 USD.
  3. Refund = max(3326, 0) = 3326 USD.

Pay-as-you-go Instances

  • The Pay-as-you-go purchase granularity is by hour, and the billing precision is by minute.
  • The used amount is settled by Usage Minutes ÷ 60 and does not generate a refund.
  • Deleting the instance stops billing.

Used Amount = Pay-as-you-go Input Unit price × Input Capacity(kTPM) ÷ 10 × Used Minutes ÷ 60 + Pay-as-you-go Output Unit price × Output Capacity(kTPM) × Used Minutes ÷ 60

Deleted After 30 Minutes·Stop Billing

Take Qwen3.8-Max (China (Beijing)) High-Speed mode as an example (assume Pay-as-you-go Input Unit price $3.96/10k TPM·hour, Output Unit price $1.188/1k TPM·hour), original 100 Input + 50 Output kTPM, deleted after 30 minutes of usage:

  1. Used Amount = 3.96 × (100 ÷ 10) × 30 ÷ 60 + 1.188 × 50 × 30 ÷ 60 = 49.5 USD (actual consumption).
  2. Billing stops after deletion, and no refund is generated.

Add-on Capacity Package Billing

Additional purchased capacity instances (Add-on Capacity Packages) reuse the same dedicated Model Code; the total capacity increases accordingly, and no API changes are required.

Billing Item

Description

Billing Cycle

The selectable Billing Cycle of an add-on package is constrained by the base reservation's billing method: when the base is Subscription, only Subscription add-on packages can be selected; when the base is Pay-as-you-go, either Subscription or Pay-as-you-go add-on packages can be selected.

Subscription (By Day): supported by both Standard mode and High-Speed mode, billed by natural day, with billing rules consistent with a new purchase.

Subscription (By 8-Hour Block): only Standard mode, with rules same as 8-Hour Block Reservation Billing. After adding, the capacity of the same dedicated Model Code increases within that 8-hour window, and the 8h package capacity is invalid outside the window.

Pay-as-you-go (By Hour): only High-Speed mode, billed by actual duration, no need to purchase a Duration.

Input Capacity

Starting at 200 kTPM, with a step size of 10 kTPM.

Output Capacity

Starting at 20 kTPM, with a step size of 1 kTPM.

Duration

Only required for Subscription/By Day: 1~30, 60, 90, 120, 365 Days.

Auto-Renew on Expiration

Only applicable to Subscription/By Day; auto-deducts and renews 1 day Before expiration, with up to 4 attempts between 8:20–20:00 on that day.

NoteA single Model Code can have at most one undeleted Pay-as-you-go (By Hour) capacity instance: if a Pay-as-you-go instance already exists, the Billing Cycle of the Add-on Capacity Package can only be Subscription/By Day.

Overflow Strategy and Billing

The overflow strategy determines how the exceeding portion is handled and billed after the capacity is exhausted:

Overflow Strategy

Exceeding Portion Behavior

Billing Method

Auto Overflow (Default)

Exceeding requests automatically degrade to the model's standard pay-as-you-go billing (Token billing), without service interruption.

The exceeding portion is billed according to the Pay-as-you-go for Model Calls standard and no longer enjoys the reserved pricing. You can view the number of degradations in the Overage Degradation Statistics on the details page.

Only Use Reserved Capacity

Exceeding requests return a 429 Error, and the business needs to retry or degrade on its own.

The exceeding portion incurs no additional Fee.

In the auto overflow scenario, when a single input exceeds the model's Max Input Token, that call also switches to pay-as-you-go billing.

Bills and Usage Query

  • Fee Bills: Subscription orders can be viewed in the Expenses Center, supporting viewing of order details and consumption records.
  • Usage Monitor: In the Bailian Console details page, the Overview tab shows usage trends, Utilization, and Overage Degradation Statistics; the Monitor tab shows the number of calls inside and outside the quota and the cache hit volume. For more information, see Monitoring.
  • Utilization Description: Utilization = Converted consumption ÷ Purchased quota. The long input tier coefficient has been included in the converted consumption.

FAQ

Q: Why is the Refund 0?

When the Original specification Remaining fee ≤ the New specification purchase fee, the Refund is 0. Settled Amount = (Daily Unit price/24) × Rounded up(used hours) × Penalty coefficient × Original Capacity(kTPM) ÷ 10; usage ≤30 Days is settled with a 1.2 Times penalty coefficient, and >30 Days with 1.0 Times (no penalty). The longer the usage, the higher the Settled Amount and the lower the Remaining fee, so the Refund may be 0.

Q: What is the difference between Standard mode and High-Speed mode?

Standard mode has the same TPS as the standard API; High-Speed mode has a 1.5~2 Times TPS improvement over the standard API. The billing rules of the two modes are similar, but High-Speed mode additionally supports Pay-as-you-go/By Hour, with different Billing Cycle options and pricing.

Q: How to estimate the Reserved Fee?

Use the formula Reserved Fee = Duration ( Days) × (Input Unit price × Input kTPM ÷ 10 + Output Unit price × Output kTPM); the unit price is taken from the corresponding mode column of the Pricing table. For more information, see the example in Billing Rules.

Q: Can I get a refund after purchase? How much?

You can scale in or unsubscribe. Refund = max(Original specification Remaining fee - New specification purchase fee, 0); the used portion is settled by the penalty coefficient. For more information, see Scale In.

Q: What happens when it expires?

After expiration, the service is immediately disabled and resources are released; the dedicated Model Code becomes invalid and cannot be recovered. It is recommended to enable Auto-Renew on Expiration in advance to avoid service interruption. For more information, see Expiration and Lifecycle.

Q: Can the dedicated Model Code still be used after being set to zero?

After being set to zero, no capacity fees are incurred anymore, and the dedicated Model Code is retained, but the capacity is 0 and cannot be called. You need to Scale Out again to restore calls.