Inference calls to built-in models incur model invocation fees billed by token usage, independent of workspace CU compute resource fees.
Billing Overview
-
Compute Resources (CU): The compute resources required for job execution, billed the same way as existing CU billing.
-
AI Service Invocation Fee: Fees incurred by calling built-in models, billed independently by token usage.
For billing of related cloud products, see Billable Items.
Billing Rules
-
Billing Unit: USD/1K tokens. Input tokens and output tokens are billed separately.
-
Billing Cycle: Settled once per minute (based on UTC+8 time), generating bills and deducting fees.
-
Billing Dimensions: Billed by the combination of
Region × Model × Input Volume Tier × Token Type. Per-minute request cost = Σ(token usage per type × unit price). -
Free Tier: The first 1 million tokens per primary account per region per calendar month (based on UTC+8 time) are free (input and output combined). Exceeding usage is billed at listed prices.
Token Type Description
Each inference request generates tokens metered by the following types:
-
Standard Input Token: Billed at the base input unit price.
-
Standard Output Token: Billed at the base output unit price.
-
Implicit Cache Hit Input Token: Billed at 20% of the base input price.
-
Explicit Cache Creation Input Token: Billed at 125% of the base input price.
-
Explicit Cache Hit Input Token: Billed at 10% of the base input price.
Per-minute request cost = Σ(token usage per type × unit price). The unit price of cache-type tokens = base input unit price × billing coefficient.
Pricing
Actual prices are subject to billing statements.
Base Pricing
The following table shows the base unit prices of standard input tokens and standard output tokens. Cache-type tokens are billed at the base prices multiplied by the corresponding coefficients. For details, see Cache Token Billing Coefficients.
-
Chinese mainland: such as China (Hangzhou) and China (Beijing).
-
Outside the Chinese mainland: includes all regions outside the Chinese mainland, such as China (Hong Kong), Singapore, and US.
For the region list, see Supported regions.
Qwen
Chinese mainland
|
Model |
Input volume tier (1K tokens) |
Input (USD/1K tokens) |
Output (USD/1K tokens) |
|
qwen3.8-max |
0~1,000 |
0.00198 |
0.005941 |
|
qwen3.7-max |
0~1,000 |
0.00198 |
0.005941 |
|
qwen3.7-plus |
0~256 |
0.00033 |
0.00132 |
|
256~1,000 |
0.00099 |
0.003961 |
|
|
qwen3.6-plus |
0~256 |
0.00033 |
0.00198 |
|
256~1,000 |
0.00132 |
0.007921 |
|
|
qwen3.5-plus |
0~128 |
0.000132 |
0.000792 |
|
128~256 |
0.00033 |
0.00198 |
|
|
256~1,000 |
0.00066 |
0.003961 |
|
|
qwen3.7-flash |
0~32 |
0.000033 |
0.000132 |
|
32~256 |
0.000099 |
0.000396 |
|
|
256~1,000 |
0.000198 |
0.000792 |
|
|
qwen3.6-flash |
0~256 |
0.000198 |
0.001188 |
|
256~1,000 |
0.000792 |
0.004753 |
|
|
qwen3.5-flash |
0~128 |
0.000033 |
0.00033 |
|
128~256 |
0.000132 |
0.00132 |
|
|
256~1,000 |
0.000198 |
0.00198 |
|
|
qwen3.5-ocr |
0~1,000 |
0.000083 |
0.00033 |
|
qwen-vl-ocr |
0~1,000 |
0.00005 |
0.000083 |
|
qwen-mt-plus |
0~1,000 |
0.000297 |
0.000891 |
|
qwen-mt-flash |
0~1,000 |
0.000116 |
0.000322 |
Outside the Chinese mainland
|
Model |
Input volume tier (1K tokens) |
Input (USD/1K tokens) |
Output (USD/1K tokens) |
|
qwen3.8-max |
0~1,000 |
0.002473 |
0.00742 |
|
qwen3.7-max |
0~1,000 |
0.003093 |
0.009276 |
|
qwen3.7-plus |
0~256 |
0.000495 |
0.001979 |
|
256~1,000 |
0.001484 |
0.005936 |
|
|
qwen3.6-plus |
0~256 |
0.000619 |
0.00371 |
|
256~1,000 |
0.002474 |
0.007421 |
|
|
qwen3.5-plus |
0~128 |
0.000485 |
0.002906 |
|
128~256 |
0.000485 |
0.002906 |
|
|
256~1,000 |
0.000606 |
0.003634 |
|
|
qwen3.7-flash |
0~32 |
0.000037 |
0.000161 |
|
32~256 |
0.000124 |
0.000495 |
|
|
256~1,000 |
0.000247 |
0.000989 |
|
|
qwen3.6-flash |
0~256 |
0.000309 |
0.001855 |
|
256~1,000 |
0.001236 |
0.004947 |
|
|
qwen3.5-flash |
0~128 |
0.00012 |
0.000485 |
|
128~256 |
0.00012 |
0.000485 |
|
|
256~1,000 |
0.000198 |
0.000485 |
|
|
qwen-vl-ocr |
0~1,000 |
0.000085 |
0.000194 |
|
qwen-mt-plus |
0~1,000 |
0.00298 |
0.008926 |
|
qwen-mt-flash |
0~1,000 |
0.000194 |
0.000593 |
DeepSeek
Chinese mainland
|
Model |
Input volume tier (1K tokens) |
Input (USD/1K tokens) |
Output (USD/1K tokens) |
|
deepseek-v4-pro |
0~1,000 |
0.00198 |
0.003961 |
|
deepseek-v4-flash-0731 |
0~1,000 |
0.000165 |
0.00033 |
Outside the Chinese mainland
|
Model |
Input volume tier (1K tokens) |
Input (USD/1K tokens) |
Output (USD/1K tokens) |
|
deepseek-v4-pro |
0~1,000 |
0.002968 |
0.005936 |
|
deepseek-v4-flash-0731 |
0~1,000 |
0.000247 |
0.000495 |
GLM
Chinese mainland
|
Model |
Input volume tier (1K tokens) |
Input (USD/1K tokens) |
Output (USD/1K tokens) |
|
glm-5.2 |
0~1,000 |
0.00132 |
0.004621 |
Outside the Chinese mainland
|
Model |
Input volume tier (1K tokens) |
Input (USD/1K tokens) |
Output (USD/1K tokens) |
|
glm-5.2 |
0~1,000 |
0.001731 |
0.005442 |
Embedding
Chinese mainland
|
Model |
Input type |
USD/1K tokens |
|
qwen3.7-text-embedding |
Text |
0.000083 |
|
text-embedding-v4 |
Text |
0.000083 |
|
qwen3-vl-embedding |
Text |
0.000116 |
|
Image |
0.000297 |
Outside the Chinese mainland
|
Model |
Input type |
USD/1K tokens |
|
text-embedding-v4 |
Text |
0.000085 |
Cache Token Billing Coefficients
Cost of cache tokens = base input unit price of the corresponding model × billing coefficient × token usage. For more information, see Context Cache.
|
Token type |
Billing coefficient |
Description |
|
Implicit cache hit |
×0.2 |
Billed at 20% of the base input price |
|
Explicit cache creation |
×1.25 |
Billed at 125% of the base input price |
|
Explicit cache hit |
×0.1 |
Billed at 10% of the base input price |
Billing Example
Scenario: Using qwen3.6-plus in the China North 2 (Beijing) region (input volume tier 0~256K tokens), with cumulative consumption of 5,000K input tokens and 1,200K output tokens.
Fee Calculation:
-
Input fee: 5,000 × 0.00033 = 1.65 USD
-
Output fee: 1,200 × 0.00198 = 2.376 USD
-
Total: 4.026 USD
Bill Details
Bills are itemized by region, workspace ID, Flink usage type, model, token type, and input volume tier.
Download Bills
-
Log on to the Bill Details page and select the billing period.
-
Set Product to Realtime Compute for Apache Flink, and set Product Name to Realtime Compute for Apache Flink_AI_PAYG_International, and click Search.
-
Click the export icon in the upper-right corner of the bill list to download the bill.
-
Open the bill file and locate the Instance ID (billing granularity) column. This field is separated by semicolons (;) in the following format:
Region;Workspace ID;Project Name;Flink Usage;Model;Token Type;Input Volume TierSegment
Description
Region
Region of the workspace
Workspace ID
Unique identifier of the Flink workspace
Project Name
Currently empty; reserved field
Flink Usage
Token usage scenario; valid values:
ai_functionorflink_agentsModel
Name of the invoked LLM
Token Type
Input Token or Output Token
Input Volume Tier
Upper bound marker of the input token range, such as 128, 256, 1000, etc.
Overdue Payment
-
After an overdue payment occurs, the service is suspended and all jobs that use built-in models cannot invoke models.
-
The service is automatically restored after you top up your account.
-
Model invocation requests during the overdue period are not processed, and the corresponding jobs may fail.
Stop Billing
To stop AI service billing, click Disable Service on the AI Service page in the console. After you disable the service:
-
All jobs will be unable to call the built-in model service.
-
No further AI service invocation fees will be incurred.
-
Workspace CU compute resource fees are not affected.
Before disabling the AI service, ensure that no jobs depend on built-in models. Otherwise, jobs may fail at runtime.