PAI Token Service provides access to open-source and proprietary large models and integrates seamlessly with PAI model evaluation, model distillation, DSW, Feature Store, and multimodal data processing. You are billed based on token usage.
Activate PAI Token Service and review its pay-as-you-go billing rules, per-model pricing, and context-cache discounts.
Activate PAI Token Service
-
Activation prerequisites:
Alibaba Cloud account: You can activate the service directly.
-
RAM user: You can use the service in the following ways.
Method 1: Contact an Alibaba Cloud account, or an account that has the
AliyunPAIFullAccesspermission policy or themodelservice:*permission, to activate the service. After activation, the RAM user can use the service.Method 2: After the RAM user is granted the
AliyunPAIFullAccesspermission policy or themodelservice:*permission, the RAM user can activate the service.
-
How to activate:
Log in to the PAI console, choose QuickStart > Token Service in the left-side navigation pane, and then click Activate Now.
Billing rules
PAI Token Service uses pay-as-you-go (postpaid) billing by default. You are charged for the input and output tokens you use. Fees are based on actual usage and billed after use.
Billing rules: Usage fee = input token count × input price + output token count × output price
Context cache discount: For requests that hit the context cache, the input tokens are discounted. For more information, see Appendix: Context cache discounts
Text generation - Qwen
Qwen Max
Billing rules: This model is billed by input and output tokens.
China North 2 (Beijing)
Model ID | Mode | Input tokens per request | Input price (per million tokens) | Output price (per million tokens) Chain-of-thought + answer |
qwen3.8-max | thinking / non-thinking mode | 0<token≤1M | $1.98 | $5.9412 |
qwen3.7-max | thinking / non-thinking mode | 0<token≤1M | $1.98 | $5.9412 |
Singapore
Model ID | Service deployment scope | Mode | Input tokens per request | Input price (per million tokens) | Output price (per million tokens) Chain-of-thought + answer |
qwen3.8-max | International | thinking / non-thinking mode | 0<token≤1M | $2.4 | $7.2 |
qwen3.7-max | International | thinking / non-thinking mode | 0<token≤1M | $3 | $9 |
Germany (Frankfurt)
Model ID | Service deployment scope | Mode | Input tokens per request | Input price (per million tokens) | Output price (per million tokens) Chain-of-thought + answer |
qwen3.8-max | Global | thinking / non-thinking mode | 0<token≤1M | $1.98 | $5.9412 |
qwen3.7-max | Global | thinking / non-thinking mode | 0<token≤1M | $1.98 | $5.9412 |
Japan (Tokyo)
Model ID | Service deployment scope | Mode | Input tokens per request | Input price (per million tokens) | Output price (per million tokens) Chain-of-thought + answer |
qwen3.8-max | Global | thinking / non-thinking mode | 0<token≤1M | $1.98 | $5.9412 |
qwen3.7-max | Global | thinking / non-thinking mode | 0<token≤1M | $1.98 | $5.9412 |
Qwen Plus
Billing rules: This model is billed by input and output tokens.
China North 2 (Beijing)
Model ID | Input token range per request | Input price (per million tokens) | Output price (per million tokens) | |
Non-thinking mode | Thinking mode (chain-of-thought + answer) | |||
qwen3.7-plus | 0<token≤256K | $0.3444 | $1.3764 | $1.3764 |
256K<token≤1M | $1.032 | $4.128 | $4.128 | |
qwen3.6-plus | 0<token≤256K | $0.3312 | $1.9812 | $1.9812 |
256K<token≤1M | $1.3212 | $7.9224 | $7.9224 | |
qwen3.5-plus | 0<token≤128K | $0.138 | $0.8256 | $0.8256 |
128K<token≤256K | $0.3444 | $2.064 | $2.064 | |
256K<token≤1M | $0.6876 | $4.128 | $4.128 | |
Singapore
Model ID | Service deployment scope | Input token range per request | Input price (per million tokens) | Output price (per million tokens) | |
Non-thinking mode | Thinking mode (chain-of-thought + answer) | ||||
qwen3.7-plus | International | 0<token≤256K | $0.48 | $1.92 | $1.92 |
256K<token≤1M | $1.44 | $5.76 | $5.76 | ||
qwen3.6-plus | International | 0<token≤256K | $0.6 | $3.6 | $3.6 |
256K<token≤1M | $2.4 | $7.2 | $7.2 | ||
qwen3.5-plus | International | 0<token≤256K | $0.48 | $2.88 | $2.88 |
256K<token≤1M | $0.6 | $3.6 | $3.6 | ||
Germany (Frankfurt)
Model ID | Service deployment scope | Input token range per request | Input price (per million tokens) | Output price (per million tokens) | |
Non-thinking mode | Thinking mode (chain-of-thought + answer) | ||||
qwen3.7-plus | Global | 0<token≤256K | $0.3444 | $1.3764 | $1.3764 |
256K<token≤1M | $1.032 | $4.128 | $4.128 | ||
qwen3.6-plus | Global | 0<token≤256K | $0.3312 | $1.9812 | $1.9812 |
256K<token≤1M | $1.3212 | $7.9224 | $7.9224 | ||
qwen3.5-plus | Global | 0<token≤128K | $0.138 | $0.8256 | $0.8256 |
128K<token≤256K | $0.3444 | $2.064 | $2.064 | ||
256K<token≤1M | $0.6876 | $4.128 | $4.128 | ||
Japan (Tokyo)
Model ID | Service deployment scope | Input token range per request | Input price (per million tokens) | Output price (per million tokens) | |
Non-thinking mode | Thinking mode (chain-of-thought + answer) | ||||
qwen3.7-plus | Global | 0<token≤256K | $0.3444 | $1.3764 | $1.3764 |
256K<token≤1M | $1.032 | $4.128 | $4.128 | ||
qwen3.6-plus | Global | 0<token≤256K | $0.3312 | $1.9812 | $1.9812 |
256K<token≤1M | $1.3212 | $7.9224 | $7.9224 | ||
Qwen Flash
Billing rules: This model is billed by input and output tokens.
China North 2 (Beijing)
Model ID | Mode | Input tokens per request | Input price (per million tokens) | Output price (per million tokens) Chain-of-thought + answer |
qwen3.7-flash | thinking / non-thinking mode | 0<token≤32K | $0.0336 | $0.132 |
32K<token≤256K | $0.0996 | $0.396 | ||
256K<token≤1M | $0.198 | $0.792 | ||
qwen3.6-flash | thinking / non-thinking mode | 0<token≤256K | $0.198 | $1.188 |
256K<token≤1M | $0.792 | $4.7532 | ||
qwen3.5-flash | thinking / non-thinking mode | 0<token≤128K | $0.0348 | $0.3444 |
128K<token≤256K | $0.138 | $1.3764 | ||
256K<token≤1M | $0.2064 | $2.064 |
Singapore
Model ID | Service deployment scope | Mode | Input tokens per request | Input price (per million tokens) | Output price (per million tokens) Chain-of-thought + answer |
qwen3.7-flash | International | thinking / non-thinking mode | 0<token≤32K | $0.036 | $0.156 |
32K<token≤256K | $0.12 | $0.48 | |||
256K<token≤1M | $0.24 | $0.96 | |||
qwen3.6-flash | International | thinking / non-thinking mode | 0<token≤256K | $0.3 | $1.8 |
256K<token≤1M | $1.2 | $4.8 | |||
qwen3.5-flash | International | thinking / non-thinking mode | 0<token≤1M | $0.12 | $0.48 |
Germany (Frankfurt)
Model ID | Service deployment scope | Mode | Input tokens per request | Input price (per million tokens) | Output price (per million tokens) Chain-of-thought + answer |
qwen3.7-flash | Global | thinking / non-thinking mode | 0<token≤32K | $0.0336 | $0.132 |
32K<token≤256K | $0.0996 | $0.396 | |||
256K<token≤1M | $0.198 | $0.792 | |||
qwen3.6-flash | Global | thinking / non-thinking mode | 0<token≤256K | $0.198 | $1.188 |
256K<token≤1M | $0.792 | $4.7532 | |||
qwen3.5-flash | Global | thinking / non-thinking mode | 0<token≤128K | $0.0348 | $0.3444 |
128K<token≤256K | $0.138 | $1.3764 | |||
256K<token≤1M | $0.2064 | $2.064 |
Japan (Tokyo)
Model ID | Service deployment scope | Mode | Input tokens per request | Input price (per million tokens) | Output price (per million tokens) Chain-of-thought + answer |
qwen3.7-flash | Global | thinking / non-thinking mode | 0<token≤32K | $0.0336 | $0.132 |
32K<token≤256K | $0.0996 | $0.396 | |||
256K<token≤1M | $0.198 | $0.792 | |||
qwen3.6-flash | Global | thinking / non-thinking mode | 0<token≤256K | $0.198 | $1.188 |
256K<token≤1M | $0.792 | $4.7532 |
Qwen Omni
Billing rules: This model is billed by input and output tokens.
China North 2 (Beijing)
Model ID | Input price (per million tokens) | Output price (per million tokens) | ||
Text/Image/Video | Audio | Text Multimodal input | Text + audio Audio-only billing | |
qwen3.5-omni-plus | $1.152 | $8.748 | $6.6 | $35.148 |
qwen3.5-omni-flash | $0.36 | $2.976 | $2.196 | $11.88 |
Singapore
Model ID | Service deployment scope | Input price (per million tokens) | Output price (per million tokens) | ||
Text/Image/Video | Audio | Text Multimodal input | Text + audio Audio-only billing | ||
qwen3.5-omni-plus | International | $1.68 | $13.2 | $9.96 | $52.8 |
qwen3.5-omni-flash | International | $0.48 | $3.6 | $2.64 | $14.28 |
Text generation - Third-party models
DeepSeek
Billing rules: This model is billed by input and output tokens.
China North 2 (Beijing)
Model ID | Input price (per million tokens) | Output price (per million tokens) Chain-of-thought + answer |
deepseek-v4-pro | $1.98 | $3.9612 |
deepseek-v4-flash-0731 | $0.1656 | $0.33 |
deepseek-v4-flash | $0.1656 | $0.33 |
Singapore
Model ID | Service deployment scope | Input price (per million tokens) | Output price (per million tokens) Chain-of-thought + answer |
deepseek-v4-pro | International | $2.88 | $5.76 |
deepseek-v4-flash-0731 | International | $0.24 | $0.48 |
deepseek-v4-flash | International | $0.24 | $0.48 |
Germany (Frankfurt)
Model ID | Service deployment scope | Input price (per million tokens) | Output price (per million tokens) Chain-of-thought + answer |
deepseek-v4-pro | Global | $1.98 | $3.9612 |
deepseek-v4-flash-0731 | Global | $0.1656 | $0.33 |
deepseek-v4-flash | Global | $0.1656 | $0.33 |
Japan (Tokyo)
Model ID | Service deployment scope | Input price (per million tokens) | Output price (per million tokens) Chain-of-thought + answer |
deepseek-v4-pro | Global | $1.98 | $3.9612 |
deepseek-v4-flash-0731 | Global | $0.1656 | $0.33 |
deepseek-v4-flash | Global | $0.1656 | $0.33 |
Kimi
Billing rules: This model is billed by input and output tokens.
China North 2 (Beijing)
|
Model ID |
Mode |
Input price (per million tokens) |
Output price (per million tokens) |
|
kimi-k2.7-code |
Reasoning-only mode |
$1.0728 |
$4.4556 |
|
kimi-k2.6 |
thinking / non-thinking mode |
$1.07268 |
$4.45572 |
Singapore
|
Model ID |
Service deployment scope |
Mode |
Input price (per million tokens) |
Output price (per million tokens) |
|
kimi-k2.7-code |
International |
Reasoning-only mode |
$1.14 |
$4.8 |
Germany (Frankfurt)
|
Model ID |
Service deployment scope |
Mode |
Input price (per million tokens) |
Output price (per million tokens) |
|
kimi-k2.7-code |
Global |
Reasoning-only mode |
$1.0728 |
$4.4556 |
GLM
Billing rules: This model is billed by input and output tokens.
China North 2 (Beijing)
Model ID | Mode | Input tokens per request | Input price (per million tokens) | Output price (per million tokens) Chain-of-thought and answer |
glm-5.2 | thinking / non-thinking mode | No tiered pricing | $1.32 | $4.6212 |
glm-5.1 | thinking / non-thinking mode | 0<token≤32K | $0.99 | $3.9612 |
32K<token≤200K | $1.32 | $4.6212 | ||
glm-5 | thinking / non-thinking mode | 0<token≤32K | $0.6876 | $3.096 |
32K<token≤198K | $1.032 | $3.7848 |
Singapore
Model ID | Service deployment scope | Mode | Input tokens per request | Input price (per million tokens) | Output price (per million tokens) Chain-of-thought and answer |
glm-5.2 | International | thinking / non-thinking mode | No tiered pricing | $1.68 | $5.28 |
glm-5.1 | International | thinking / non-thinking mode | 0<token≤200K | $1.68 | $5.28 |
Germany (Frankfurt)
Model ID | Service deployment scope | Mode | Input tokens per request | Input price (per million tokens) | Output price (per million tokens) Chain-of-thought and answer |
glm-5.2 | Global | thinking / non-thinking mode | No tiered pricing | $1.32 | $4.6212 |
glm-5.1 | Global | thinking / non-thinking mode | 0<token≤32K | $0.99 | $3.9612 |
32K<token≤200K | $1.32 | $4.6212 |
Japan (Tokyo)
Model ID | Service deployment scope | Mode | Input tokens per request | Input price (per million tokens) | Output price (per million tokens) Chain-of-thought and answer |
glm-5.1 | Global | thinking / non-thinking mode | 0<token≤32K | $0.99 | $3.9612 |
32K<token≤200K | $1.32 | $4.6212 |
Image generation
Billing rules: This model is billed by the number of input images and generated output images. For models whose input image price is not specified, input is not billed; only output is billed.
Billing formula: Cost = Input image unit price × Number of input images + Output image unit price × Number of output images
Qwen image generation and editing
Billed based on the number of input and output images
China North 2 (Beijing)
Model ID | Output image resolution | Input price | Output price |
qwen-image-3.0-pro | 1k | $0.0033/image | $0.041256/image |
2k | $0.0825132/image | ||
qwen-image-3.0 | 1k | $0.0033/image | $0.0297048/image |
2k | $0.0297048/image |
Singapore
Model ID | Service deployment scope | Output image resolution | Input price | Output price |
qwen-image-3.0-pro | International | 1k | $0.0036/image | $0.048/image |
2k | $0.09/image | |||
qwen-image-3.0 | International | 1k | $0.0036/image | $0.036/image |
2k | $0.036/image |
Germany (Frankfurt)
Model ID | Service deployment scope | Output image resolution | Input price | Output price |
qwen-image-3.0-pro | Global | 1k | $0.0033/image | $0.041256/image |
2k | $0.0825132/image | |||
qwen-image-3.0 | Global | 1k | $0.0033/image | $0.0297048/image |
2k | $0.0297048/image |
Japan (Tokyo)
Model ID | Service deployment scope | Output image resolution | Input price | Output price |
qwen-image-3.0-pro | Global | 1k | $0.0033/image | $0.041256/image |
2k | $0.0825132/image | |||
qwen-image-3.0 | Global | 1k | $0.0033/image | $0.0297048/image |
2k | $0.0297048/image |
Qwen text-to-image
Output-only billing
China North 2 (Beijing)
|
Model ID |
Output price |
|
qwen-image-2.0-pro |
$0.0860112/image |
|
qwen-image-2.0 |
$0.0344052/image |
Singapore
|
Model ID |
Service deployment scope |
Output price |
|
qwen-image-2.0-pro |
International |
$0.09/image |
|
qwen-image-2.0 |
International |
$0.042/image |
Video generation
Billing rules: This model is billed for output only. You are charged by the second of generated video.
Billing formula: Cost = Video unit price × Output video duration (unit: seconds)
HappyHorse text-to-video
Output-only billing
China North 2 (Beijing)
Model ID | Output video resolution | Output price |
happyhorse-1.1-t2v | 480P | $0.074262/second |
720P | $0.1485228/second | |
1080P | $0.1980312/second |
Singapore
Model ID | Service deployment scope | Output video resolution | Output price |
happyhorse-1.1-t2v | International | 480P | $0.084/second |
720P | $0.168/second | ||
1080P | $0.216/second |
Germany (Frankfurt)
Model ID | Service deployment scope | Output video resolution | Output price |
happyhorse-1.1-t2v | Global | 480P | $0.074262/second |
720P | $0.1485228/second | ||
1080P | $0.1980312/second |
HappyHorse image-to-video (first-frame)
Output-only billing
China North 2 (Beijing)
Model ID | Output video resolution | Output price |
happyhorse-1.1-i2v | 480P | $0.074262/second |
720P | $0.1485228/second | |
1080P | $0.1980312/second |
Singapore
Model ID | Service deployment scope | Output video resolution | Output price |
happyhorse-1.1-i2v | International | 480P | $0.084/second |
720P | $0.168/second | ||
1080P | $0.216/second |
Germany (Frankfurt)
Model ID | Service deployment scope | Output video resolution | Output price |
happyhorse-1.1-i2v | Global | 480P | $0.074262/second |
720P | $0.1485228/second | ||
1080P | $0.1980312/second |
Wanxiang 3.0 video generation
Billing rules: This model is billed by input and output video duration.
Billing formula: Billed duration = input video duration + output video duration.
China North 2 (Beijing)
Model ID | Output video resolution | Input and output price |
wan3.0-video | 480P | $0.0495072/second |
720P | $0.0990156/second | |
1080P | $0.19803/second |
Singapore
Model ID | Service deployment scope | Output video resolution | Input and output price |
wan3.0-video | International | 480P | $0.06/second |
720P | $0.12/second | ||
1080P | $0.24/second |
Wanxiang text-to-video
Output-only billing
China North 2 (Beijing)
Model ID | Output video resolution | Output price |
wan2.7-t2v | 720P | $0.1032144/second |
1080P | $0.1720236/second |
Singapore
Model ID | Service deployment scope | Output video resolution | Output price |
wan2.7-t2v | International | 720P | $0.12/second |
1080P | $0.18/second |
Wanxiang image-to-video
Output-only billing
China North 2 (Beijing)
Model ID | Output video type | Output video resolution | Output price |
wan2.7-i2v | Video with audio | 720P | $0.1032144/second |
1080P | $0.1720236/second |
Singapore
Model ID | Service deployment scope | Output video type | Output video resolution | Output price |
wan2.7-i2v | International | Video with audio | 720P | $0.12/second |
1080P | $0.18/second |
Text embedding
Billing rules: This model is billed by input tokens only; output is not billed.
China North 2 (Beijing)
|
Model ID |
Input price (per million tokens) |
|
text-embedding-v4 |
$0.0864 |
Singapore
|
Model ID |
Input price (per million tokens) |
|
text-embedding-v4 |
$0.084 |
Multimodal embedding
Billing rules: This model is billed by input tokens only; output is not billed.
China North 2 (Beijing)
|
Model ID |
Input price (per million tokens) |
|
qwen3-vl-embedding |
Image/Video:$0.3096 Text:$0.12 |
Singapore
|
Model ID |
Service deployment scope |
Input price (per million tokens) |
|
tongyi-embedding-vision-plus |
International |
$0.108 |
|
tongyi-embedding-vision-flash |
International |
Image/Video:$0.036 Text:$0.108 |
Reranking models
Text reranking model
Billing rules: This model is billed by input tokens only; output is not billed.
China North 2 (Beijing)
|
Model ID |
Input price (per million tokens) |
|
qwen3-rerank |
Text input:$0.0828 |
Singapore
|
Model ID |
Service deployment scope |
Input price (per million tokens) |
|
qwen3-rerank |
International |
$0.12 |
Appendix: Context cache discounts
|
Item |
Explicit caching |
Implicit caching |
|
token billing for cache creation |
125% of the input token unit price |
100% of the input token unit price |
|
token billing for cache hits |
10% of the input token unit price |
20% of the input token unit price |
FAQ
Q: Does PAI Token Service support savings plan? Can savings plan from Bailian be used for PAI Token Service?
PAI Token Service does not currently support savings plan, and savings plan from Bailian cannot be used for PAI Token Service.