All Products
Search
Document Center

Alibaba Cloud Model Studio:Rate limiting

Last Updated:Sep 24, 2026

Alibaba Cloud Model Studio applies rate limiting to model calls at the Alibaba Cloud account level, aggregating usage across all RAM users, workspaces, and API keys under the account. Requests are rejected when the limit is exceeded and typically recover automatically within one minute.

Rate Limiting Rules

  • Account-level Rate Limiting: Rate limiting is calculated by primary account dimension. The call volume of all RAM sub-accounts, workspaces, and API Keys under the account is calculated cumulatively.
  • Model-independent Rate Limiting: Different models have independent rate limit quotas. For more information, see the table below.

Request Body Size Limits

When sending a request via API, the gateway has the following limits on the request body (Request Body) size:

Request TypeLimit
Requests without Base64Request body Max 16 MiB
Requests with Base64 (Anthropic protocol)Single Base64 content Max 32 MiB
Requests with Base64 (other protocols)Single Base64 content Max 20 MiB
Requests with Base64 (entire request body)Max 64 MiB

NoteAfter server-side decoding of requests containing Base64, the entire request body size must still not exceed 16 MiB. When the limit is exceeded, the request will be rejected by the gateway and Return HTTP 413 Error.

FAQ

Why is rate limiting triggered?

Based on the error message, determine which type of rate limit was triggered:

  • Requests rate limit exceeded Or You exceeded your current requests list: Requests per minute (RPM) rate limit triggered.
  • Allocated quota exceeded Or You exceeded your current quota: Tokens per minute (TPM) rate limit triggered.
  • Request rate increased too quickly: Request frequency spiked in a short period, triggering system stability protection—this is triggered even if the total number of calls does not reach the RPM or TPM upper limit.
  • For other errors, see Error Code to confirm the cause.

In addition to RPM and TPM, the rate limit policy may be executed based on per-second RPS (RPM/60) and TPS (TPM/60). Even if the total number of calls per minute does not exceed the limit, a burst of requests in a short period may also trigger rate limiting.

How to view model call volume?

After a model call occursminute-levelyou can view it in the monitoring chart. Log in to the Monitoring page (SingaporeOrBeijing), in the left navigation pane, selectO&M Management > MonitoringGo to the Monitoring overview page, set query conditions (for example, select a time range, workspace, etc.), then in theModelsarea, find the target Model and clickView Detailsto view the invocation statistics of the Model. For details, seeMonitoringDocs.

Monitoring data is updated at the Minute(s)-level and is for reference only; it is not used as a billing basis.

How long does it take to recover after Rate Limiting is triggered?

It usually recovers within one Minute(s). If other Errors occur, seeError Code进行处理。

模型响应速度与付费有关吗?

Model response speed is unrelated to whether you pay. Free quota and paid calls use the same Model service infrastructure, and the generation speed of the same Model is identical. Free quota only controls whether a request is accepted—when the quota is exhausted, a 403 Error is returned (ActivateFree Quota Onlyafter), and it does not reduce the generation speed of already accepted requests.

什么因素影响模型响应速度?

  • Model Type:轻量模型(如 qwen-flash)比大型模型(如 qwen-max)生成更快。
  • Output length:The more Output Tokens, the longer the total time.
  • Server load:Slight fluctuations may occur during peak periods.

Does rate limiting reduce the generation speed of accepted requests?

No. Rate limiting (RPM/TPM) only causes requests exceeding the limit to be rejected (Return 429 Error). When the free quota is exhausted (Activate Free Quota Only) Return 403 Error, which also does not affect the speed of accepted requests. 429 is rate limiting, 403 is quota exhaustion, both are rejection mechanisms.

How to avoid rate limiting?

  1. Choose models with higher rate limits:The stable version Or latest version has more lenient rate limiting than dated snapshot versions.

  2. Optimize invocation strategy
    • Reduce invocation frequency: Received Requests rate limit exceeded Or You exceeded your current requests list, reduce the API invocation frequency.
    • Reduce Token consumption: Received Allocated quota exceeded Or You exceeded your current quota, shorten the input or limit the Output length.
    • Smooth request rate: Receive Request rate increased too quickly, use uniform-rate scheduling, exponential backoff, or request queues to evenly distribute requests and avoid instantaneous peaks.
  3. Add backup model

    After triggering rate limit, switch to a backup model to continue generation, which can reduce the failure probability and improve throughput. The following code calls qwen-plus-2025-07-28 after triggering rate limit, automatically switches to qwen-plus-2025-07-14 to retry. You can also choose models from different series (such as qwen-flash) as backup models to further reduce the risk of both primary and backup models being rate-limited at the same time.

    示例代码

    import os
    import asyncio
    from openai import AsyncOpenAI, APIStatusError
    
    # 配置
    API_KEY = os.getenv("DASHSCOPE_API_KEY")
    # 主用模型
    MODEL = "qwen-plus-2025-07-28"
    # 备选模型
    BACKUP_MODEL = "qwen-plus-2025-07-14"
    # 测试问题
    QUESTION = "你是谁?"
    # 并发设置
    NUM_REQUESTS = 10
    
    client = AsyncOpenAI(
        api_key=API_KEY,
        # 调用时请将WorkspaceId替换为真实的业务空间ID
        base_url="https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1"
    )
    
    async def send_request(model):
        """发送单个请求"""
        try:
            await client.chat.completions.create(
                model=model,
                messages=[{"role": "user", "content": QUESTION}]
            )
            return True
        except APIStatusError as e:
            if e.status_code == 429:
                print(f"[限流触发] 模型 {model}")
                return False
            raise
        except Exception as e:
            print(f"[请求失败] 模型 {model},错误:{e}")
            return False
    
    async def task(i):
        # 尝试主模型
        if await send_request(MODEL):
            return True
        # 限流时尝试备用模型
        return await send_request(BACKUP_MODEL)
    
    async def main():
        results = await asyncio.gather(*(task(i) for i in range(NUM_REQUESTS)))
        print(f"成功请求: {sum(results)}, 失败请求: {len(results) - sum(results)}")
    
    if __name__ == "__main__":
        asyncio.run(main())
    
  4. Split tasks: Long conversations or large documents quickly consume a large number of Tokens. Split large batch tasks into smaller batches and submit them in time slots.

  5. Batches: When real-time response is not needed, useBatches(Batch API)。Batch requests are not subject to real-time rate limiting constraints, but you need to consider queuing and processing time.

How to control Token Usage Or cost expenditure?

Rate limiting only constrains the call rate per unit time and does not limit cumulative Usage. If you need to control Token Usage Or cost expenditure, you can manage it through the following methods:

  • Set consumption limits and fee alerts: InBills and feescard, setfee alerts, enable monthly consumption limits and Configure threshold notifications. You will be alerted when the threshold is reached to avoid overspending. For more information, see Billing Query and Cost Management。
  • Enable Free Quota Only: For models that support the free quota, you can enableFree Quota Only, after the free quota is exhausted, calling automatically stops to avoid additional costs. For more information, see New User Free Quota。
  • Monitor Model Usage: Regularly check the Token usage of each model, and detect abnormal growth in a timely manner. See above How to view model usage?。

Will you still be rate-limited after topping up?

Top Up does not change the model's default RPM and TPM rate limit thresholds. Rate limits are configured at the Alibaba Cloud primary account level and are independent of billing; Top Up (Pay-as-you-go) only ensures the account is not suspended due to overdue payment and does not increase the rate limit.

If you need higher rate limit quotas, contact your business manager to apply.

Text Generation-Qwen

Qwen Language Model

新加坡

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions; the service may also limit by RPS (RPM/60) and TPS (TPM/60)
Calls per minute (RPM)Tokens consumed per minute (TPM)
Including Input and Output Token

qwen3.8-max

International

Dynamic Rate Limiting

qwen3.8-max-0902

International

Dynamic Rate Limiting

qwen3.8-flash

International

Dynamic Rate Limiting

qwen3.7-max

International

600

1,000,000

qwen3.7-max-2026-06-08

International

60

1,000,000

qwen3.7-max-2026-05-20

International

60

1,000,000

qwen3.7-max-preview

International

600

1,000,000

qwen3.7-max-2026-05-17

International

600

1,000,000

qwen3.6-max-preview

International

600

1,000,000

qwen3-max

International

600

1,000,000

qwen3-max-2026-01-23

International

600

1,000,000

qwen3-max-2025-09-23

International

60

100,000

qwen3-max-preview

International

600

1,000,000

qwen-max

UseBatch APIWhen calling the service, it is not subject to rate limiting.

International

600

1,000,000

qwen3.7-plus

International

15,000

5,000,000

qwen3.7-plus-2026-05-26

International

60

1,000,000

qwen3.6-plus

International

15,000

5,000,000

qwen3.6-plus-2026-04-02

International

60

1,000,000

qwen3.7-flash

International

15,000

5,000,000

qwen3.7-flash-2026-07-15

International

15,000

5,000,000

qwen3.6-flash

International

15,000

5,000,000

qwen3.6-flash-2026-04-16

International

60

1,000,000

qwen3.5-plus

International

15,000

5,000,000

qwen3.5-plus-2026-04-20

International

600

1,000,000

qwen3.5-plus-2026-02-15

International

60

1,000,000

qwen-plus

UseBatch APIWhen calling the service, it is not subject to rate limiting.

International

600

1,000,000

qwen-plus-latest

International

600

1,000,000

qwen-plus-2025-12-01

International

120

1,000,000

qwen-plus-2025-09-11

International

120

1,000,000

qwen-plus-2025-07-28

International

60

100,000

qwen-plus-2025-07-14

(qwen-plus-0714)

International

60

100,000

qwen-plus-2025-04-28

(qwen-plus-0428)

International

60

1,000,000

qwen-plus-2025-01-25

(qwen-plus-0125)

International

60

100,000

qwen3.5-flash

International

15,000

5,000,000

qwen3.5-flash-2026-02-23

International

60

1,000,000

qwen-flash

UseBatch APIWhen calling the service, it is not subject to rate limiting.

International

600

5,000,000

qwen-flash-2025-07-28

International

600

5,000,000

qwq-plus

International

60

100,000

qwen-turbo

UseBatch APIWhen calling the service, it is not subject to rate limiting.

International

600

5,000,000

美国(弗吉尼亚)

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may impose limits based on RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output tokens

qwen3.8-max

Global

30,000

5,000,000

qwen3.8-max-0902

Global

30,000

150,000

qwen3.8-flash

Global

30,000

5,000,000

qwen3.7-max

Global

30,000

5,000,000

qwen3.7-max

United States

600

1,000,000

qwen3.7-max-2026-06-08

Global

600

1,000,000

qwen3.7-max-2026-05-20

Global

600

1,000,000

qwen3-max

Global

600

1,000,000

qwen3-max-preview

Global

600

1,000,000

qwen3-max-2025-09-23

Global

60

100,000

qwen3.7-plus

Global

30,000

5,000,000

qwen3.7-plus

United States

15,000

5,000,000

qwen3.7-plus-2026-05-26

Global

600

1,000,000

qwen3.6-plus

Global

30,000

5,000,000

qwen3.6-plus-2026-04-02

Global

600

1,000,000

qwen3.7-flash

Global

15,000

5,000,000

qwen3.7-flash-2026-07-15

Global

60

1,000,000

qwen3.6-flash

Global

15,000

5,000,000

qwen3.6-flash-2026-04-16

Global

60

1,000,000

qwen3.6-flash

United States

15,000

5,000,000

qwen3.5-plus

Global

30,000

5,000,000

qwen3.5-plus-2026-02-15

Global

600

1,000,000

qwen-plus

Global

15,000

5,000,000

qwen-plus

United States

600

1,000,000

qwen-plus-2025-12-01

Global

60

1,000,000

qwen-plus-2025-09-11

Global

60

1,000,000

qwen-plus-2025-07-28

Global

60

1,000,000

qwen-plus-2025-12-01

United States

60

1,000,000

qwen3.5-flash

Global

30,000

10,000,000

qwen3.5-flash-2026-02-23

Global

600

1,000,000

qwen-flash

Global

15,000

10,000,000

qwen-flash

United States

30000

10,000,000

qwen-flash-2025-07-28

Global

60

1,000,000

qwen-flash-2025-07-28

United States

600

5,000,000

华北2(北京)

ModelRate limit conditions (triggered when any value is exceeded)
Below are the rate limit conditions per minute. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

qwen3.8-max

UseBatch APIWhen calling the service, it is not subject to rate limits.

Dynamic rate limiting

qwen3.8-max-0902

Dynamic rate limiting

qwen3.8-flash

UseBatch APIWhen calling the service, it is not subject to rate limiting.

Dynamic rate limiting

qwen3.7-max

UseBatch APIWhen calling the service, it is not subject to rate limiting.

30,000

5,000,000

qwen3.7-max-2026-06-08

600

1,000,000

qwen3.7-max-2026-05-20

600

1,000,000

qwen3.6-max-preview

600

1,000,000

qwen3-max

UseBatch APIWhen calling the service, it is not subject to rate limiting.

30,000

5,000,000

qwen3-max-2026-01-23

600

1,000,000

qwen3-max-2025-09-23

60

100,000

qwen3-max-preview

600

1,000,000

qwen-max

UseBatch APIWhen calling the service, it is not subject to rate limiting.

1,200

1,000,000

qwen3.7-plus

30,000

5,000,000

qwen3.7-plus-2026-05-26

600

1,000,000

qwen3.6-plus

UseBatch APIWhen calling the service, it is not subject to rate limiting.

30,000

5,000,000

qwen3.6-plus-2026-04-02

600

1,000,000

qwen3.7-flash

UseBatch APIWhen calling the service, it is not subject to rate limiting.

30,000

5,000,000

qwen3.7-flash-2026-07-15

30,000

5,000,000

qwen3.6-flash

UseBatch APIWhen calling the service, it is not subject to rate limiting.

30,000

10,000,000

qwen3.6-flash-2026-04-16

600

1,000,000

qwen3.5-plus

UseBatch APIWhen calling the service, it is not subject to rate limiting.

30,000

5,000,000

qwen3.5-plus-2026-04-20

600

1,000,000

qwen3.5-plus-2026-02-15

600

1,000,000

qwen-plus

UseBatch APIWhen calling the service, it is not subject to rate limiting.

30,000

5,000,000

qwen-plus-latest

UseBatch APIWhen calling the service, it is not subject to rate limiting.

15,000

1,200,000

qwen-plus-2025-12-01

120

1,000,000

qwen-plus-2025-09-11

60

1,000,000

qwen-plus-2025-07-28

(qwen-plus-0728)

60

1,000,000

qwen-plus-2025-07-14

(qwen-plus-0714)

60

100,000

qwen-plus-2025-04-28

(qwen-plus-0428)

60

1,000,000

qwen-plus-2025-01-25

(qwen-plus-0125)

60

150,000

qwen-plus-2025-01-12

(qwen-plus-0112)

60

150,000

qwen-plus-2024-12-20

(qwen-plus-1220)

60

150,000

qwen3.5-flash

UseBatch APIWhen calling the service, it is not subject to rate limiting.

30,000

10,000,000

qwen3.5-flash-2026-02-23

600

1,000,000

qwen-flash

UseBatch APIWhen calling the service, it is not subject to rate limiting.

30,000

10,000,000

qwen-flash-2025-07-28

60

1,000,000

qwq-plus

UseBatch APIWhen calling the service, it is not subject to rate limiting.

600

1,000,000

qwen-turbo

1,200

5,000,000

qwen-long-latest

UseBatch API调用服务时,不受限流限制。

1,200

60,000

qwen-long-2025-01-25

(qwen-long-0125)

3

7,500

德国(法兰克福)

ModelService Deployment Scope限流条件(超出任一数值时触发限流)
以下为每分钟限流条件,服务可能按 RPS(RPM/60)与 TPS(TPM/60)限制
每分钟调用次数(RPM)每分钟消耗Token数(TPM)
含输入与Output Token

qwen3.8-max

Global

30,000

5,000,000

qwen3.8-max-0902

Global

30,000

150,000

qwen3.8-flash

Global

30,000

5,000,000

qwen3.7-max

Global

30,000

5,000,000

qwen3.7-max-2026-06-08

Global

600

1,000,000

qwen3.7-max-2026-05-20

Global

600

1,000,000

qwen3-max

Global

600

1,000,000

qwen3-max

European Union

600

1,000,000

qwen3-max-preview

Global

600

1,000,000

qwen3-max-2026-01-23

European Union

60

100,000

qwen3-max-2025-09-23

Global

60

100,000

qwen3.7-plus

Global

30,000

5,000,000

qwen3.7-plus-2026-05-26

Global

600

1,000,000

qwen3.6-plus

Global

30,000

5,000,000

qwen3.6-plus-2026-04-02

Global

600

1,000,000

qwen3.7-flash

Global

15,000

5,000,000

qwen3.7-flash-2026-07-15

Global

15,000

5,000,000

qwen3.6-flash

Global

15,000

5,000,000

qwen3.6-flash-2026-04-16

Global

60

1,000,000

qwen3.5-plus

Global

30,000

5,000,000

qwen3.5-plus-2026-02-15

Global

600

1,000,000

qwen-plus

Global

600

1,000,000

qwen-plus

European Union

600

1,000,000

qwen-plus-2025-12-01

Global

120

1,000,000

qwen-plus-2025-12-01

European Union

120

1,000,000

qwen-plus-2025-09-11

Global

60

1,000,000

qwen-plus-2025-07-28

Global

60

1,000,000

qwen3.5-flash

Global

35,000

10,000,000

qwen3.5-flash

European Union

35,000

10,000,000

qwen3.5-flash-2026-02-23

Global

600

1,000,000

qwen3.5-flash-2026-02-23

European Union

600

1,000,000

qwen-flash

Global

30,000

10,000,000

qwen-flash-2025-07-28

Global

60

1,000,000

中国香港

ModelService Deployment ScopeRate limiting conditions (triggered when any value is exceeded)
The following are per-minute rate limiting conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

qwen3.8-max

Hong Kong (China)

600

1,000,000

qwen3.8-max-0902

Hong Kong (China)

600

150,000

qwen3.8-flash

Global

30,000

5,000,000

qwen3-max

Hong Kong (China)

600

1,000,000

qwen3-max-2026-01-23

Hong Kong (China)

60

100,000

qwen3.6-plus

Global

30,000

5,000,000

qwen3.7-flash

Global

15,000

5,000,000

qwen3.7-flash-2026-07-15

Global

15,000

5,000,000

qwen3.6-flash

Global

30,000

10,000,000

qwen-plus

Hong Kong (China)

600

1,000,000

qwen-plus-2025-12-01

Hong Kong (China)

60

100,000

qwen3.5-flash

Hong Kong (China)

30,000

10,000,000

qwen3.5-flash-2026-02-23

Hong Kong (China)

600

1,000,000

日本(东京)

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

qwen3.8-max

Global

30,000

5,000,000

qwen3.8-max-0902

Global

30,000

150,000

qwen3.8-flash

Global

30,000

5,000,000

qwen3.7-max

Global

600

1,000,000

qwen3.7-max-2026-05-20

Global

60

1,000,000

qwen3.7-plus

Global

15,000

5,000,000

qwen3.7-plus-2026-05-26

Global

60

1,000,000

qwen3.7-plus

Japan

15,000

5,000,000

qwen3.7-plus-2026-05-26

Japan

60

1,000,000

qwen3.6-plus

Global

15,000

5,000,000

qwen3.6-plus-2026-04-02

Global

60

1,000,000

qwen3.7-flash

Global

15,000

5,000,000

qwen3.7-flash-2026-07-15

Global

15,000

5,000,000

qwen3.6-flash

Global

15,000

10,000,000

qwen3.6-flash-2026-04-16

Global

60

1,000,000

Qwen-VL (Visual Understanding/Image-to-Text)

新加坡

ModelService Deployment ScopeThrottling Conditions (Triggered When Any Value Is Exceeded)
The following are per-minute throttling conditions. The service may throttle by RPS (RPM/60) and TPS (TPM/60).
Calls Per Minute (RPM)Tokens Consumed Per Minute (TPM)
Includes input and output Tokens

qwen3-vl-plus

International

1,200

1,000,000

qwen3-vl-plus-2025-12-19

International

60

100,000

qwen3-vl-plus-2025-09-23

International

120

1,000,000

qwen3-vl-flash

International

1,200

1,000,000

qwen3-vl-flash-2026-01-22

International

60

100,000

qwen3-vl-flash-2025-10-15

International

120

1,000,000

qwen-vl-max

International

1,200

1,000,000

qwen-vl-plus

International

1,200

1,000,000

qvq-max

International

60

100,000

美国(弗吉尼亚)

ModelService Deployment ScopeThrottling conditions (throttling is triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may throttle based on RPS (RPM/60) and TPS (TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes input and Output Tokens

qwen3-vl-plus

Global

6,000

10,000,000

qwen3-vl-plus-2025-09-23

Global

60

100,000

qwen3-vl-flash

Global

1,200

1,000,000

qwen3-vl-flash

United States

1,200

1,000,000

qwen3-vl-flash-2025-10-15

Global

60

100,000

qwen3-vl-flash-2026-01-22

United States

120

1,000,000

qwen3-vl-flash-2025-10-15

United States

120

1,000,000

华北2(北京)

ModelThrottling conditions (triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes Input and Output Tokens

qwen3-vl-plus

UseBatch APIWhen calling the service, it is not subject to rate limiting.

3,000

5,000,000

qwen3-vl-plus-2025-12-19

60

100,000

qwen3-vl-plus-2025-09-23

60

100,000

qwen3-vl-flash

UseBatch APIWhen calling the service, it is not subject to rate limiting.

3,000

5,000,000

qwen3-vl-flash-2026-01-22

60

100,000

qwen3-vl-flash-2025-10-15

60

100,000

qwen-vl-max

UseBatch APIWhen calling the service, it is not subject to rate limiting.

1,200

1,000,000

qwen-vl-plus

UseBatch APIWhen calling the service, it is not subject to rate limiting.

1,200

1,000,000

qvq-max

60

100,000

qvq-plus

60

100,000

德国(法兰克福)

ModelService Deployment Scope限流条件(超出任一数值时触发限流)
以下为每分钟限流条件,服务可能按 RPS(RPM/60)与 TPS(TPM/60)限制
每分钟调用次数(RPM)每分钟消耗Token数(TPM)
含输入与输出Token

qwen3-vl-plus

全球

2,000

1,000,000

qwen3-vl-plus

欧盟

2,000

1,000,000

qwen3-vl-plus-2025-09-23

全球

60

100,000

qwen3-vl-flash

全球

12,000

10,000,000

qwen3-vl-flash

欧盟

12,000

10,000,000

qwen3-vl-flash-2026-01-22

European Union

60

100,000

qwen3-vl-flash-2025-10-15

Global

600

1,000,000

qwen3-vl-flash-2025-10-15

European Union

600

1,000,000

中国香港

ModelService Deployment ScopeRate limiting conditions (triggered when any threshold is exceeded)
The following are per-minute rate limiting conditions. The service may enforce limits based on RPS (RPM/60) and TPS (TPM/60).
Requests Per Minute (RPM)Tokens Per Minute (TPM)
Includes Input and Output Tokens

qwen3-vl-plus

Hong Kong (China)

6,000

1,000,000

qwen3-vl-plus-2025-12-19

Hong Kong (China)

6,000

1,000,000

Qwen Omni (Omni-modality)

新加坡

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60)
Calls per minute (RPM)Tokens consumed per minute (TPM)
Including input and output Tokens

qwen3.8-omni-flash

International

30000

Dynamic Rate Limiting

qwen3.5-omni-flash

International

60

100,000

qwen3.5-omni-flash-2026-03-15

International

60

100,000

qwen3.5-omni-plus

International

60

100,000

qwen3.5-omni-plus-2026-03-15

International

60

100,000

qwen3-omni-flash

International

60

100,000

qwen3-omni-flash-2025-12-01

International

60

100,000

qwen3-omni-flash-2025-09-15

International

60

100,000

qwen-omni-turbo

International

60

100,000

qwen-omni-turbo-latest

International

60

100,000

qwen-omni-turbo-2025-03-26

International

60

100,000

华北2(北京)

ModelThrottling Conditions (triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may throttle by RPS (RPM/60) and TPS (TPM/60).
Calls per Minute (RPM)Tokens Consumed per Minute (TPM)
Includes Input and Output Tokens

qwen3.8-omni-flash

30000

Dynamic Throttling

qwen3.5-omni-flash

60

100,000

qwen3.5-omni-flash-2026-03-15

60

100,000

qwen3.5-omni-plus

60

100,000

qwen3.5-omni-plus-2026-03-15

60

100,000

qwen3-omni-flash

60

100,000

qwen3-omni-flash-2025-12-01

60

100,000

qwen3-omni-flash-2025-09-15

60

100,000

qwen-omni-turbo

60

100,000

qwen-omni-turbo-latest

60

100,000

qwen-omni-turbo-2025-03-26

(qwen-omni-turbo-0326)

60

100,000

qwen-omni-turbo-2025-01-19

(qwen-omni-turbo-0119)

60

100,000

中国香港

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
Rate limit conditions per minute, the service may limit by RPS (RPM/60) and TPS (TPM/60)
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and Output Tokens

qwen3.8-omni-flash

Global

30000

2000000

日本(东京)

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60)
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

qwen3.8-omni-flash

Global

30000

2000000

德国(法兰克福)

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Including input and output Tokens

qwen3.8-omni-flash

Global

30000

2000000

美国(弗吉尼亚)

ModelService Deployment ScopeThrottling conditions (triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Including input and output Tokens

qwen3.8-omni-flash

Global

30000

2000000

Qwen-Omni-Realtime

新加坡

ModelService Deployment ScopeRate limiting conditions (triggered when any value is exceeded)
The following are per-minute rate limiting conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output tokens

qwen3.8-omni-flash-realtime

International

60

2,000,000

qwen3.5-omni-plus-realtime

International

60

100,000

qwen3.5-omni-plus-realtime-2026-03-15

International

60

100,000

qwen3.5-omni-flash-realtime

International

60

100,000

qwen3.5-omni-flash-realtime-2026-03-15

International

60

100,000

qwen3-omni-flash-realtime

International

60

100,000

qwen3-omni-flash-realtime-2025-12-01

International

60

100,000

qwen3-omni-flash-realtime-2025-09-15

International

60

100,000

qwen-omni-turbo-realtime

International

60

10,000

qwen-omni-turbo-realtime-latest

International

60

10,000

qwen-omni-turbo-realtime-2025-05-08

International

60

10,000

华北2(北京)

ModelRate limiting conditions (triggered when any value is exceeded)
The following are per-minute rate limiting conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60)
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output tokens

qwen3.8-omni-flash-realtime

60

2,000,000

qwen3.5-omni-plus-realtime

60

100,000

qwen3.5-omni-plus-realtime-2026-03-15

60

100,000

qwen3.5-omni-flash-realtime

60

100,000

qwen3.5-omni-flash-realtime-2026-03-15

60

100,000

qwen3-omni-flash-realtime

60

100,000

qwen3-omni-flash-realtime-2025-12-01

60

100,000

qwen3-omni-flash-realtime-2025-09-15

60

100,000

qwen-omni-turbo-realtime

60

100,000

qwen-omni-turbo-realtime-latest

60

100,000

qwen-omni-turbo-realtime-2025-05-08

60

100,000

Qwen-OCR (Text Extraction)

新加坡

ModelService Deployment ScopeRate limiting conditions (triggered when any value is exceeded)
The following are per-minute rate limiting conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60)
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

qwen-vl-ocr

International

600

6,000,000

qwen-vl-ocr-2025-11-20

International

1,200

6,000,000

美国(弗吉尼亚)

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Token count consumed per minute (TPM)
Includes input and output Tokens

qwen-vl-ocr

Global

600

6,000,000

qwen-vl-ocr-2025-11-20

Global

1,200

6,000,000

华北2(北京)

ModelRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60)
Requests per minute (RPM)Token count consumed per minute (TPM)
Includes input and output Tokens

qwen3.5-ocr

6,000

30,000,000

qwen-vl-ocr

UseBatch APIWhen calling the service, it is not subject to rate limits.

600

6,000,000

qwen-vl-ocr-latest

6,000

30,000,000

qwen-vl-ocr-2025-11-20

6,000

30,000,000

qwen-vl-ocr-2025-04-13

600

6,000,000

qwen-vl-ocr-2024-10-28

600

6,000,000

德国(法兰克福)

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes input and output tokens

qwen-vl-ocr

Global

600

6,000,000

qwen-vl-ocr-2025-11-20

Global

1,200

6,000,000

Qwen Math Model

华北2(北京)

ModelRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output tokens

qwen-math-plus

1,200

1,000,000

qwen-math-plus-latest

1,200

1,000,000

qwen-math-plus-2024-09-19

(qwen-math-plus-0919)

60

100,000

qwen-math-plus-2024-08-16

(qwen-math-plus-0816)

10

20,000

qwen-math-turbo

1200

1,000,000

Qwen-Coder

新加坡

ModelService Deployment ScopeThrottling conditions (throttling is triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may throttle based on RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

qwen3-coder-plus

International

2,400

2,000,000

qwen3-coder-plus-2025-09-23

International

600

1,000,000

qwen3-coder-plus-2025-07-22

International

60

1,000,000

qwen3-coder-flash

International

600

5,000,000

qwen3-coder-flash-2025-07-28

International

600

5,000,000

美国(弗吉尼亚)

ModelService Deployment ScopeRate limiting conditions (triggered when any value is exceeded)
The following are per-minute rate limiting conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output tokens

qwen3-coder-plus

Global

2,400

2,000,000

qwen3-coder-plus-2025-09-23

Global

60

1,000,000

qwen3-coder-plus-2025-07-22

Global

60

1,000,000

qwen3-coder-flash

Global

1,200

1,000,000

qwen3-coder-flash-2025-07-28

Global

60

1,000,000

华北2(北京)

ModelRate limiting conditions (triggered when any value is exceeded)
The following are per-minute rate limiting conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Requests per minute (RPM)Tokens consumed per minute (TPM)
Including input and Output tokens

qwen3-coder-plus

5,000

5,000,000

qwen3-coder-plus-2025-09-23

60

1,000,000

qwen3-coder-plus-2025-07-22

60

1,000,000

qwen3-coder-flash

5,000

5,000,000

qwen3-coder-flash-2025-07-28

60

1,000,000

qwen-coder-plus

1,200

1,000,000

qwen-coder-turbo

1,200

1,000,000

德国(法兰克福)

ModelService Deployment Scope限流条件(超出任一数值时触发限流)
以下为每分钟限流条件,服务可能按 RPS(RPM/60)与 TPS(TPM/60)限制
每分钟调用次数(RPM)每分钟消耗Token数(TPM)
含输入与输出Token

qwen3-coder-plus

Global

5,000

5,000,000

qwen3-coder-plus-2025-09-23

Global

60

1,000,000

qwen3-coder-plus-2025-07-22

Global

60

1,000,000

qwen3-coder-flash

Global

5,000

5,000,000

qwen3-coder-flash-2025-07-28

Global

60

1,000,000

Qwen Translation Model

新加坡

ModelService Deployment ScopeThrottling Conditions (triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may throttle by RPS (RPM/60) and TPS (TPM/60).
Requests Per Minute (RPM)Tokens Per Minute (TPM)
Includes input and output tokens

qwen-mt-plus

International

60

100,000

qwen-mt-flash

International

60

100,000

qwen-mt-lite

International

60

100,000

qwen-mt-turbo

International

60

100,000

美国(弗吉尼亚)

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may enforce limits via RPS (RPM/60) and TPS (TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes input and output tokens

qwen-mt-plus

Global

60

25,000

qwen-mt-flash

Global

60

35,000

qwen-mt-lite

Global

60

100,000

qwen-mt-lite

United States

60

100,000

华北2(北京)

ModelRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

qwen-mt-plus

60

25,000

qwen-mt-flash

60

35,000

qwen-mt-lite

60

100,000

qwen-mt-turbo

60

35,000

德国(法兰克福)

ModelService Deployment ScopeThrottling conditions (triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may also limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

qwen-mt-plus

Global

60

25,000

qwen-mt-flash

Global

60

35,000

qwen-mt-lite

Global

60

100,000

Qwen Data Mining Model

华北2(北京)

ModelRate limit conditions (rate limit triggered when any value is exceeded)
The following are rate limit conditions per minute. The service may limit based on RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

qwen-doc-turbo

600

3,000,000

Qwen In-Depth Research Model

华北2(北京)

ModelRate limit conditions (rate limit triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60)
Number of calls per minute (RPM)Number of Tokens consumed per minute (TPM)
Includes Input and Output Tokens

qwen-deep-research

120

1,200,000

Text Generation-Qwen-Open Source Edition

Qwen Language Model Open Source Edition

新加坡

ModelService Deployment ScopeRate limit conditions (rate limit triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60)
每分钟调用次数(RPM)每分钟消耗Token数(TPM)
含输入与输出Token

qwen3.8-2.4t-a95b

国际

动态限流

qwen3.8-27b

国际

动态限流

qwen3.6-35b-a3b

国际

600

1,000,000

qwen3.6-27b

国际

600

1,000,000

qwen3.5-397b-a17b

国际

600

1,000,000

qwen3.5-122b-a10b

International

600

1,000,000

qwen3.5-27b

International

600

1,000,000

qwen3.5-35b-a3b

International

600

5,000,000

qwen3-next-80b-a3b-thinking

International

600

1,000,000

qwen3-next-80b-a3b-instruct

International

600

1,000,000

qwen3-235b-a22b-thinking-2507

International

600

1,000,000

qwen3-235b-a22b-instruct-2507

International

600

1,000,000

qwen3-30b-a3b-thinking-2507

International

600

5,000,000

qwen3-30b-a3b-instruct-2507

International

600

5,000,000

qwen3-235b-a22b

International

600

1,000,000

qwen3-32b

International

600

1,000,000

qwen3-30b-a3b

International

600

1,000,000

qwen3-14b

International

600

1,000,000

qwen3-8b

International

600

1,000,000

美国(弗吉尼亚)

ModelService Deployment ScopeThrottling Conditions (triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may throttle based on RPS (RPM/60) and TPS (TPM/60).
Requests per Minute (RPM)Tokens per Minute (TPM)
Includes Input and Output Tokens

qwen3.5-397b-a17b

Global

600

1,000,000

qwen3.5-122b-a10b

Global

600

1,000,000

qwen3.5-27b

Global

600

1,000,000

qwen3.6-35b-a3b

Global

600

1,000,000

qwen3.5-35b-a3b

Global

600

1,000,000

qwen3-next-80b-a3b-thinking

Global

600

1,000,000

qwen3-next-80b-a3b-instruct

Global

600

1,000,000

qwen3-235b-a22b-thinking-2507

Global

600

1,000,000

qwen3-235b-a22b-instruct-2507

Global

600

1,000,000

qwen3-30b-a3b-thinking-2507

Global

600

1,000,000

qwen3-30b-a3b-instruct-2507

Global

600

1,000,000

qwen3-235b-a22b

Global

600

1,000,000

qwen3-30b-a3b

Global

600

1,000,000

qwen3-32b

Global

600

1,000,000

qwen3-14b

Global

600

1,000,000

qwen3-8b

Global

600

1,000,000

华北2(北京)

ModelRate limit conditions (rate limiting is triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes input and Output Tokens

qwen3.8-2.4t-a95b

Dynamic Rate Limiting

qwen3.8-27b

Dynamic Rate Limiting

qwen3.6-35b-a3b

600

1,000,000

qwen3.6-27b

600

1,000,000

qwen3.5-397b-a17b

600

1,000,000

qwen3.5-122b-a10b

600

1,000,000

qwen3.5-27b

600

1,000,000

qwen3.5-35b-a3b

600

1,000,000

qwen3-next-80b-a3b-thinking

600

1,000,000

qwen3-next-80b-a3b-instruct

600

1,000,000

qwen3-235b-a22b-thinking-2507

600

1,000,000

qwen3-235b-a22b-instruct-2507

600

1,000,000

qwen3-30b-a3b-thinking-2507

600

1,000,000

qwen3-30b-a3b-instruct-2507

600

1,000,000

qwen3-235b-a22b

600

1,000,000

qwen3-30b-a3b

600

1,000,000

qwen3-32b

2400

1,000,000

qwen3-14b

600

1,000,000

qwen3-8b

600

1,000,000

德国(法兰克福)

ModelService Deployment ScopeRate Limiting Conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit based on RPS (RPM/60) and TPS (TPM/60)
Requests per minute (RPM)Tokens Consumed Per Minute (TPM)
Including Input and Output Tokens

qwen3.5-397b-a17b

Global

600

1,000,000

qwen3.5-122b-a10b

Global

600

1,000,000

qwen3.5-27b

Global

600

1,000,000

qwen3.6-35b-a3b

Global

600

1,000,000

qwen3.5-35b-a3b

Global

600

1,000,000

qwen3-next-80b-a3b-thinking

Global

600

1,000,000

qwen3-next-80b-a3b-instruct

Global

600

1,000,000

qwen3-235b-a22b-thinking-2507

Global

600

1,000,000

qwen3-235b-a22b-instruct-2507

Global

600

1,000,000

qwen3-30b-a3b-thinking-2507

Global

600

1,000,000

qwen3-30b-a3b-instruct-2507

Global

600

1,000,000

qwen3-235b-a22b

Global

600

1,000,000

qwen3-30b-a3b

Global

600

1,000,000

qwen3-32b

Global

2,400

1,000,000

qwen3-14b

Global

600

1,000,000

qwen3-8b

Global

600

1,000,000

Qwen-VL(Visual Understanding/Image-to-Text)

新加坡

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output tokens

qwen3-vl-32b-thinking

International

60

100,000

qwen3-vl-32b-instruct

International

60

100,000

qwen3-vl-30b-a3b-thinking

International

60

100,000

qwen3-vl-30b-a3b-instruct

International

60

100,000

qwen3-vl-8b-thinking

International

60

100,000

qwen3-vl-8b-instruct

International

60

100,000

qwen3-vl-235b-a22b-thinking

International

60

100,000

qwen3-vl-235b-a22b-instruct

International

60

100,000

美国(弗吉尼亚)

ModelService Deployment ScopeThrottling conditions (triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Token

qwen3-vl-235b-a22b-thinking

Global

60

100,000

qwen3-vl-235b-a22b-instruct

Global

60

100,000

qwen3-vl-32b-thinking

Global

600

1,000,000

qwen3-vl-32b-instruct

Global

600

1,000,000

qwen3-vl-30b-a3b-thinking

Global

600

1,000,000

qwen3-vl-30b-a3b-instruct

Global

600

1,000,000

qwen3-vl-8b-thinking

Global

600

1,000,000

qwen3-vl-8b-instruct

Global

600

1,000,000

华北2(北京)

ModelRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit based on RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

qwen3-vl-32b-thinking

600

1,000,000

qwen3-vl-32b-instruct

600

1,000,000

qwen3-vl-30b-a3b-thinking

600

1,000,000

qwen3-vl-30b-a3b-instruct

600

1,000,000

qwen3-vl-8b-thinking

600

1,000,000

qwen3-vl-8b-instruct

600

1,000,000

qwen3-vl-235b-a22b-thinking

60

100,000

qwen3-vl-235b-a22b-instruct

60

100,000

德国(法兰克福)

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit based on RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)每分钟消耗Token数(TPM)
含输入与输出Token

qwen3-vl-235b-a22b-thinking

Global

60

100,000

qwen3-vl-235b-a22b-instruct

Global

60

100,000

qwen3-vl-32b-thinking

Global

600

1,000,000

qwen3-vl-32b-instruct

Global

600

1,000,000

qwen3-vl-30b-a3b-thinking

Global

600

1,000,000

qwen3-vl-30b-a3b-instruct

Global

600

1,000,000

qwen3-vl-8b-thinking

Global

600

1,000,000

qwen3-vl-8b-instruct

Global

600

1,000,000

Qwen3-Omni

新加坡

ModelService Deployment ScopeThrottling conditions (triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may impose limits based on RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

qwen2.5-omni-7b

International

60

100,000

华北2(北京)

ModelThrottling conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Including input and Output tokens

qwen2.5-omni-7b

60

100,000

Qwen3-Omni-Captioner

新加坡

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and Output Tokens

qwen3-omni-30b-a3b-captioner

International

60

100,000

华北2(北京)

ModelThrottling conditions (throttling is triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may enforce limits based on RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and Output Tokens

qwen3-omni-30b-a3b-captioner

60

100,000

Qwen-Coder

新加坡

ModelService Deployment ScopeRate Limit Conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60)
Calls Per Minute (RPM)Tokens Per Minute (TPM)
Including Input and Output Tokens

qwen3-coder-next

International

600

1,000,000

qwen3-coder-480b-a35b-instruct

International

600

1,000,000

qwen3-coder-30b-a3b-instruct

International

600

1,000,000

美国(弗吉尼亚)

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes Input and Output tokens

qwen3-coder-480b-a35b-instruct

Global

600

1,000,000

qwen3-coder-30b-a3b-instruct

Global

600

1,000,000

华北2(北京)

ModelRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes Input and Output Token

qwen3-coder-next

600

1,000,000

qwen3-coder-480b-a35b-instruct

600

1,000,000

qwen3-coder-30b-a3b-instruct

600

1,000,000

德国(法兰克福)

ModelService Deployment ScopeThrottling conditions (triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may also limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes Input and Output Token

qwen3-coder-480b-a35b-instruct

Global

600

1,000,000

qwen3-coder-30b-a3b-instruct

Global

600

1,000,000

qwen3-coder-next

European Union

600

1,000,000

Text Generation-Third-party Models

DeepSeek

新加坡

ModelService Deployment ScopeThrottling Conditions (triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may throttle by RPS (RPM/60) and TPS (TPM/60).
Calls per Minute (RPM)Tokens Consumed per Minute (TPM)
Includes Input and Output Tokens

deepseek-v4.1-flash

International

10,000

1,200,000

deepseek-v4-pro

International

10,000

1,200,000

deepseek-v4-pro-0813

International

Dynamic Rate Limiting

deepseek-v4-flash-0731

International

15,000

1,200,000

deepseek-v4-flash

International

10,000

1,200,000

deepseek-v3.2

International

10,000

1,200,000

美国(弗吉尼亚)

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes Input and Output Tokens

deepseek-v4.1-flash

Global

15,000

1,200,000

deepseek-v4-pro

Global

15,000

1,200,000

deepseek-v4-pro-0813

Global

15,000

1,200,000

deepseek-v4-pro

United States

10,000

1,200,000

deepseek-v4-flash-0731

Global

15,000

1,200,000

deepseek-v4-flash

Global

15,000

1,200,000

deepseek-v4-flash

United States

10,000

1,200,000

deepseek-v4-flash-0731

United States

15,000

1,200,000

华北2(北京)

ModelRate limit conditions (rate limit is triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Including input and Output Tokens

deepseek-v4.1-flash

15,000

1,200,000

deepseek-v4-pro

15,000

1,200,000

deepseek-v4-pro-0813

Dynamic rate limiting

deepseek-v4-flash-0731

15,000

1,200,000

deepseek-v4-flash

15,000

1,200,000

deepseek-v3.2

UseBatch APIWhen calling the service, it is not subject to rate limiting.

15,000

1,000,000

deepseek-v3.2-exp

15,000

1,200,000

deepseek-v3.1

15,000

1,200,000

deepseek-r1-0528

60

100,000

deepseek-r1

UseBatch APIWhen calling the service, it is not subject to rate limiting.

15,000

1,200,000

deepseek-v3

UseBatch APIWhen calling the service, it is not subject to rate limiting.

15,000

1,200,000

deepseek-r1-distill-qwen-7b

15,000

1,200,000

deepseek-r1-distill-qwen-14b

15,000

1,200,000

deepseek-r1-distill-qwen-32b

15,000

1,200,000

deepseek-r1-distill-qwen-1.5b

60

100,000

deepseek-r1-distill-llama-8b

60

100,000

deepseek-r1-distill-llama-70b

60

100,000

德国(法兰克福)

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Including input and Output Token

deepseek-v4.1-flash

Global

15,000

1,200,000

deepseek-v4-pro

Global

15,000

1,200,000

deepseek-v4-pro-0813

Global

15,000

1,200,000

deepseek-v4-flash-0731

Global

15,000

1,200,000

deepseek-v4-flash

Global

15,000

1,200,000

日本(东京)

ModelService Deployment ScopeRate limiting conditions (triggered when any value is exceeded)
The following are rate limiting conditions per minute; the service may limit by RPS (RPM/60) and TPS (TPM/60)
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

deepseek-v4-pro

Japan

10,000

1,200,000

deepseek-v4-pro-0813

Japan

15,000

1,200,000

deepseek-v4-flash

Japan

10,000

1,200,000

deepseek-v4-flash-0731

Japan

15,000

1,200,000

deepseek-v4.1-flash

Global

15,000

1,200,000

deepseek-v4-pro

Global

10,000

1,200,000

deepseek-v4-pro-0813

Global

15,000

1,200,000

deepseek-v4-flash-0731

Global

15,000

1,200,000

deepseek-v4-flash

Global

10,000

1,200,000

中国香港

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may apply limits based on RPS (RPM/60) and TPS (TPM/60)
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

deepseek-v4.1-flash

Hong Kong (China)

15,000

1,200,000

deepseek-v4-flash-0731

Hong Kong (China)

15,000

1,200,000

deepseek-v4-pro-0813

Hong Kong (China)

15,000

1,200,000

Kimi

华北2(北京)

ModelRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit requests by RPS (RPM/60) and TPS (TPM/60)
Requests per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

kimi-k3

Dynamic Rate Limiting

kimi-k2.7-code

500

1,000,000

kimi-k2.6

500

1,000,000

kimi-k2.5

500

1,000,000

kimi-k2-thinking

500

1,000,000

Moonshot-Kimi-K2-Instruct

500

1,000,000

美国(弗吉尼亚)

ModelService Deployment ScopeRate Limiting Conditions (triggered when any value is exceeded)
The following are per-minute rate limiting conditions. The service may impose limits based on RPS (RPM/60) and TPS (TPM/60).
Requests per Minute (RPM)Tokens per Minute (TPM)
Includes Input and Output Tokens

kimi-k3

Global

15,000

1,200,000

kimi-k3

International

10,000

1,200,000

kimi-k2.7-code

Global

500

1,000,000

kimi-k2.5

Global

500

1,000,000

德国(法兰克福)

ModelService Deployment ScopeThrottling conditions (throttling is triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may impose limits based on RPS (RPM/60) and TPS (TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes Input and Output Tokens

kimi-k3

Global

15,000

1,200,000

kimi-k2.7-code

Global

500

1,000,000

kimi-k2.5

Global

500

1,000,000

中国香港

ModelService Deployment ScopeRate Limit Conditions (rate limit is triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

kimi-k3

Global

15,000

1,200,000

kimi-k2.7-code

Global

500

1,000,000

日本(东京)

ModelService Deployment ScopeRate Limit Conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls Per Minute (RPM)Tokens Per Minute (TPM)
Includes input and output tokens

kimi-k3

Global

15,000

1,200,000

kimi-k2.7-code

Global

500

1,000,000

kimi-k2.5

Global

500

1,000,000

新加坡

ModelService Deployment ScopeThrottling conditions (triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may throttle based on RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes Input and Output Tokens

kimi-k3

International

Dynamic Throttling

kimi-k2.7-code

International

500

1,000,000

MiniMax

华北2(北京)

ModelRate limit conditions (triggered when any value is exceeded)
Below are the rate limit conditions per minute. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Including input and output tokens

MiniMax-M2.5

500

1,000,000

GLM

美国(弗吉尼亚)

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Including input and output Tokens

glm-5.3

Global

15,000

1,200,000

glm-5.2

Global

500

1,000,000

glm-5.2

United States

500

1,000,000

glm-5.1

Global

500

1,000,000

华北2(北京)

ModelRate limit conditions (triggered when any value is exceeded)
The following are rate limit conditions per minute. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

glm-5.3

15,000

5,000,000

glm-5.2

500

2,000,000

glm-5.1

500

1,000,000

glm-5

500

1,000,000

glm-4.7

500

1,000,000

glm-4.6

60

1,000,000

德国(法兰克福)

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are rate limit conditions per minute. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Including Input and Output Tokens

glm-5.3

Global

15,000

1,200,000

glm-5.2

Global

500

1,000,000

glm-5.1

Global

500

1,000,000

新加坡

ModelService Deployment ScopeRate Limit Conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may apply limits based on RPS (RPM/60) and TPS (TPM/60).
Calls Per Minute (RPM)Tokens Consumed Per Minute (TPM)
Includes input and output Tokens

glm-5.3

International

10,000

5,000,000

glm-5.2

International

500

1,000,000

glm-5.1

International

500

1,000,000

中国香港

ModelService Deployment ScopeThrottling conditions (throttling is triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may limit requests by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Token

glm-5.3

Global

15,000

1,200,000

glm-5.2

Global

500

1,000,000

日本(东京)

ModelService Deployment ScopeRate limiting conditions (triggered when any value is exceeded)
The following are per-minute rate limiting conditions. The service may be limited by RPS (RPM/60) and TPS (TPM/60)
Calls per minute (RPM)Token consumption per minute (TPM)
Includes input and output Token

glm-5.3

Global

15,000

1,200,000

glm-5.1

Global

500

1,000,000

GLM-Z.AI直供

新加坡

ModelRate limiting conditions (triggered when any value is exceeded)
The following are per-minute rate limiting conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

ZHIPU/GLM-5.3

200

3,000,000

ZHIPU/GLM-5.2

200

3,000,000

Image Generation

Qwen(Qwen-Image)

新加坡

Model

Service Deployment Scope

Rate limit conditions (triggered when any value is exceeded)

Task submission API call limit

Number of tasks being processed simultaneously (concurrency)

qwen-image-3.0-pro

International

5 Times/Minutes

Synchronous interface unlimited / Asynchronous interface 10

qwen-image-3.0

International

20 Times/ Minutes

Synchronous interface unlimited / Asynchronous interface 10

qwen-image-2.0-pro

International

2 Times/ Minutes

Synchronous interface unlimited

qwen-image-2.0-pro-2026-06-22

International

2 Times/ Minutes

Synchronous interface unlimited

qwen-image-2.0-pro-2026-04-22

International

2 Times/ Minutes

同步接口无限制

qwen-image-2.0-pro-2026-03-03

International

2 Times/ Minutes

同步接口无限制

qwen-image-2.0

International

2 Times/ s

同步接口无限制

qwen-image-2.0-2026-03-03

International

2 Times/ s

同步接口无限制

qwen-image-max

International

2 Times/ Minutes

Synchronous interface no limit

qwen-image-max-2025-12-30

International

2 Times/ Minutes

Synchronous interface no limit

qwen-image-plus

International

2 Times/ s

Synchronous interface no limit / Asynchronous interface 2

qwen-image-plus-2026-01-09

International

2 Times/ s

Synchronous interface unlimited

qwen-image

International

2 Times/ s

Synchronous interface unlimited / Asynchronous interface 2

qwen-image-edit-max

International

2 Times/ Minutes

Synchronous interface unlimited

qwen-image-edit-max-2026-01-16

International

2 Times/ Minutes

Sync API: no limit

qwen-image-edit-plus

International

2 Times/s

Sync API: no limit

qwen-image-edit-plus-2025-12-15

International

2 Times/s

Sync API: no limit

qwen-image-edit-plus-2025-10-30

International

2 Times/s

Sync API: no limit

qwen-image-edit

International

2 Times/ s

Synchronous interface unlimited

qwen-mt-image-2.0

International

60 Times/ Minutes

Synchronous interface unlimited / Asynchronous interface 2

美国(弗吉尼亚)

Model

Service Deployment Scope

Rate Limit Conditions (rate limit triggered when any value is exceeded)

Task Submission API Call Limit

Number of Tasks Being Processed Simultaneously (Concurrency)

qwen-image-3.0-pro

Global

5 Times/Minutes

Synchronous interface unlimited / Asynchronous interface 10

qwen-image-3.0

Global

20 Times/Minutes

Synchronous interface unlimited / Asynchronous interface 10

华北2(北京)

Model

Rate limit conditions (triggered when any value is exceeded)

Task submission interface call limit

Number of tasks being processed simultaneously (concurrency)

qwen-image-3.0-pro

5 Times/ Minutes

Synchronous interface unlimited / Asynchronous interface 10

qwen-image-3.0

20 Times/ Minutes

Synchronous interface unlimited / Asynchronous interface 10

qwen-image-2.0-pro

2 Times/ Minutes

Synchronous interface unlimited

qwen-image-2.0-pro-2026-06-22

2 Times/ Minutes

Synchronous interface unlimited

qwen-image-2.0-pro-2026-04-22

2 Times/ Minutes

Synchronous interface unlimited

qwen-image-2.0-pro-2026-03-03

2 Times/ Minutes

Synchronous interface unlimited

qwen-image-2.0

2 Times/ s

Synchronous interface unlimited

qwen-image-2.0-2026-03-03

2 Times/ s

Synchronous interface unlimited

qwen-image-max

2Times/ Minutes

Synchronous interface unlimited

qwen-image-max-2025-12-30

2Times/ Minutes

Synchronous interface unlimited

qwen-image-plus

2 Times/s

Synchronous interface unlimited / Asynchronous interface 2

qwen-image-plus-2026-01-09

2 Times/s

Synchronous interface unlimited

qwen-image

2 Times/s

Synchronous interface unlimited / Asynchronous interface 2

qwen-image-edit-max

2 Times/ Minutes

Synchronous interface unlimited

qwen-image-edit-max-2026-01-16

2 Times/ Minutes

Synchronous interface unlimited

qwen-image-edit-plus

2 Times/s

Synchronous interface: unlimited

qwen-image-edit-plus-2025-12-15

2 Times/s

Synchronous interface: unlimited

qwen-image-edit-plus-2025-10-30

2 Times/s

Synchronous interface: unlimited

qwen-image-edit

2 Times/s

Synchronous interface: unlimited

qwen-mt-image-2.0

60 Times/Minutes

Synchronous interface unlimited / Asynchronous interface 2

qwen-mt-image

1 Times/s

Synchronous interface unlimited / Asynchronous interface 2

德国(法兰克福)

Model

Service Deployment Scope

Throttling conditions (throttling is triggered when any value is exceeded)

Task submission interface call limit

Concurrent tasks being processed (concurrency)

qwen-image-3.0-pro

Global

5 Times/ Minutes

Synchronous interface: no limit / Asynchronous interface: 10

qwen-image-3.0

Global

20 Times/ Minutes

Synchronous interface unlimited / Asynchronous interface 10

中国香港

Model

Service Deployment Scope

Rate limit conditions (triggered when any value is exceeded)

Task submission interface call limit

Number of tasks being processed simultaneously (concurrency)

qwen-image-3.0-pro

Global

5 Times/ Minutes

Synchronous interface unlimited / Asynchronous interface 10

qwen-image-3.0

Global

20 Times/ Minute

Synchronous interface unlimited / Asynchronous interface 10

日本(东京)

Model

Service Deployment Scope

Rate limit conditions (triggered when any value is exceeded)

Task submission interface call limit

Number of tasks being processed simultaneously (concurrency)

qwen-image-3.0-pro

Global

5 Times/ Minute

Synchronous interface unlimited / Asynchronous interface 10

qwen-image-3.0

Global

20 Times/Minutes

Sync interface unlimited / Async interface 10

Image Generation-Z-Image

新加坡

Model

Service Deployment Scope

Rate limit conditions (triggered when any value is exceeded)

Calls per second (RPS)

Number of tasks being processed simultaneously (concurrency)

z-image-turbo

International

2

Synchronous interface: unlimited

华北2(北京)

Model

Rate limit conditions (rate limiting is triggered when any value is exceeded)

Calls per second (RPS)

Number of concurrent tasks being processed (concurrency)

z-image-turbo

2

Synchronous interface: unlimited

Wanxiang

新加坡

Model

Service Deployment Scope

Rate limit conditions (rate limiting is triggered when any value is exceeded)

Requests per second (RPS)

Number of tasks being processed simultaneously (concurrency)

wan2.7-image-pro

International

5

5

wan2.7-image

International

5

5

wan2.6-image

International

5

5

wan2.6-t2i

International

5

5

wan2.5-t2i-preview

International

5

5

wan2.2-t2i-flash

International

2

2

wan2.2-t2i-plus

International

2

2

wan2.1-t2i-turbo

International

2

2

wan2.1-t2i-plus

International

2

2

wan2.5-i2i-preview

International

5

5

美国(弗吉尼亚)

Model

Service Deployment Scope

Rate limiting conditions (rate limiting triggered when any value is exceeded)

Calls per second (RPS)

Number of tasks being processed concurrently (concurrency)

wan2.6-t2i

Global

5

5

wan2.6-image

Global

5

5

华北2(北京)

Model

Rate limit conditions (triggered when any value is exceeded)

Requests per second (RPS)

Number of concurrent tasks being processed (concurrency)

wan2.7-image-pro

5

5

wan2.7-image

5

5

wan2.6-image

5

5

wan2.6-t2i

1

5

wan2.5-t2i-preview

5

5

wanx2.0-t2i-turbo

2

2

wanx2.1-t2i-turbo

2

2

wanx2.1-t2i-plus

2

2

wan2.2-t2i-flash

2

2

wan2.2-t2i-plus

2

2

wan2.5-i2i-preview

5

5

wanx2.1-imageedit

2

2

德国(法兰克福)

Model

Service Deployment Scope

Rate limit conditions (triggered when any value is exceeded)

Requests per second (RPS)

Number of concurrent tasks being processed (concurrency)

wan2.6-t2i

Global

1

5

wan2.6-image

Global

5

5

AI试衣OutfitAnyone

华北2(北京)

Model

限流条件(超出任一数值时触发限流)

每秒钟调用次数(RPS)

同时处理中任务数量

aitryon-plus

10

5

aitryon-parsing-v1

10

同步接口无限制

Vidu系列

新加坡

Model限流条件(超出任一数值时触发限流)
每分钟调用次数(RPM)Number of concurrent tasks being processed

vidu/vidu-image_reference2image

300

5

Under a single Bailian API Key, the Vidu reference image generation series models share 5 concurrent slots. That is, the total number of tasks in running status across all models combined must not exceed 5.

Video Generation

HappyHorse Series

新加坡

Model

Service Deployment Scope

Throttling conditions (throttling triggers when any value is exceeded)

Calls per second (RPS)

Number of concurrent tasks being processed (concurrency)

happyhorse-1.1-t2v

International

5

5

happyhorse-1.1-i2v

International

5

5

happyhorse-1.1-r2v

International

5

5

happyhorse-1.0-t2v

International

5

5

happyhorse-1.0-i2v

International

5

5

happyhorse-1.0-r2v

International

5

5

happyhorse-1.0-video-edit

International

5

5

美国(弗吉尼亚)

Model

Service Deployment Scope

Rate Limit Conditions (triggered when any value is exceeded)

Requests per second (RPS)

Concurrent Tasks Being Processed (Concurrency)

happyhorse-1.1-t2v

Global

5

5

happyhorse-1.1-i2v

Global

5

5

happyhorse-1.1-r2v

Global

5

5

happyhorse-1.0-t2v

Global

5

5

happyhorse-1.0-i2v

Global

5

5

happyhorse-1.0-r2v

Global

5

5

happyhorse-1.0-video-edit

Global

5

5

华北2(北京)

Model

Throttling Conditions (triggered when any value is exceeded)

Requests per second (RPS)

Number of tasks being processed simultaneously (concurrency)

happyhorse-1.1-t2v

5

5

happyhorse-1.1-i2v

5

5

happyhorse-1.1-r2v

5

5

happyhorse-1.0-t2v

5

5

happyhorse-1.0-i2v

5

5

happyhorse-1.0-r2v

5

5

happyhorse-1.0-video-edit

5

5

德国(法兰克福)

Model

Service Deployment Scope

Throttling conditions (triggered when any value is exceeded)

Requests per second (RPS)

Number of tasks being processed simultaneously (concurrency)

happyhorse-1.1-t2v

Global

5

5

happyhorse-1.1-i2v

Global

5

5

happyhorse-1.1-r2v

Global

5

5

happyhorse-1.0-t2v

Global

5

5

happyhorse-1.0-i2v

Global

5

5

happyhorse-1.0-r2v

Global

5

5

happyhorse-1.0-video-edit

Global

5

5

日本(东京)

Model

Service Deployment Scope

Throttling conditions (triggered when any value is exceeded)

Calls per second (RPS)

Concurrent tasks being processed (concurrency)

happyhorse-1.0-video-edit

Global

5

5

Wanx Series

新加坡

Model

Service Deployment Scope

Throttling conditions (triggered when any value is exceeded)

Requests per second (RPS)

Concurrent requests in progress

wan3.0-video-prime

International

5

5

wan3.0-video

International

5

5

wan2.7-t2v-2026-06-12

International

5

5

wan2.7-t2v-2026-04-25

International

5

5

wan2.7-t2v

International

5

5

wan2.6-t2v

International

5

5

wan2.5-t2v-preview

International

5

5

wan2.2-t2v-plus

International

2

2

wan2.1-t2v-turbo

International

2

2

wan2.1-t2v-plus

International

2

2

wan2.7-i2v-2026-04-25

International

5

5

wan2.7-i2v

International

5

5

wan2.6-i2v-flash

International

5

5

wan2.6-i2v

International

5

5

wan2.5-i2v-preview

International

5

5

wan2.2-i2v-flash

International

2

2

wan2.1-i2v-plus

International

2

2

wan2.1-i2v-turbo

International

2

2

wan2.2-i2v-plus

International

2

2

wan2.2-kf2v-flash

International

2

2

wan2.1-kf2v-plus

International

1

2

wan2.1-vace-plus

International

2

2

wan2.7-videoedit

International

5

5

wan2.7-r2v

International

5

5

wan2.6-r2v-flash

International

5

5

wan2.6-r2v

International

5

5

wan2.2-animate-move

International

5

1

wan2.2-animate-mix

International

5

1

美国(弗吉尼亚)

Model

Service Deployment Scope

Rate limiting conditions (triggered when any value is exceeded)

Calls per second (RPS)

Number of concurrent tasks being processed (concurrency)

wan3.0-video-prime

Global

5

5

wan3.0-video

Global

5

5

wan2.7-r2v-2026-06-12

Global

5

5

wan2.6-t2v

Global

5

5

wan2.6-i2v

Global

5

5

wan2.6-r2v

Global

5

5

wan2.6-t2v

United States

5

5

wan2.6-i2v

United States

5

5

华北2(北京)

Model

Rate limit conditions (triggered when exceeding any of the values)

Requests per second (RPS)

Number of tasks being processed concurrently (concurrency)

wan3.0-video-prime

5

5

wan3.0-video

5

5

wan2.7-r2v-2026-06-12

5

5

wan2.7-t2v-2026-06-12

5

5

wan2.7-t2v-2026-04-25

5

5

wan2.7-t2v

5

5

wan2.6-t2v

5

5

wan2.5-t2v-preview

5

5

wan2.2-t2v-plus

2

2

wanx2.1-t2v-turbo

2

2

wanx2.1-t2v-plus

2

2

wan2.7-i2v-2026-04-25

5

5

wan2.7-i2v

5

5

wan2.6-i2v-flash

5

5

wan2.6-i2v

5

5

wan2.5-i2v-preview

5

5

wan2.2-i2v-plus

2

2

wanx2.1-i2v-turbo

2

2

wanx2.1-i2v-plus

2

2

wan2.2-kf2v-flash

2

2

wanx2.1-kf2v-plus

2

2

wanx2.1-vace-plus

2

2

wan2.7-videoedit

5

5

wan2.7-r2v

5

5

wan2.6-r2v-flash

5

5

wan2.6-r2v

5

5

wan2.2-s2v-detect

5

Synchronous interface has no limit

wan2.2-s2v

5

1

wan2.2-animate-move

5

1

wan2.2-animate-mix

5

1

德国(法兰克福)

Model

Service Deployment Scope

Throttling conditions (triggered when any value is exceeded)

Calls per second (RPS)

Number of tasks being processed concurrently (concurrency)

wan3.0-video-prime

Global

5

5

wan3.0-video

Global

5

5

wan2.6-t2v

Global

5

5

wan2.6-i2v

Global

5

5

wan2.6-r2v

Global

5

5

日本(东京)

Model

Service Deployment Scope

Rate limit conditions (triggered when any value is exceeded)

Requests per second (RPS)

Number of concurrent tasks being processed (concurrency)

wan3.0-video-prime

Global

5

5

wan3.0-video

Global

5

5

中国香港

Model

Service Deployment Scope

Rate limit conditions (rate limit is triggered when any value is exceeded)

Requests per second (RPS)

Number of concurrent tasks being processed (concurrency)

wan3.0-video-prime

Global

5

5

wan3.0-video

Global

5

5

舞动人像AnimateAnyone

华北2(北京)

Model

Requests per second (RPS)

Number of concurrent tasks being processed

animate-anyone-detect-gen2

5

No limit for synchronous API

animate-anyone-template-gen2

5

1

At any given time, only 1 job is actually running, and other jobs in the queue are waiting.

animate-anyone-gen2

5

1

At any given time, only 1 job is actually running, and other jobs in the queue are waiting.

EMO

华北2(北京)

Model

Requests per second (RPS)

Number of tasks being processed simultaneously

emo-detect-v1

5

No limit for synchronous API

emo-v1

5

1

At any given time, only 1 job is actually running, and other jobs in the queue are waiting.

LivePortrait

华北2(北京)

Model

Number of calls per second(RPS)

Number of tasks being processed concurrently

liveportrait-detect

5

No limit for synchronous interface

liveportrait

5

1

At any given moment, only 1 job is actually running, while other jobs in the queue are in a queued state.

Shengdongrenxiang VideoRetalk

华北2(北京)

Model

Number of calls per second(RPS)

Number of tasks being processed concurrently

videoretalk

1

1

At any given moment, only 1 job is actually running, while other jobs in the queue are in a queued state.

Emoji

华北2(北京)

Model

Requests per second (RPS)

Concurrent processing tasks

emoji-detect-v1

1

No limit on synchronous interfaces

emoji-v1

1

1

At any given moment, only 1 job is actually running, and other jobs in the queue are in a pending state.

Video Style Repainting

华北2(北京)

Model

Requests per second (RPS)

Concurrent processing tasks

video-style-transform

20

2

At the same time, only 1 job is actually running, and the jobs in other queues are in the queueing state.

Vidu Series

新加坡

ModelRate Limit Conditions (triggered when any value is exceeded)
Calls per second (RPS)Number of tasks being processed simultaneously (concurrency)

vidu/viduq3-mix_reference2video

5

5

Under a single Bailian API Key, the models of the Vidu video generation series share 5 concurrent slots. That is, the total number of tasks in running state across all models cannot exceed 5.

vidu/viduq3-ad_reference2video

5

vidu/viduq3-drama_reference2video

5

vidu/viduq2-pro-fast_img2video

5

Audio generation

China (Beijing)

ModelRequests per second (RPS)
qwen-audio-3.1-tts-next3

Music generation

华北2(北京)

Model

Calls per minute (RPM)

fun-music-preview

180

fun-music-v1

180

Speech Dialogue

Real-time Speech Dialogue

新加坡

ModelService Deployment ScopeRate Limit Conditions (Triggered When Any Value Is Exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls Per Minute (RPM)Tokens Consumed Per Minute (TPM)
Includes Input and Output Tokens

qwen-audio-3.1-realtime-plus

International

60

100,000

qwen-audio-3.0-realtime-plus

International

60

100,000

qwen-audio-3.0-realtime-flash

International

60

100,000

华北2(北京)

ModelRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes Input and Output tokens

qwen-audio-3.1-realtime-plus

60

100,000

qwen-audio-3.0-realtime-plus

60

100,000

qwen-audio-3.0-realtime-flash

60

100,000

Speech Synthesis

Qwen-Audio-TTS

新加坡

Model

Service Deployment Scope

Requests per second (RPS)

qwen-audio-3.0-tts-plus

International

3

qwen-audio-3.0-tts-flash

International

3

华北2(北京)

Model

Requests per second (RPS)

qwen-audio-3.0-tts-plus

3

qwen-audio-3.0-tts-flash

3

Qwen-TTS

新加坡

Qwen3-TTS-Instruct-Flash

Model

Service Deployment Scope

Requests per minute (RPM)

qwen3-tts-instruct-flash

International

180

qwen3-tts-instruct-flash-2026-01-26

International

180

Qwen3-TTS-VD

Model

Service Deployment Scope

Calls per minute (RPM)

qwen3-tts-vd-2026-01-26

International

180

Qwen3-TTS-VC

Model

Service Deployment Scope

Calls per minute (RPM)

qwen3-tts-vc-2026-01-22

International

180

Qwen3-TTS-Flash

Model

Service Deployment Scope

Requests Per Minute (RPM)

qwen3-tts-flash

International

180

qwen3-tts-flash-2025-11-27

International

180

qwen3-tts-flash-2025-09-18

International

10

华北2(北京)

Qwen3-TTS-Instruct-Flash

Model

Requests Per Minute (RPM)

qwen3-tts-instruct-flash

180

qwen3-tts-instruct-flash-2026-01-26

180

Qwen3-TTS-VD

Model

Requests Per Minute (RPM)

qwen3-tts-vd-2026-01-26

180

Qwen3-TTS-VC

Model

Requests per minute (RPM)

qwen3-tts-vc-2026-01-22

180

Qwen3-TTS-Flash

Model

Requests per minute (RPM)

qwen3-tts-flash

180

qwen3-tts-flash-2025-11-27

180

qwen3-tts-flash-2025-09-18

10

Qwen-TTS

ModelThrottling conditions (triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may throttle based on RPS (RPM/60) and TPS (TPM/60)
Requests per minute (RPM)Tokens consumed per minute (TPM)
Includes Input and Output tokens

qwen-tts

10

100,000

qwen-tts-latest

qwen-tts-2025-05-22

qwen-tts-2025-04-10

Qwen-TTS-Realtime

新加坡

Qwen3-TTS-Instruct-Flash-Realtime

Model

Service Deployment Scope

每分钟调用次数(RPM)

qwen3-tts-instruct-flash-realtime

International

180

qwen3-tts-instruct-flash-realtime-2026-01-22

International

180

Qwen3-TTS-VD-Realtime

Model

Service Deployment Scope

每分钟调用次数(RPM)

qwen3-tts-vd-realtime-2026-01-15

International

180

qwen3-tts-vd-realtime-2025-12-16

International

Qwen3-TTS-VC-Realtime

Model

Service Deployment Scope

每分钟调用次数(RPM)

qwen3-tts-vc-realtime-2026-01-15

International

180

qwen3-tts-vc-realtime-2025-11-27

International

Qwen3-TTS-Flash-Realtime

Model

Service Deployment Scope

每分钟调用次数(RPM)

qwen3-tts-flash-realtime

International

180

qwen3-tts-flash-realtime-2025-11-27

International

180

qwen3-tts-flash-realtime-2025-09-18

International

10

华北2(北京)

Qwen3-TTS-Instruct-Flash-Realtime

Model

Calls per minute (RPM)

qwen3-tts-instruct-flash-realtime

180

qwen3-tts-instruct-flash-realtime-2026-01-22

180

Qwen3-TTS-VD-Realtime

Model

Calls per minute (RPM)

qwen3-tts-vd-realtime-2026-01-15

180

qwen3-tts-vd-realtime-2025-12-16

Qwen3-TTS-VC-Realtime

Model

Calls per minute (RPM)

qwen3-tts-vc-realtime-2026-01-15

180

qwen3-tts-vc-realtime-2025-11-27

Qwen3-TTS-Flash-Realtime

Model

Calls per minute (RPM)

qwen3-tts-flash-realtime

180

qwen3-tts-flash-realtime-2025-11-27

180

qwen3-tts-flash-realtime-2025-09-18

10

Qwen-TTS-Realtime

ModelThrottling conditions (throttling is triggered when any value is exceeded)
以下为每分钟限流条件,服务可能按 RPS(RPM/60)与 TPS(TPM/60)限制
每分钟调用次数(RPM)每分钟消耗Token数(TPM)
含输入与输出Token

qwen-tts-realtime

10

100,000

qwen-tts-realtime-latest

qwen-tts-realtime-2025-07-15

Qwen-TTS Voice Cloning

新加坡

Model

Service Deployment Scope

每分钟调用次数(RPM)

qwen-voice-enrollment

International

180

华北2(北京)

Model

Calls per minute (RPM)

qwen-voice-enrollment

180

Qwen-TTS Voice Design

新加坡

Model

Service Deployment Scope

Calls per minute (RPM)

qwen-voice-design

International

180

华北2(北京)

Model

Calls per minute (RPM)

qwen-voice-design

180

CosyVoice

新加坡

Model

Service Deployment Scope

Requests per second (RPS)

cosyvoice-v3-plus

International

3

cosyvoice-v3-flash

International

华北2(北京)

Model

Requests per second (RPS)

cosyvoice-v3.5-plus

3

cosyvoice-v3.5-flash

cosyvoice-v3-plus

cosyvoice-v3-flash

cosyvoice-v2

Qwen-Audio-TTS/CosyVoice Voice Cloning/Design

Qwen-Audio-TTS/CosyVoice Voice Cloning/Design share one model and share the rate limit quota.

新加坡

Model

Service Deployment Scope

Requests per second (RPS)

voice-enrollment

International

10

华北2(北京)

Model

Requests per second (RPS)

voice-enrollment

10

Speech Recognition (speech to text) and Translation (speech converted into text in a specified language)

Qwen3-LiveTranslate-Flash

新加坡

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Requests per minute (RPM)Tokens consumed per minute (TPM)
Includes Input and Output Tokens

qwen3-livetranslate-flash

International

100

100,000

qwen3-livetranslate-flash-2025-12-01

International

100

100,000

华北2(北京)

ModelRate Limit Conditions (Triggered When Any Value Is Exceeded)
The following are rate limit conditions per minute. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls Per Minute (RPM)Tokens Per Minute (TPM)
Includes Input and Output Tokens

qwen3-livetranslate-flash

100

100,000

qwen3-livetranslate-flash-2025-12-01

Qwen-LiveTranslate-Flash-Realtime

新加坡

ModelService Deployment ScopeRate limit conditions (rate limit is triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes Input and Output Tokens

qwen3.8-livetranslate-flash-realtime

International

10

100,000

qwen3.5-livetranslate-flash-realtime

International

qwen3.5-livetranslate-flash-realtime-2026-05-19

International

qwen3-livetranslate-flash-realtime

International

qwen3-livetranslate-flash-realtime-2025-09-22

International

华北2(北京)

ModelRate limit conditions (rate limit triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes Input and Output Token

qwen3.8-livetranslate-flash-realtime

10

100,000

qwen3.5-livetranslate-flash-realtime

qwen3.5-livetranslate-flash-realtime-2026-05-19

qwen3-livetranslate-flash-realtime

qwen3-livetranslate-flash-realtime-2025-09-22

Qwen-Audio-3.1-ASR-Flash-Message

Singapore

Model nameDeployment scopeRequests per minute (RPM)
qwen-audio-3.1-asr-flash-messageInternational1200

China (Beijing)

Model nameRequests per minute (RPM)
qwen-audio-3.1-asr-flash-message1200

Qwen-Audio-3.x-ASR-Flash-Streaming

Singapore

Model name

Service deployment scope

Requests per second (RPS)

Requests per minute (RPM)

qwen-audio-3.1-asr-flash-streaming

International

—600

qwen-audio-3.0-asr-flash-streaming

International

20

—

China (Beijing)

Model name

Requests per second (RPS)

Requests per minute (RPM)

qwen-audio-3.1-asr-flash-streaming

—600

qwen-audio-3.0-asr-flash-streaming

20

—

Qwen-Audio-3.x-ASR-Flash-Filetrans

Singapore

Model name

Service deployment scope

Requests per minute (RPM)

qwen-audio-3.1-asr-flash-filetrans

International

600

qwen-audio-3.0-asr-flash-filetrans

International

600

China (Beijing)

Model name

Requests per minute (RPM)

qwen-audio-3.1-asr-flash-filetrans

600

qwen-audio-3.0-asr-flash-filetrans

600

Qwen-Audio-3.x-ASR-Flash

Singapore

Model name

Service deployment scope

Requests per minute (RPM)

qwen-audio-3.1-asr-flash

International

600

qwen-audio-3.0-asr-flash

International

600

China (Beijing)

Model name

Requests per minute (RPM)

qwen-audio-3.1-asr-flash

600

qwen-audio-3.0-asr-flash

600

Qwen-ASR

新加坡

Qwen3-ASR-Flash-Filetrans

Model

Service Deployment Scope

Calls per minute (RPM)

qwen3-asr-flash-filetrans

International

100

qwen3-asr-flash-filetrans-2025-11-17

International

Qwen3-ASR-Flash

Model

Service Deployment Scope

Calls per minute (RPM)

qwen3-asr-flash

International

100

qwen3-asr-flash-2026-02-10

International

qwen3-asr-flash-2025-09-08

International

美国(弗吉尼亚)

Model

Service Deployment Scope

Calls per minute (RPM)

qwen3-asr-flash

United States

100

qwen3-asr-flash-2025-09-08

United States

华北2(北京)

Qwen3-ASR-Flash-Filetrans

Model

每分钟调用次数(RPM)

qwen3-asr-flash-filetrans

100

qwen3-asr-flash-filetrans-2025-11-17

Qwen3-ASR-Flash

Model

每分钟调用次数(RPM)

qwen3-asr-flash

100

qwen3-asr-flash-2026-02-10

qwen3-asr-flash-2025-09-08

Qwen-ASR-Realtime

新加坡

Model

Service Deployment Scope

每秒钟调用次数(RPS)

qwen3-asr-flash-realtime

International

20

qwen3-asr-flash-realtime-2026-02-10

International

qwen3-asr-flash-realtime-2025-10-27

International

华北2(北京)

Model

Requests per second (RPS)

qwen3-asr-flash-realtime

20

qwen3-asr-flash-realtime-2026-02-10

qwen3-asr-flash-realtime-2025-10-27

Paraformer

华北2(北京)

Model

Requests per second (RPS)

paraformer-realtime-v2

20

paraformer-realtime-8k-v2

Model

Requests per minute (RPM)

paraformer-v2

1,200

Model

Requests per second (RPS)

Concurrent processing tasks (concurrency)

paraformer-8k-v2

20

100

Fun-ASR

新加坡

Model

Service Deployment Scope

每分钟调用次数(RPM)

fun-asr

International

600

fun-asr-2025-11-07

International

600

fun-asr-2025-08-25

International

600

fun-asr-mtl

International

100

fun-asr-mtl-2025-08-25

International

100

fun-asr-flash-2026-06-15

International

600

华北2(北京)

Model

Requests Per Minute (RPM)

fun-asr

600

fun-asr-2025-11-07

fun-asr-2025-08-25

fun-asr-mtl

fun-asr-mtl-2025-08-25

fun-asr-flash-2026-06-15

Fun-ASR-Realtime

新加坡

Model

Service Deployment Scope

Requests Per Second (RPS)

fun-asr-realtime

International

20

fun-asr-realtime-2025-11-07

International

华北2(北京)

Model

Requests Per Second (RPS)

fun-asr-realtime

20

fun-asr-realtime-2026-02-28

fun-asr-realtime-2025-11-07

fun-asr-realtime-2025-09-15

fun-asr-flash-8k-realtime

fun-asr-flash-8k-realtime-2026-01-28

Text Embedding

新加坡

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

qwen3.7-text-embedding

International

24,000

1,000,000

text-embedding-v4

International

1,800

1,000,000

text-embedding-v3

International

6,000

24,000,000

华北2(北京)

ModelRate limit conditions (triggered when any value is exceeded)
Requests per second (RPS)Tokens consumed per minute (TPM)
Includes input and Output tokens

text-embedding-v4

UseBatch APIWhen calling the service, it is not subject to rate limit.

30

1,200,000

中国香港

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60)
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and Output Token

text-embedding-v4

Hong Kong (China)

1,800

1,200,000

Multimodal Embedding

新加坡

ModelService Deployment ScopeRate limit conditions
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Input only Token

tongyi-embedding-vision-plus

International

600

200,000

tongyi-embedding-vision-flash

International

600

200,000

华北2(北京)

ModelRate limit conditions
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Input only Token

qwen3-vl-embedding

2,400

1,200,000

multimodal-embedding-v1

120

1,000,000

Reranking Model

新加坡

ModelService Deployment ScopeRate Limit Conditions
The following are per-minute rate limit conditions. The service may be limited by RPS (RPM/60) and TPS (TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Input Tokens Only

qwen3-rerank

International

5,400

5,000,000,000

华北2(北京)

ModelRate Limit Conditions
The following are rate limit conditions per minute. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Requests Per Minute (RPM)Tokens Per Minute (TPM)
Only Input Tokens

qwen3-rerank

5,400

5,000,000,000

qwen3-vl-rerank

600

9,000,000

gte-rerank-v2

5,040

4,980,000,000

Industry

Intent Understanding

华北2(北京)

ModelRate Limit Conditions (triggered when any value is exceeded)
The following are rate limit conditions per minute. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Requests per minute (RPM)Tokens consumed per minute (TPM)
Includes Input and Output Tokens

tongyi-intent-detect-v3

1,200

1,000,000

Role-Playing

新加坡

ModelService Deployment ScopeThrottling conditions (triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may throttle based on RPS (RPM/60) and TPS (TPM/60).
Requests per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Token

qwen-plus-character

International

120

500,000

qwen-flash-character

International

120

500,000

qwen-plus-character-ja

International

120

500,000

美国(弗吉尼亚)

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

qwen-plus-character

Global

120

500,000

qwen-flash-character

Global

120

500,000

华北2(北京)

ModelRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes input and output Tokens

qwen-plus-character

120

500,000

qwen-flash-character

120

500,000

Decision model

Singapore

Model nameRate limits (triggered if any value is exceeded)
The limits below are specified per minute. The service may enforce limits based on records per second (RPS) (RPM/60) and Transactions Per Second (TPS) (TPM/60).
Requests Per Minute (RPM)Tokens Per Minute (TPM)
Includes input and output tokens

decision-model-preview

1,200

2,000,000

China (Beijing)

Model nameRate limits (triggered if any value is exceeded)
The limits below are specified per minute. The service may enforce limits based on records per second (RPS) (RPM/60) and Transactions Per Second (TPS) (TPM/60).
Requests Per Minute (RPM)Tokens Per Minute (TPM)
Includes input and output tokens

decision-model-preview

1,200

2,000,000

Offline Models

For more information, see Model Offline Mechanism 。

2026年1月30日下线

CategoryModelRate limit conditions (triggered when any value is exceeded)
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and Output Tokens

Qwen Plus

qwen-plus-2024-11-27

0

0

qwen-plus-2024-11-25

qwen-plus-2024-09-19

qwen-plus-2024-08-06

Qwen Turbo

qwen-turbo-2024-09-19

Qwen-VL

qwen-vl-max-2024-10-30

qwen-vl-max-2024-08-09

qwen-vl-plus-2024-08-09

2025年8月20日下线

CategoryModelRate Limit Conditions (triggered when any value is exceeded)
Calls Per Minute (RPM)Tokens Per Minute (TPM)
Including Input and Output Tokens

Text Generation-Qwen

qwen2-72b-instruct

0

0

qwen2-57b-a14b-instruct

qwen2-7b-instruct

qwen1.5-110b-chat

qwen1.5-72b-chat

qwen1.5-32b-chat

qwen1.5-14b-chat

qwen1.5-7b-chat