All Products
Search
Document Center

Alibaba Cloud Model Studio:Rate limiting

Last Updated:Oct 03, 2026

Alibaba Cloud Model Studio applies rate limiting to model calls at the Alibaba Cloud account level, aggregating usage across all RAM users, workspaces, and API keys under the account. Requests are rejected when the limit is exceeded and typically recover automatically within one minute.

Rate Limiting Rules

  • Account-level Rate Limiting: Rate limiting is calculated by primary account dimension. The call volume of all RAM sub-accounts, workspaces, and API Keys under the account is calculated cumulatively.
  • Model-independent Rate Limiting: Different models have independent rate limit quotas. For more information, see the table below.

FAQ

Why is rate limiting triggered?

Based on the error message, determine which type of rate limit was triggered:

  • Requests rate limit exceeded Or You exceeded your current requests list: Requests per minute (RPM) rate limit triggered.
  • Allocated quota exceeded Or You exceeded your current quota: Tokens per minute (TPM) rate limit triggered.
  • Request rate increased too quickly: Request frequency spiked in a short period, triggering system stability protection—this is triggered even if the total number of calls does not reach the RPM or TPM upper limit.
  • For other errors, see Error Code to confirm the cause.

In addition to RPM and TPM, the rate limit policy may be executed based on per-second RPS (RPM/60) and TPS (TPM/60). Even if the total number of calls per minute does not exceed the limit, a burst of requests in a short period may also trigger rate limiting.

How to view model call volume?

After a model call occursminute-levelyou can view it in the monitoring chart. Log in to the Monitoring page (SingaporeOrBeijing), in the left navigation pane, selectO&M Management > MonitoringGo to the Monitoring overview page, set query conditions (for example, select a time range, workspace, etc.), then in theModelsarea, find the target Model and clickView Detailsto view the invocation statistics of the Model. For details, seeMonitoringDocs.

Monitoring data is updated at the Minute(s)-level and is for reference only; it is not used as a billing basis.

How long does it take to recover after Rate Limiting is triggered?

It usually recovers within one Minute(s). If other Errors occur, see Error Code to handle the issue.

Does model response speed relate to payment?

Model response speed is unrelated to whether you pay. Free quota and paid calls use the same Model service infrastructure, and the generation speed of the same Model is identical. Free quota only controls whether a request is accepted—when the quota is exhausted, a 403 Error is returned (ActivateFree Quota Onlyafter), and it does not reduce the generation speed of already accepted requests.

What factors affect model response speed?

  • Model Type: Lightweight models (such as qwen-flash) generate faster than large models (such as qwen-max).
  • Output length: The more Output Tokens, the longer the total time.
  • Server load: Slight fluctuations may occur during peak periods.

Does rate limiting reduce the generation speed of accepted requests?

No. Rate limiting (RPM/TPM) only causes requests exceeding the limit to be rejected (Return 429 Error). When the free quota is exhausted (Activate Free Quota Only) Return 403 Error, which also does not affect the speed of accepted requests. 429 is rate limiting, 403 is quota exhaustion, both are rejection mechanisms.

How to avoid rate limiting?

  1. Choose models with higher rate limits: The stable version Or latest version has more lenient rate limiting than dated snapshot versions.

  2. Optimize invocation strategy
    • Reduce invocation frequency: Received Requests rate limit exceeded Or You exceeded your current requests list, reduce the API invocation frequency.
    • Reduce Token consumption: Received Allocated quota exceeded Or You exceeded your current quota, shorten the input or limit the Output length.
    • Smooth request rate: Receive Request rate increased too quickly, use uniform-rate scheduling, exponential backoff, or request queues to evenly distribute requests and avoid instantaneous peaks.
  3. Add backup model

    After triggering rate limit, switch to a backup model to continue generation, which can reduce the failure probability and improve throughput. The following code calls qwen-plus-2025-07-28 after triggering rate limit, automatically switches to qwen-plus-2025-07-14 to retry. You can also choose models from different series (such as qwen-flash) as backup models to further reduce the risk of both primary and backup models being rate-limited at the same time.

    Example code

    import os
    import asyncio
    from openai import AsyncOpenAI, APIStatusError
    
    # 配置
    API_KEY = os.getenv("DASHSCOPE_API_KEY")
    # 主用模型
    MODEL = "qwen-plus-2025-07-28"
    # 备选模型
    BACKUP_MODEL = "qwen-plus-2025-07-14"
    # 测试问题
    QUESTION = "你是谁?"
    # 并发设置
    NUM_REQUESTS = 10
    
    client = AsyncOpenAI(
        api_key=API_KEY,
        # 调用时请将WorkspaceId替换为真实的业务空间ID
        base_url="https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1"
    )
    
    async def send_request(model):
        """发送单个请求"""
        try:
            await client.chat.completions.create(
                model=model,
                messages=[{"role": "user", "content": QUESTION}]
            )
            return True
        except APIStatusError as e:
            if e.status_code == 429:
                print(f"[限流触发] 模型 {model}")
                return False
            raise
        except Exception as e:
            print(f"[请求失败] 模型 {model},错误:{e}")
            return False
    
    async def task(i):
        # 尝试主模型
        if await send_request(MODEL):
            return True
        # 限流时尝试备用模型
        return await send_request(BACKUP_MODEL)
    
    async def main():
        results = await asyncio.gather(*(task(i) for i in range(NUM_REQUESTS)))
        print(f"成功请求: {sum(results)}, 失败请求: {len(results) - sum(results)}")
    
    if __name__ == "__main__":
        asyncio.run(main())
    
  4. Split tasks: Long conversations or large documents quickly consume a large number of Tokens. Split large batch tasks into smaller batches and submit them in time slots.

  5. Batches: When real-time response is not needed, useBatches (Batch API). Batch requests are not subject to real-time rate limiting constraints, but you need to consider queuing and processing time.

How to control Token Usage Or cost expenditure?

Rate limiting only constrains the call rate per unit time and does not limit cumulative Usage. If you need to control Token Usage Or cost expenditure, you can manage it through the following methods:

  • Set consumption limits and fee alerts: InBills and feescard, setfee alerts, enable monthly consumption limits and Configure threshold notifications. You will be alerted when the threshold is reached to avoid overspending. For more information, see Billing Query and Cost Management.
  • Enable Free Quota Only: For models that support the free quota, you can enableFree Quota Only, after the free quota is exhausted, calling automatically stops to avoid additional costs. For more information, see New User Free Quota.
  • Monitor Model Usage: Regularly check the Token usage of each model, and detect abnormal growth in a timely manner. See above How to view model usage?.

Will you still be rate-limited after topping up?

Top Up does not change the model's default RPM and TPM rate limit thresholds. Rate limits are configured at the Alibaba Cloud primary account level and are independent of billing; Top Up (Pay-as-you-go) only ensures the account is not suspended due to overdue payment and does not increase the rate limit.

If you need higher rate limit quotas, contact your business manager to apply.

Text Generation-Qwen

Qwen Language Model

Singapore

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions; the service may also limit by RPS (RPM/60) and TPS (TPM/60)
Calls per minute (RPM)Tokens consumed per minute (TPM)
Including Input and Output Token

qwen3.8-max

International

Dynamic Rate Limiting

qwen3.8-max-0902

International

Dynamic Rate Limiting

qwen3.8-flash

International

Dynamic Rate Limiting

qwen3.7-max

International

600

1,000,000

qwen3.7-max-2026-06-08

International

60

1,000,000

qwen3.7-max-2026-05-20

International

60

1,000,000

qwen3.7-max-preview

International

600

1,000,000

qwen3.7-max-2026-05-17

International

600

1,000,000

qwen3.6-max-preview

International

600

1,000,000

qwen3-max

International

600

1,000,000

qwen3-max-2026-01-23

International

600

1,000,000

qwen3-max-2025-09-23

International

60

100,000

qwen3-max-preview

International

600

1,000,000

qwen-max

UseBatch APIWhen calling the service, it is not subject to rate limiting.

International

600

1,000,000

qwen3.7-plus

International

15,000

5,000,000

qwen3.7-plus-2026-05-26

International

60

1,000,000

qwen3.6-plus

International

15,000

5,000,000

qwen3.6-plus-2026-04-02

International

60

1,000,000

qwen3.7-flash

International

15,000

5,000,000

qwen3.7-flash-2026-07-15

International

15,000

5,000,000

qwen3.6-flash

International

15,000

5,000,000

qwen3.6-flash-2026-04-16

International

60

1,000,000

qwen3.5-plus

International

15,000

5,000,000

qwen3.5-plus-2026-04-20

International

600

1,000,000

qwen3.5-plus-2026-02-15

International

60

1,000,000

qwen-plus

UseBatch APIWhen calling the service, it is not subject to rate limiting.

International

600

1,000,000

qwen-plus-latest

International

600

1,000,000

qwen-plus-2025-12-01

International

120

1,000,000

qwen-plus-2025-09-11

International

120

1,000,000

qwen-plus-2025-07-28

International

60

100,000

qwen-plus-2025-07-14

(qwen-plus-0714)

International

60

100,000

qwen-plus-2025-04-28

(qwen-plus-0428)

International

60

1,000,000

qwen-plus-2025-01-25

(qwen-plus-0125)

International

60

100,000

qwen3.5-flash

International

15,000

5,000,000

qwen3.5-flash-2026-02-23

International

60

1,000,000

qwen-flash

UseBatch APIWhen calling the service, it is not subject to rate limiting.

International

600

5,000,000

qwen-flash-2025-07-28

International

600

5,000,000

qwq-plus

International

60

100,000

qwen-turbo

UseBatch APIWhen calling the service, it is not subject to rate limiting.

International

600

5,000,000

US (Virginia)

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may impose limits based on RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output tokens

qwen3.8-max

Global

30,000

5,000,000

qwen3.8-max-0902

Global

30,000

150,000

qwen3.8-flash

Global

30,000

5,000,000

qwen3.7-max

Global

30,000

5,000,000

qwen3.7-max

United States

600

1,000,000

qwen3.7-max-2026-06-08

Global

600

1,000,000

qwen3.7-max-2026-05-20

Global

600

1,000,000

qwen3-max

Global

600

1,000,000

qwen3-max-preview

Global

600

1,000,000

qwen3-max-2025-09-23

Global

60

100,000

qwen3.7-plus

Global

30,000

5,000,000

qwen3.7-plus

United States

15,000

5,000,000

qwen3.7-plus-2026-05-26

Global

600

1,000,000

qwen3.6-plus

Global

30,000

5,000,000

qwen3.6-plus-2026-04-02

Global

600

1,000,000

qwen3.7-flash

Global

15,000

5,000,000

qwen3.7-flash-2026-07-15

Global

60

1,000,000

qwen3.6-flash

Global

15,000

5,000,000

qwen3.6-flash-2026-04-16

Global

60

1,000,000

qwen3.6-flash

United States

15,000

5,000,000

qwen3.5-plus

Global

30,000

5,000,000

qwen3.5-plus-2026-02-15

Global

600

1,000,000

qwen-plus

Global

15,000

5,000,000

qwen-plus

United States

600

1,000,000

qwen-plus-2025-12-01

Global

60

1,000,000

qwen-plus-2025-09-11

Global

60

1,000,000

qwen-plus-2025-07-28

Global

60

1,000,000

qwen-plus-2025-12-01

United States

60

1,000,000

qwen3.5-flash

Global

30,000

10,000,000

qwen3.5-flash-2026-02-23

Global

600

1,000,000

qwen-flash

Global

15,000

10,000,000

qwen-flash

United States

30000

10,000,000

qwen-flash-2025-07-28

Global

60

1,000,000

qwen-flash-2025-07-28

United States

600

5,000,000

North China 2 (Beijing)

ModelRate limit conditions (triggered when any value is exceeded)
Below are the rate limit conditions per minute. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

qwen3.8-max

UseBatch APIWhen calling the service, it is not subject to rate limits.

Dynamic rate limiting

qwen3.8-max-0902

Dynamic rate limiting

qwen3.8-flash

UseBatch APIWhen calling the service, it is not subject to rate limiting.

Dynamic rate limiting

qwen3.7-max

UseBatch APIWhen calling the service, it is not subject to rate limiting.

30,000

5,000,000

qwen3.7-max-2026-06-08

600

1,000,000

qwen3.7-max-2026-05-20

600

1,000,000

qwen3.6-max-preview

600

1,000,000

qwen3-max

UseBatch APIWhen calling the service, it is not subject to rate limiting.

30,000

5,000,000

qwen3-max-2026-01-23

600

1,000,000

qwen3-max-2025-09-23

60

100,000

qwen3-max-preview

600

1,000,000

qwen-max

UseBatch APIWhen calling the service, it is not subject to rate limiting.

1,200

1,000,000

qwen3.7-plus

30,000

5,000,000

qwen3.7-plus-2026-05-26

600

1,000,000

qwen3.6-plus

UseBatch APIWhen calling the service, it is not subject to rate limiting.

30,000

5,000,000

qwen3.6-plus-2026-04-02

600

1,000,000

qwen3.7-flash

UseBatch APIWhen calling the service, it is not subject to rate limiting.

30,000

5,000,000

qwen3.7-flash-2026-07-15

30,000

5,000,000

qwen3.6-flash

UseBatch APIWhen calling the service, it is not subject to rate limiting.

30,000

10,000,000

qwen3.6-flash-2026-04-16

600

1,000,000

qwen3.5-plus

UseBatch APIWhen calling the service, it is not subject to rate limiting.

30,000

5,000,000

qwen3.5-plus-2026-04-20

600

1,000,000

qwen3.5-plus-2026-02-15

600

1,000,000

qwen-plus

UseBatch APIWhen calling the service, it is not subject to rate limiting.

30,000

5,000,000

qwen-plus-latest

UseBatch APIWhen calling the service, it is not subject to rate limiting.

15,000

1,200,000

qwen-plus-2025-12-01

120

1,000,000

qwen-plus-2025-09-11

60

1,000,000

qwen-plus-2025-07-28

(qwen-plus-0728)

60

1,000,000

qwen-plus-2025-07-14

(qwen-plus-0714)

60

100,000

qwen-plus-2025-04-28

(qwen-plus-0428)

60

1,000,000

qwen-plus-2025-01-25

(qwen-plus-0125)

60

150,000

qwen-plus-2025-01-12

(qwen-plus-0112)

60

150,000

qwen-plus-2024-12-20

(qwen-plus-1220)

60

150,000

qwen3.5-flash

UseBatch APIWhen calling the service, it is not subject to rate limiting.

30,000

10,000,000

qwen3.5-flash-2026-02-23

600

1,000,000

qwen-flash

UseBatch APIWhen calling the service, it is not subject to rate limiting.

30,000

10,000,000

qwen-flash-2025-07-28

60

1,000,000

qwq-plus

UseBatch APIWhen calling the service, it is not subject to rate limiting.

600

1,000,000

qwen-turbo

1,200

5,000,000

qwen-long-latest

When calling the service via the Batch API, it is not subject to rate limiting.

1,200

60,000

qwen-long-2025-01-25

(qwen-long-0125)

3

7,500

Germany (Frankfurt)

ModelService Deployment ScopeRate limiting conditions (triggered when any value is exceeded)
The following are per-minute rate limiting conditions. The service may also limit by RPS (RPM/60) and TPS (TPM/60)
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and Output tokens

qwen3.8-max

Global

30,000

5,000,000

qwen3.8-max-0902

Global

30,000

150,000

qwen3.8-flash

Global

30,000

5,000,000

qwen3.7-max

Global

30,000

5,000,000

qwen3.7-max-2026-06-08

Global

600

1,000,000

qwen3.7-max-2026-05-20

Global

600

1,000,000

qwen3-max

Global

600

1,000,000

qwen3-max

European Union

600

1,000,000

qwen3-max-preview

Global

600

1,000,000

qwen3-max-2026-01-23

European Union

60

100,000

qwen3-max-2025-09-23

Global

60

100,000

qwen3.7-plus

Global

30,000

5,000,000

qwen3.7-plus-2026-05-26

Global

600

1,000,000

qwen3.6-plus

Global

30,000

5,000,000

qwen3.6-plus-2026-04-02

Global

600

1,000,000

qwen3.7-flash

Global

15,000

5,000,000

qwen3.7-flash-2026-07-15

Global

15,000

5,000,000

qwen3.6-flash

Global

15,000

5,000,000

qwen3.6-flash-2026-04-16

Global

60

1,000,000

qwen3.5-plus

Global

30,000

5,000,000

qwen3.5-plus-2026-02-15

Global

600

1,000,000

qwen-plus

Global

600

1,000,000

qwen-plus

European Union

600

1,000,000

qwen-plus-2025-12-01

Global

120

1,000,000

qwen-plus-2025-12-01

European Union

120

1,000,000

qwen-plus-2025-09-11

Global

60

1,000,000

qwen-plus-2025-07-28

Global

60

1,000,000

qwen3.5-flash

Global

35,000

10,000,000

qwen3.5-flash

European Union

35,000

10,000,000

qwen3.5-flash-2026-02-23

Global

600

1,000,000

qwen3.5-flash-2026-02-23

European Union

600

1,000,000

qwen-flash

Global

30,000

10,000,000

qwen-flash-2025-07-28

Global

60

1,000,000

Hong Kong (China)

ModelService Deployment ScopeRate limiting conditions (triggered when any value is exceeded)
The following are per-minute rate limiting conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

qwen3.8-max

Hong Kong (China)

600

1,000,000

qwen3.8-max-0902

Hong Kong (China)

600

150,000

qwen3.8-flash

Global

30,000

5,000,000

qwen3-max

Hong Kong (China)

600

1,000,000

qwen3-max-2026-01-23

Hong Kong (China)

60

100,000

qwen3.6-plus

Global

30,000

5,000,000

qwen3.7-flash

Global

15,000

5,000,000

qwen3.7-flash-2026-07-15

Global

15,000

5,000,000

qwen3.6-flash

Global

30,000

10,000,000

qwen-plus

Hong Kong (China)

600

1,000,000

qwen-plus-2025-12-01

Hong Kong (China)

60

100,000

qwen3.5-flash

Hong Kong (China)

30,000

10,000,000

qwen3.5-flash-2026-02-23

Hong Kong (China)

600

1,000,000

Japan (Tokyo)

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

qwen3.8-max

Global

30,000

5,000,000

qwen3.8-max-0902

Global

30,000

150,000

qwen3.8-flash

Global

30,000

5,000,000

qwen3.7-max

Global

600

1,000,000

qwen3.7-max-2026-05-20

Global

60

1,000,000

qwen3.7-plus

Global

15,000

5,000,000

qwen3.7-plus-2026-05-26

Global

60

1,000,000

qwen3.7-plus

Japan

15,000

5,000,000

qwen3.7-plus-2026-05-26

Japan

60

1,000,000

qwen3.6-plus

Global

15,000

5,000,000

qwen3.6-plus-2026-04-02

Global

60

1,000,000

qwen3.7-flash

Global

15,000

5,000,000

qwen3.7-flash-2026-07-15

Global

15,000

5,000,000

qwen3.6-flash

Global

15,000

10,000,000

qwen3.6-flash-2026-04-16

Global

60

1,000,000

Qwen-VL (Visual Understanding/Image-to-Text)

Singapore

ModelService Deployment ScopeThrottling Conditions (Triggered When Any Value Is Exceeded)
The following are per-minute throttling conditions. The service may throttle by RPS (RPM/60) and TPS (TPM/60).
Calls Per Minute (RPM)Tokens Consumed Per Minute (TPM)
Includes input and output Tokens

qwen3-vl-plus

International

1,200

1,000,000

qwen3-vl-plus-2025-12-19

International

60

100,000

qwen3-vl-plus-2025-09-23

International

120

1,000,000

qwen3-vl-flash

International

1,200

1,000,000

qwen3-vl-flash-2026-01-22

International

60

100,000

qwen3-vl-flash-2025-10-15

International

120

1,000,000

qwen-vl-max

International

1,200

1,000,000

qwen-vl-plus

International

1,200

1,000,000

qvq-max

International

60

100,000

US (Virginia)

ModelService Deployment ScopeThrottling conditions (throttling is triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may throttle based on RPS (RPM/60) and TPS (TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes input and Output Tokens

qwen3-vl-plus

Global

6,000

10,000,000

qwen3-vl-plus-2025-09-23

Global

60

100,000

qwen3-vl-flash

Global

1,200

1,000,000

qwen3-vl-flash

United States

1,200

1,000,000

qwen3-vl-flash-2025-10-15

Global

60

100,000

qwen3-vl-flash-2026-01-22

United States

120

1,000,000

qwen3-vl-flash-2025-10-15

United States

120

1,000,000

North China 2 (Beijing)

ModelThrottling conditions (triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes Input and Output Tokens

qwen3-vl-plus

UseBatch APIWhen calling the service, it is not subject to rate limiting.

3,000

5,000,000

qwen3-vl-plus-2025-12-19

60

100,000

qwen3-vl-plus-2025-09-23

60

100,000

qwen3-vl-flash

UseBatch APIWhen calling the service, it is not subject to rate limiting.

3,000

5,000,000

qwen3-vl-flash-2026-01-22

60

100,000

qwen3-vl-flash-2025-10-15

60

100,000

qwen-vl-max

UseBatch APIWhen calling the service, it is not subject to rate limiting.

1,200

1,000,000

qwen-vl-plus

UseBatch APIWhen calling the service, it is not subject to rate limiting.

1,200

1,000,000

qvq-max

60

100,000

qvq-plus

60

100,000

Germany (Frankfurt)

ModelService Deployment ScopeRate limiting conditions (triggered when any value is exceeded)
The following are per-minute rate limiting conditions. The service may also limit by RPS (RPM/60) and TPS (TPM/60)
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output tokens

qwen3-vl-plus

Global

2,000

1,000,000

qwen3-vl-plus

EU

2,000

1,000,000

qwen3-vl-plus-2025-09-23

Global

60

100,000

qwen3-vl-flash

Global

12,000

10,000,000

qwen3-vl-flash

EU

12,000

10,000,000

qwen3-vl-flash-2026-01-22

European Union

60

100,000

qwen3-vl-flash-2025-10-15

Global

600

1,000,000

qwen3-vl-flash-2025-10-15

European Union

600

1,000,000

Hong Kong (China)

ModelService Deployment ScopeRate limiting conditions (triggered when any threshold is exceeded)
The following are per-minute rate limiting conditions. The service may enforce limits based on RPS (RPM/60) and TPS (TPM/60).
Requests Per Minute (RPM)Tokens Per Minute (TPM)
Includes Input and Output Tokens

qwen3-vl-plus

Hong Kong (China)

6,000

1,000,000

qwen3-vl-plus-2025-12-19

Hong Kong (China)

6,000

1,000,000

Qwen Omni (Omni-modality)

Singapore

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60)
Calls per minute (RPM)Tokens consumed per minute (TPM)
Including input and output Tokens

qwen3.8-omni-flash

International

30000

Dynamic Rate Limiting

qwen3.5-omni-flash

International

60

100,000

qwen3.5-omni-flash-2026-03-15

International

60

100,000

qwen3.5-omni-plus

International

60

100,000

qwen3.5-omni-plus-2026-03-15

International

60

100,000

qwen3-omni-flash

International

60

100,000

qwen3-omni-flash-2025-12-01

International

60

100,000

qwen3-omni-flash-2025-09-15

International

60

100,000

qwen-omni-turbo

International

60

100,000

qwen-omni-turbo-latest

International

60

100,000

qwen-omni-turbo-2025-03-26

International

60

100,000

North China 2 (Beijing)

ModelThrottling Conditions (triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may throttle by RPS (RPM/60) and TPS (TPM/60).
Calls per Minute (RPM)Tokens Consumed per Minute (TPM)
Includes Input and Output Tokens

qwen3.8-omni-flash

30000

Dynamic Throttling

qwen3.5-omni-flash

60

100,000

qwen3.5-omni-flash-2026-03-15

60

100,000

qwen3.5-omni-plus

60

100,000

qwen3.5-omni-plus-2026-03-15

60

100,000

qwen3-omni-flash

60

100,000

qwen3-omni-flash-2025-12-01

60

100,000

qwen3-omni-flash-2025-09-15

60

100,000

qwen-omni-turbo

60

100,000

qwen-omni-turbo-latest

60

100,000

qwen-omni-turbo-2025-03-26

(qwen-omni-turbo-0326)

60

100,000

qwen-omni-turbo-2025-01-19

(qwen-omni-turbo-0119)

60

100,000

Hong Kong (China)

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
Rate limit conditions per minute, the service may limit by RPS (RPM/60) and TPS (TPM/60)
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and Output Tokens

qwen3.8-omni-flash

Global

30000

2000000

Japan (Tokyo)

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60)
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

qwen3.8-omni-flash

Global

30000

2000000

Germany (Frankfurt)

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Including input and output Tokens

qwen3.8-omni-flash

Global

30000

2000000

US (Virginia)

ModelService Deployment ScopeThrottling conditions (triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Including input and output Tokens

qwen3.8-omni-flash

Global

30000

2000000

Qwen-Omni-Realtime

Singapore

ModelService Deployment ScopeRate limiting conditions (triggered when any value is exceeded)
The following are per-minute rate limiting conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output tokens

qwen3.8-omni-flash-realtime

International

60

2,000,000

qwen3.5-omni-plus-realtime

International

60

100,000

qwen3.5-omni-plus-realtime-2026-03-15

International

60

100,000

qwen3.5-omni-flash-realtime

International

60

100,000

qwen3.5-omni-flash-realtime-2026-03-15

International

60

100,000

qwen3-omni-flash-realtime

International

60

100,000

qwen3-omni-flash-realtime-2025-12-01

International

60

100,000

qwen3-omni-flash-realtime-2025-09-15

International

60

100,000

qwen-omni-turbo-realtime

International

60

10,000

qwen-omni-turbo-realtime-latest

International

60

10,000

qwen-omni-turbo-realtime-2025-05-08

International

60

10,000

North China 2 (Beijing)

ModelRate limiting conditions (triggered when any value is exceeded)
The following are per-minute rate limiting conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60)
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output tokens

qwen3.8-omni-flash-realtime

60

2,000,000

qwen3.5-omni-plus-realtime

60

100,000

qwen3.5-omni-plus-realtime-2026-03-15

60

100,000

qwen3.5-omni-flash-realtime

60

100,000

qwen3.5-omni-flash-realtime-2026-03-15

60

100,000

qwen3-omni-flash-realtime

60

100,000

qwen3-omni-flash-realtime-2025-12-01

60

100,000

qwen3-omni-flash-realtime-2025-09-15

60

100,000

qwen-omni-turbo-realtime

60

100,000

qwen-omni-turbo-realtime-latest

60

100,000

qwen-omni-turbo-realtime-2025-05-08

60

100,000

Qwen-OCR (Text Extraction)

Singapore

ModelService Deployment ScopeRate limiting conditions (triggered when any value is exceeded)
The following are per-minute rate limiting conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60)
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

qwen-vl-ocr

International

600

6,000,000

qwen-vl-ocr-2025-11-20

International

1,200

6,000,000

US (Virginia)

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Token count consumed per minute (TPM)
Includes input and output Tokens

qwen-vl-ocr

Global

600

6,000,000

qwen-vl-ocr-2025-11-20

Global

1,200

6,000,000

North China 2 (Beijing)

ModelRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60)
Requests per minute (RPM)Token count consumed per minute (TPM)
Includes input and output Tokens

qwen3.5-ocr

6,000

30,000,000

qwen-vl-ocr

UseBatch APIWhen calling the service, it is not subject to rate limits.

600

6,000,000

qwen-vl-ocr-latest

6,000

30,000,000

qwen-vl-ocr-2025-11-20

6,000

30,000,000

qwen-vl-ocr-2025-04-13

600

6,000,000

qwen-vl-ocr-2024-10-28

600

6,000,000

Germany (Frankfurt)

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes input and output tokens

qwen-vl-ocr

Global

600

6,000,000

qwen-vl-ocr-2025-11-20

Global

1,200

6,000,000

Qwen Math Model

North China 2 (Beijing)

ModelRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output tokens

qwen-math-plus

1,200

1,000,000

qwen-math-plus-latest

1,200

1,000,000

qwen-math-plus-2024-09-19

(qwen-math-plus-0919)

60

100,000

qwen-math-plus-2024-08-16

(qwen-math-plus-0816)

10

20,000

qwen-math-turbo

1200

1,000,000

Qwen-Coder

Singapore

ModelService Deployment ScopeThrottling conditions (throttling is triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may throttle based on RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

qwen3-coder-plus

International

2,400

2,000,000

qwen3-coder-plus-2025-09-23

International

600

1,000,000

qwen3-coder-plus-2025-07-22

International

60

1,000,000

qwen3-coder-flash

International

600

5,000,000

qwen3-coder-flash-2025-07-28

International

600

5,000,000

US (Virginia)

ModelService Deployment ScopeRate limiting conditions (triggered when any value is exceeded)
The following are per-minute rate limiting conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output tokens

qwen3-coder-plus

Global

2,400

2,000,000

qwen3-coder-plus-2025-09-23

Global

60

1,000,000

qwen3-coder-plus-2025-07-22

Global

60

1,000,000

qwen3-coder-flash

Global

1,200

1,000,000

qwen3-coder-flash-2025-07-28

Global

60

1,000,000

North China 2 (Beijing)

ModelRate limiting conditions (triggered when any value is exceeded)
The following are per-minute rate limiting conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Requests per minute (RPM)Tokens consumed per minute (TPM)
Including input and Output tokens

qwen3-coder-plus

5,000

5,000,000

qwen3-coder-plus-2025-09-23

60

1,000,000

qwen3-coder-plus-2025-07-22

60

1,000,000

qwen3-coder-flash

5,000

5,000,000

qwen3-coder-flash-2025-07-28

60

1,000,000

qwen-coder-plus

1,200

1,000,000

qwen-coder-turbo

1,200

1,000,000

Germany (Frankfurt)

ModelService Deployment ScopeRate limiting conditions (triggered when any value is exceeded)
The following are per-minute rate limiting conditions. The service may also limit by RPS (RPM/60) and TPS (TPM/60)
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output tokens

qwen3-coder-plus

Global

5,000

5,000,000

qwen3-coder-plus-2025-09-23

Global

60

1,000,000

qwen3-coder-plus-2025-07-22

Global

60

1,000,000

qwen3-coder-flash

Global

5,000

5,000,000

qwen3-coder-flash-2025-07-28

Global

60

1,000,000

Qwen Translation Model

Singapore

ModelService Deployment ScopeThrottling Conditions (triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may throttle by RPS (RPM/60) and TPS (TPM/60).
Requests Per Minute (RPM)Tokens Per Minute (TPM)
Includes input and output tokens

qwen-mt-plus

International

60

100,000

qwen-mt-flash

International

60

100,000

qwen-mt-lite

International

60

100,000

qwen-mt-turbo

International

60

100,000

US (Virginia)

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may enforce limits via RPS (RPM/60) and TPS (TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes input and output tokens

qwen-mt-plus

Global

60

25,000

qwen-mt-flash

Global

60

35,000

qwen-mt-lite

Global

60

100,000

qwen-mt-lite

United States

60

100,000

North China 2 (Beijing)

ModelRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

qwen-mt-plus

60

25,000

qwen-mt-flash

60

35,000

qwen-mt-lite

60

100,000

qwen-mt-turbo

60

35,000

Germany (Frankfurt)

ModelService Deployment ScopeThrottling conditions (triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may also limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

qwen-mt-plus

Global

60

25,000

qwen-mt-flash

Global

60

35,000

qwen-mt-lite

Global

60

100,000

Qwen Data Mining Model

North China 2 (Beijing)

ModelRate limit conditions (rate limit triggered when any value is exceeded)
The following are rate limit conditions per minute. The service may limit based on RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

qwen-doc-turbo

600

3,000,000

Qwen In-Depth Research Model

North China 2 (Beijing)

ModelRate limit conditions (rate limit triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60)
Number of calls per minute (RPM)Number of Tokens consumed per minute (TPM)
Includes Input and Output Tokens

qwen-deep-research

120

1,200,000

Text Generation-Qwen-Open Source Edition

Qwen Language Model Open Source Edition

Singapore

ModelService Deployment ScopeRate limit conditions (rate limit triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60)
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output tokens

qwen3.8-2.4t-a95b

International

Dynamic rate limiting

qwen3.8-27b

International

Dynamic rate limiting

qwen3.6-35b-a3b

International

600

1,000,000

qwen3.6-27b

International

600

1,000,000

qwen3.5-397b-a17b

International

600

1,000,000

qwen3.5-122b-a10b

International

600

1,000,000

qwen3.5-27b

International

600

1,000,000

qwen3.5-35b-a3b

International

600

5,000,000

qwen3-next-80b-a3b-thinking

International

600

1,000,000

qwen3-next-80b-a3b-instruct

International

600

1,000,000

qwen3-235b-a22b-thinking-2507

International

600

1,000,000

qwen3-235b-a22b-instruct-2507

International

600

1,000,000

qwen3-30b-a3b-thinking-2507

International

600

5,000,000

qwen3-30b-a3b-instruct-2507

International

600

5,000,000

qwen3-235b-a22b

International

600

1,000,000

qwen3-32b

International

600

1,000,000

qwen3-30b-a3b

International

600

1,000,000

qwen3-14b

International

600

1,000,000

qwen3-8b

International

600

1,000,000

US (Virginia)

ModelService Deployment ScopeThrottling Conditions (triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may throttle based on RPS (RPM/60) and TPS (TPM/60).
Requests per Minute (RPM)Tokens per Minute (TPM)
Includes Input and Output Tokens

qwen3.5-397b-a17b

Global

600

1,000,000

qwen3.5-122b-a10b

Global

600

1,000,000

qwen3.5-27b

Global

600

1,000,000

qwen3.6-35b-a3b

Global

600

1,000,000

qwen3.5-35b-a3b

Global

600

1,000,000

qwen3-next-80b-a3b-thinking

Global

600

1,000,000

qwen3-next-80b-a3b-instruct

Global

600

1,000,000

qwen3-235b-a22b-thinking-2507

Global

600

1,000,000

qwen3-235b-a22b-instruct-2507

Global

600

1,000,000

qwen3-30b-a3b-thinking-2507

Global

600

1,000,000

qwen3-30b-a3b-instruct-2507

Global

600

1,000,000

qwen3-235b-a22b

Global

600

1,000,000

qwen3-30b-a3b

Global

600

1,000,000

qwen3-32b

Global

600

1,000,000

qwen3-14b

Global

600

1,000,000

qwen3-8b

Global

600

1,000,000

North China 2 (Beijing)

ModelRate limit conditions (rate limiting is triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes input and Output Tokens

qwen3.8-2.4t-a95b

Dynamic Rate Limiting

qwen3.8-27b

Dynamic Rate Limiting

qwen3.6-35b-a3b

600

1,000,000

qwen3.6-27b

600

1,000,000

qwen3.5-397b-a17b

600

1,000,000

qwen3.5-122b-a10b

600

1,000,000

qwen3.5-27b

600

1,000,000

qwen3.5-35b-a3b

600

1,000,000

qwen3-next-80b-a3b-thinking

600

1,000,000

qwen3-next-80b-a3b-instruct

600

1,000,000

qwen3-235b-a22b-thinking-2507

600

1,000,000

qwen3-235b-a22b-instruct-2507

600

1,000,000

qwen3-30b-a3b-thinking-2507

600

1,000,000

qwen3-30b-a3b-instruct-2507

600

1,000,000

qwen3-235b-a22b

600

1,000,000

qwen3-30b-a3b

600

1,000,000

qwen3-32b

2400

1,000,000

qwen3-14b

600

1,000,000

qwen3-8b

600

1,000,000

Germany (Frankfurt)

ModelService Deployment ScopeRate Limiting Conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit based on RPS (RPM/60) and TPS (TPM/60)
Requests per minute (RPM)Tokens Consumed Per Minute (TPM)
Including Input and Output Tokens

qwen3.5-397b-a17b

Global

600

1,000,000

qwen3.5-122b-a10b

Global

600

1,000,000

qwen3.5-27b

Global

600

1,000,000

qwen3.6-35b-a3b

Global

600

1,000,000

qwen3.5-35b-a3b

Global

600

1,000,000

qwen3-next-80b-a3b-thinking

Global

600

1,000,000

qwen3-next-80b-a3b-instruct

Global

600

1,000,000

qwen3-235b-a22b-thinking-2507

Global

600

1,000,000

qwen3-235b-a22b-instruct-2507

Global

600

1,000,000

qwen3-30b-a3b-thinking-2507

Global

600

1,000,000

qwen3-30b-a3b-instruct-2507

Global

600

1,000,000

qwen3-235b-a22b

Global

600

1,000,000

qwen3-30b-a3b

Global

600

1,000,000

qwen3-32b

Global

2,400

1,000,000

qwen3-14b

Global

600

1,000,000

qwen3-8b

Global

600

1,000,000

Qwen-VL(Visual Understanding/Image-to-Text)

Singapore

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output tokens

qwen3-vl-32b-thinking

International

60

100,000

qwen3-vl-32b-instruct

International

60

100,000

qwen3-vl-30b-a3b-thinking

International

60

100,000

qwen3-vl-30b-a3b-instruct

International

60

100,000

qwen3-vl-8b-thinking

International

60

100,000

qwen3-vl-8b-instruct

International

60

100,000

qwen3-vl-235b-a22b-thinking

International

60

100,000

qwen3-vl-235b-a22b-instruct

International

60

100,000

US (Virginia)

ModelService Deployment ScopeThrottling conditions (triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Token

qwen3-vl-235b-a22b-thinking

Global

60

100,000

qwen3-vl-235b-a22b-instruct

Global

60

100,000

qwen3-vl-32b-thinking

Global

600

1,000,000

qwen3-vl-32b-instruct

Global

600

1,000,000

qwen3-vl-30b-a3b-thinking

Global

600

1,000,000

qwen3-vl-30b-a3b-instruct

Global

600

1,000,000

qwen3-vl-8b-thinking

Global

600

1,000,000

qwen3-vl-8b-instruct

Global

600

1,000,000

North China 2 (Beijing)

ModelRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit based on RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

qwen3-vl-32b-thinking

600

1,000,000

qwen3-vl-32b-instruct

600

1,000,000

qwen3-vl-30b-a3b-thinking

600

1,000,000

qwen3-vl-30b-a3b-instruct

600

1,000,000

qwen3-vl-8b-thinking

600

1,000,000

qwen3-vl-8b-instruct

600

1,000,000

qwen3-vl-235b-a22b-thinking

60

100,000

qwen3-vl-235b-a22b-instruct

60

100,000

Germany (Frankfurt)

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit based on RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output tokens

qwen3-vl-235b-a22b-thinking

Global

60

100,000

qwen3-vl-235b-a22b-instruct

Global

60

100,000

qwen3-vl-32b-thinking

Global

600

1,000,000

qwen3-vl-32b-instruct

Global

600

1,000,000

qwen3-vl-30b-a3b-thinking

Global

600

1,000,000

qwen3-vl-30b-a3b-instruct

Global

600

1,000,000

qwen3-vl-8b-thinking

Global

600

1,000,000

qwen3-vl-8b-instruct

Global

600

1,000,000

Qwen3-Omni

Singapore

ModelService Deployment ScopeThrottling conditions (triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may impose limits based on RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

qwen2.5-omni-7b

International

60

100,000

North China 2 (Beijing)

ModelThrottling conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Including input and Output tokens

qwen2.5-omni-7b

60

100,000

Qwen3-Omni-Captioner

Singapore

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and Output Tokens

qwen3-omni-30b-a3b-captioner

International

60

100,000

North China 2 (Beijing)

ModelThrottling conditions (throttling is triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may enforce limits based on RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and Output Tokens

qwen3-omni-30b-a3b-captioner

60

100,000

Qwen-Coder

Singapore

ModelService Deployment ScopeRate Limit Conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60)
Calls Per Minute (RPM)Tokens Per Minute (TPM)
Including Input and Output Tokens

qwen3-coder-next

International

600

1,000,000

qwen3-coder-480b-a35b-instruct

International

600

1,000,000

qwen3-coder-30b-a3b-instruct

International

600

1,000,000

US (Virginia)

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes Input and Output tokens

qwen3-coder-480b-a35b-instruct

Global

600

1,000,000

qwen3-coder-30b-a3b-instruct

Global

600

1,000,000

North China 2 (Beijing)

ModelRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes Input and Output Token

qwen3-coder-next

600

1,000,000

qwen3-coder-480b-a35b-instruct

600

1,000,000

qwen3-coder-30b-a3b-instruct

600

1,000,000

Germany (Frankfurt)

ModelService Deployment ScopeThrottling conditions (triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may also limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes Input and Output Token

qwen3-coder-480b-a35b-instruct

Global

600

1,000,000

qwen3-coder-30b-a3b-instruct

Global

600

1,000,000

qwen3-coder-next

European Union

600

1,000,000

Text Generation-Third-party Models

DeepSeek

Singapore

ModelService Deployment ScopeThrottling Conditions (triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may throttle by RPS (RPM/60) and TPS (TPM/60).
Calls per Minute (RPM)Tokens Consumed per Minute (TPM)
Includes Input and Output Tokens

deepseek-v4.1-flash

International

10,000

1,200,000

deepseek-v4-pro

International

10,000

1,200,000

deepseek-v4-pro-0813

International

Dynamic Rate Limiting

deepseek-v4-flash-0731

International

15,000

1,200,000

deepseek-v4-flash

International

10,000

1,200,000

deepseek-v3.2

International

10,000

1,200,000

US (Virginia)

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes Input and Output Tokens

deepseek-v4.1-flash

Global

15,000

1,200,000

deepseek-v4-pro

Global

15,000

1,200,000

deepseek-v4-pro-0813

Global

15,000

1,200,000

deepseek-v4-pro

United States

10,000

1,200,000

deepseek-v4-flash-0731

Global

15,000

1,200,000

deepseek-v4-flash

Global

15,000

1,200,000

deepseek-v4-flash

United States

10,000

1,200,000

deepseek-v4-flash-0731

United States

15,000

1,200,000

North China 2 (Beijing)

ModelRate limit conditions (rate limit is triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Including input and Output Tokens

deepseek-v4.1-flash

15,000

1,200,000

deepseek-v4-pro

15,000

1,200,000

deepseek-v4-pro-0813

Dynamic rate limiting

deepseek-v4-flash-0731

15,000

1,200,000

deepseek-v4-flash

15,000

1,200,000

deepseek-v3.2

UseBatch APIWhen calling the service, it is not subject to rate limiting.

15,000

1,000,000

deepseek-v3.2-exp

15,000

1,200,000

deepseek-v3.1

15,000

1,200,000

deepseek-r1-0528

60

100,000

deepseek-r1

UseBatch APIWhen calling the service, it is not subject to rate limiting.

15,000

1,200,000

deepseek-v3

UseBatch APIWhen calling the service, it is not subject to rate limiting.

15,000

1,200,000

deepseek-r1-distill-qwen-7b

15,000

1,200,000

deepseek-r1-distill-qwen-14b

15,000

1,200,000

deepseek-r1-distill-qwen-32b

15,000

1,200,000

deepseek-r1-distill-qwen-1.5b

60

100,000

deepseek-r1-distill-llama-8b

60

100,000

deepseek-r1-distill-llama-70b

60

100,000

Germany (Frankfurt)

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Including input and Output Token

deepseek-v4.1-flash

Global

15,000

1,200,000

deepseek-v4-pro

Global

15,000

1,200,000

deepseek-v4-pro-0813

Global

15,000

1,200,000

deepseek-v4-flash-0731

Global

15,000

1,200,000

deepseek-v4-flash

Global

15,000

1,200,000

Japan (Tokyo)

ModelService Deployment ScopeRate limiting conditions (triggered when any value is exceeded)
The following are rate limiting conditions per minute; the service may limit by RPS (RPM/60) and TPS (TPM/60)
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

deepseek-v4-pro

Japan

10,000

1,200,000

deepseek-v4-pro-0813

Japan

15,000

1,200,000

deepseek-v4-flash

Japan

10,000

1,200,000

deepseek-v4-flash-0731

Japan

15,000

1,200,000

deepseek-v4.1-flash

Global

15,000

1,200,000

deepseek-v4-pro

Global

10,000

1,200,000

deepseek-v4-pro-0813

Global

15,000

1,200,000

deepseek-v4-flash-0731

Global

15,000

1,200,000

deepseek-v4-flash

Global

10,000

1,200,000

Hong Kong (China)

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may apply limits based on RPS (RPM/60) and TPS (TPM/60)
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

deepseek-v4.1-flash

Hong Kong (China)

15,000

1,200,000

deepseek-v4-flash-0731

Hong Kong (China)

15,000

1,200,000

deepseek-v4-pro-0813

Hong Kong (China)

15,000

1,200,000

Kimi

North China 2 (Beijing)

ModelRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit requests by RPS (RPM/60) and TPS (TPM/60)
Requests per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

kimi-k3

Dynamic Rate Limiting

kimi-k2.7-code

500

1,000,000

kimi-k2.6

500

1,000,000

kimi-k2.5

500

1,000,000

kimi-k2-thinking

500

1,000,000

Moonshot-Kimi-K2-Instruct

500

1,000,000

US (Virginia)

ModelService Deployment ScopeRate Limiting Conditions (triggered when any value is exceeded)
The following are per-minute rate limiting conditions. The service may impose limits based on RPS (RPM/60) and TPS (TPM/60).
Requests per Minute (RPM)Tokens per Minute (TPM)
Includes Input and Output Tokens

kimi-k3

Global

15,000

1,200,000

kimi-k3

International

10,000

1,200,000

kimi-k2.7-code

Global

500

1,000,000

kimi-k2.5

Global

500

1,000,000

Germany (Frankfurt)

ModelService Deployment ScopeThrottling conditions (throttling is triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may impose limits based on RPS (RPM/60) and TPS (TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes Input and Output Tokens

kimi-k3

Global

15,000

1,200,000

kimi-k2.7-code

Global

500

1,000,000

kimi-k2.5

Global

500

1,000,000

Hong Kong (China)

ModelService Deployment ScopeRate Limit Conditions (rate limit is triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

kimi-k3

Global

15,000

1,200,000

kimi-k2.7-code

Global

500

1,000,000

Japan (Tokyo)

ModelService Deployment ScopeRate Limit Conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls Per Minute (RPM)Tokens Per Minute (TPM)
Includes input and output tokens

kimi-k3

Global

15,000

1,200,000

kimi-k2.7-code

Global

500

1,000,000

kimi-k2.5

Global

500

1,000,000

Singapore

ModelService Deployment ScopeThrottling conditions (triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may throttle based on RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes Input and Output Tokens

kimi-k3

International

Dynamic Throttling

kimi-k2.7-code

International

500

1,000,000

MiniMax

North China 2 (Beijing)

ModelRate limit conditions (triggered when any value is exceeded)
Below are the rate limit conditions per minute. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Including input and output tokens

MiniMax-M2.5

500

1,000,000

GLM

US (Virginia)

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Including input and output Tokens

glm-5.3

Global

15,000

1,200,000

glm-5.2

Global

500

1,000,000

glm-5.2

United States

500

1,000,000

glm-5.1

Global

500

1,000,000

North China 2 (Beijing)

ModelRate limit conditions (triggered when any value is exceeded)
The following are rate limit conditions per minute. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

glm-5.3

15,000

5,000,000

glm-5.2

500

2,000,000

glm-5.1

500

1,000,000

glm-5

500

1,000,000

glm-4.7

500

1,000,000

glm-4.6

60

1,000,000

Germany (Frankfurt)

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are rate limit conditions per minute. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Including Input and Output Tokens

glm-5.3

Global

15,000

1,200,000

glm-5.2

Global

500

1,000,000

glm-5.1

Global

500

1,000,000

Singapore

ModelService Deployment ScopeRate Limit Conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may apply limits based on RPS (RPM/60) and TPS (TPM/60).
Calls Per Minute (RPM)Tokens Consumed Per Minute (TPM)
Includes input and output Tokens

glm-5.3

International

10,000

5,000,000

glm-5.2

International

500

1,000,000

glm-5.1

International

500

1,000,000

Hong Kong (China)

ModelService Deployment ScopeThrottling conditions (throttling is triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may limit requests by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Token

glm-5.3

Global

15,000

1,200,000

glm-5.2

Global

500

1,000,000

Japan (Tokyo)

ModelService Deployment ScopeRate limiting conditions (triggered when any value is exceeded)
The following are per-minute rate limiting conditions. The service may be limited by RPS (RPM/60) and TPS (TPM/60)
Calls per minute (RPM)Token consumption per minute (TPM)
Includes input and output Token

glm-5.3

Global

15,000

1,200,000

glm-5.1

Global

500

1,000,000

GLM-Z.AI Direct Supply

Singapore

ModelRate limiting conditions (triggered when any value is exceeded)
The following are per-minute rate limiting conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

ZHIPU/GLM-5.3

200

3,000,000

ZHIPU/GLM-5.2

200

3,000,000

Image Generation

Qwen(Qwen-Image)

Singapore

Model

Service Deployment Scope

Rate limit conditions (triggered when any value is exceeded)

Task submission API call limit

Number of tasks being processed simultaneously (concurrency)

qwen-image-3.0-pro

International

5 Times/Minutes

Synchronous interface unlimited / Asynchronous interface 10

qwen-image-3.0

International

20 Times/ Minutes

Synchronous interface unlimited / Asynchronous interface 10

qwen-image-2.1-pro

International

20 Times/ Minutes

Synchronous interface unlimited / Asynchronous interface 10

qwen-image-2.0-pro

International

2 Times/ Minutes

Synchronous interface unlimited

qwen-image-2.0-pro-2026-06-22

International

2 Times/ Minutes

Synchronous interface unlimited

qwen-image-2.0-pro-2026-04-22

International

2 Times/ Minutes

Synchronous interface unlimited

qwen-image-2.0-pro-2026-03-03

International

2 Times/ Minutes

Synchronous interface unlimited

qwen-image-2.0

International

2 Times/ s

Synchronous interface unlimited

qwen-image-2.0-2026-03-03

International

2 Times/ s

Synchronous interface unlimited

qwen-image-max

International

2 Times/ Minutes

Synchronous interface no limit

qwen-image-max-2025-12-30

International

2 Times/ Minutes

Synchronous interface no limit

qwen-image-plus

International

2 Times/ s

Synchronous interface no limit / Asynchronous interface 2

qwen-image-plus-2026-01-09

International

2 Times/ s

Synchronous interface unlimited

qwen-image

International

2 Times/ s

Synchronous interface unlimited / Asynchronous interface 2

qwen-image-edit-max

International

2 Times/ Minutes

Synchronous interface unlimited

qwen-image-edit-max-2026-01-16

International

2 Times/ Minutes

Sync API: no limit

qwen-image-edit-plus

International

2 Times/s

Sync API: no limit

qwen-image-edit-plus-2025-12-15

International

2 Times/s

Sync API: no limit

qwen-image-edit-plus-2025-10-30

International

2 Times/s

Sync API: no limit

qwen-image-edit

International

2 Times/ s

Synchronous interface unlimited

qwen-mt-image-2.0

International

60 Times/ Minutes

Synchronous interface unlimited / Asynchronous interface 2

US (Virginia)

Model

Service Deployment Scope

Rate Limit Conditions (rate limit triggered when any value is exceeded)

Task Submission API Call Limit

Number of Tasks Being Processed Simultaneously (Concurrency)

qwen-image-3.0-pro

Global

5 Times/Minutes

Synchronous interface unlimited / Asynchronous interface 10

qwen-image-3.0

Global

20 Times/Minutes

Synchronous interface unlimited / Asynchronous interface 10

qwen-image-2.1-pro

Global

20 Times/Minutes

Synchronous interface unlimited / Asynchronous interface 10

North China 2 (Beijing)

Model

Rate limit conditions (triggered when any value is exceeded)

Task submission interface call limit

Number of tasks being processed simultaneously (concurrency)

qwen-image-3.0-pro

5 Times/ Minutes

Synchronous interface unlimited / Asynchronous interface 10

qwen-image-3.0

20 Times/ Minutes

Synchronous interface unlimited / Asynchronous interface 10

qwen-image-2.1-pro

20 Times/ Minutes

Synchronous interface unlimited / Asynchronous interface 10

qwen-image-2.0-pro

2 Times/ Minutes

Synchronous interface unlimited

qwen-image-2.0-pro-2026-06-22

2 Times/ Minutes

Synchronous interface unlimited

qwen-image-2.0-pro-2026-04-22

2 Times/ Minutes

Synchronous interface unlimited

qwen-image-2.0-pro-2026-03-03

2 Times/ Minutes

Synchronous interface unlimited

qwen-image-2.0

2 Times/ s

Synchronous interface unlimited

qwen-image-2.0-2026-03-03

2 Times/ s

Synchronous interface unlimited

qwen-image-max

2Times/ Minutes

Synchronous interface unlimited

qwen-image-max-2025-12-30

2Times/ Minutes

Synchronous interface unlimited

qwen-image-plus

2 Times/s

Synchronous interface unlimited / Asynchronous interface 2

qwen-image-plus-2026-01-09

2 Times/s

Synchronous interface unlimited

qwen-image

2 Times/s

Synchronous interface unlimited / Asynchronous interface 2

qwen-image-edit-max

2 Times/ Minutes

Synchronous interface unlimited

qwen-image-edit-max-2026-01-16

2 Times/ Minutes

Synchronous interface unlimited

qwen-image-edit-plus

2 Times/s

Synchronous interface: unlimited

qwen-image-edit-plus-2025-12-15

2 Times/s

Synchronous interface: unlimited

qwen-image-edit-plus-2025-10-30

2 Times/s

Synchronous interface: unlimited

qwen-image-edit

2 Times/s

Synchronous interface: unlimited

qwen-mt-image-2.0

60 Times/Minutes

Synchronous interface unlimited / Asynchronous interface 2

qwen-mt-image

1 Times/s

Synchronous interface unlimited / Asynchronous interface 2

Germany (Frankfurt)

Model

Service Deployment Scope

Throttling conditions (throttling is triggered when any value is exceeded)

Task submission interface call limit

Concurrent tasks being processed (concurrency)

qwen-image-3.0-pro

Global

5 Times/ Minutes

Synchronous interface: no limit / Asynchronous interface: 10

qwen-image-3.0

Global

20 Times/ Minutes

Synchronous interface unlimited / Asynchronous interface 10

qwen-image-2.1-pro

Global

20 Times/ Minutes

Synchronous interface unlimited / Asynchronous interface 10

Hong Kong (China)

Model

Service Deployment Scope

Rate limit conditions (triggered when any value is exceeded)

Task submission interface call limit

Number of tasks being processed simultaneously (concurrency)

qwen-image-3.0-pro

Global

5 Times/ Minutes

Synchronous interface unlimited / Asynchronous interface 10

qwen-image-3.0

Global

20 Times/ Minute

Synchronous interface unlimited / Asynchronous interface 10

qwen-image-2.1-pro

Global

20 Times/ Minute

Synchronous interface unlimited / Asynchronous interface 10

Japan (Tokyo)

Model

Service Deployment Scope

Rate limit conditions (triggered when any value is exceeded)

Task submission interface call limit

Number of tasks being processed simultaneously (concurrency)

qwen-image-3.0-pro

Global

5 Times/ Minute

Synchronous interface unlimited / Asynchronous interface 10

qwen-image-3.0

Global

20 Times/Minutes

Sync interface unlimited / Async interface 10

qwen-image-2.1-pro

Global

20 Times/Minutes

Sync interface unlimited / Async interface 10

Image Generation-Z-Image

Singapore

Model

Service Deployment Scope

Rate limit conditions (triggered when any value is exceeded)

Calls per second (RPS)

Number of tasks being processed simultaneously (concurrency)

z-image-turbo

International

2

Synchronous interface: unlimited

North China 2 (Beijing)

Model

Rate limit conditions (rate limiting is triggered when any value is exceeded)

Calls per second (RPS)

Number of concurrent tasks being processed (concurrency)

z-image-turbo

2

Synchronous interface: unlimited

Wanxiang

Singapore

Model

Service Deployment Scope

Rate limit conditions (rate limiting is triggered when any value is exceeded)

Requests per second (RPS)

Number of tasks being processed simultaneously (concurrency)

wan2.7-image-pro

International

5

5

wan2.7-image

International

5

5

wan2.6-image

International

5

5

wan2.6-t2i

International

5

5

wan2.5-t2i-preview

International

5

5

wan2.2-t2i-flash

International

2

2

wan2.2-t2i-plus

International

2

2

wan2.1-t2i-turbo

International

2

2

wan2.1-t2i-plus

International

2

2

wan2.5-i2i-preview

International

5

5

US (Virginia)

Model

Service Deployment Scope

Rate limiting conditions (rate limiting triggered when any value is exceeded)

Calls per second (RPS)

Number of tasks being processed concurrently (concurrency)

wan2.6-t2i

Global

5

5

wan2.6-image

Global

5

5

North China 2 (Beijing)

Model

Rate limit conditions (triggered when any value is exceeded)

Requests per second (RPS)

Number of concurrent tasks being processed (concurrency)

wan2.7-image-pro

5

5

wan2.7-image

5

5

wan2.6-image

5

5

wan2.6-t2i

1

5

wan2.5-t2i-preview

5

5

wanx2.0-t2i-turbo

2

2

wanx2.1-t2i-turbo

2

2

wanx2.1-t2i-plus

2

2

wan2.2-t2i-flash

2

2

wan2.2-t2i-plus

2

2

wan2.5-i2i-preview

5

5

wanx2.1-imageedit

2

2

Germany (Frankfurt)

Model

Service Deployment Scope

Rate limit conditions (triggered when any value is exceeded)

Requests per second (RPS)

Number of concurrent tasks being processed (concurrency)

wan2.6-t2i

Global

1

5

wan2.6-image

Global

5

5

OutfitAnyone (AI Try-On)

North China 2 (Beijing)

Model

Rate limiting conditions (triggered when any value is exceeded)

Calls per second (RPS)

Number of tasks being processed simultaneously

aitryon-plus

10

5

aitryon-parsing-v1

10

Synchronous interface unlimited

Vidu Series

Singapore

ModelRate limiting conditions (triggered when any value is exceeded)
Calls per minute (RPM)Number of concurrent tasks being processed

vidu/vidu-image_reference2image

300

5

Under a single Bailian API Key, the Vidu reference image generation series models share 5 concurrent slots. That is, the total number of tasks in running status across all models combined must not exceed 5.

Video Generation

HappyHorse Series

Singapore

Model

Service Deployment Scope

Throttling conditions (throttling triggers when any value is exceeded)

Calls per second (RPS)

Number of concurrent tasks being processed (concurrency)

happyhorse-1.1-t2v

International

5

5

happyhorse-1.1-i2v

International

5

5

happyhorse-1.1-r2v

International

5

5

happyhorse-1.0-t2v

International

5

5

happyhorse-1.0-i2v

International

5

5

happyhorse-1.0-r2v

International

5

5

happyhorse-1.0-video-edit

International

5

5

US (Virginia)

Model

Service Deployment Scope

Rate Limit Conditions (triggered when any value is exceeded)

Requests per second (RPS)

Concurrent Tasks Being Processed (Concurrency)

happyhorse-1.1-t2v

Global

5

5

happyhorse-1.1-i2v

Global

5

5

happyhorse-1.1-r2v

Global

5

5

happyhorse-1.0-t2v

Global

5

5

happyhorse-1.0-i2v

Global

5

5

happyhorse-1.0-r2v

Global

5

5

happyhorse-1.0-video-edit

Global

5

5

North China 2 (Beijing)

Model

Throttling Conditions (triggered when any value is exceeded)

Requests per second (RPS)

Number of tasks being processed simultaneously (concurrency)

happyhorse-1.1-t2v

5

5

happyhorse-1.1-i2v

5

5

happyhorse-1.1-r2v

5

5

happyhorse-1.0-t2v

5

5

happyhorse-1.0-i2v

5

5

happyhorse-1.0-r2v

5

5

happyhorse-1.0-video-edit

5

5

Germany (Frankfurt)

Model

Service Deployment Scope

Throttling conditions (triggered when any value is exceeded)

Requests per second (RPS)

Number of tasks being processed simultaneously (concurrency)

happyhorse-1.1-t2v

Global

5

5

happyhorse-1.1-i2v

Global

5

5

happyhorse-1.1-r2v

Global

5

5

happyhorse-1.0-t2v

Global

5

5

happyhorse-1.0-i2v

Global

5

5

happyhorse-1.0-r2v

Global

5

5

happyhorse-1.0-video-edit

Global

5

5

Japan (Tokyo)

Model

Service Deployment Scope

Throttling conditions (triggered when any value is exceeded)

Calls per second (RPS)

Concurrent tasks being processed (concurrency)

happyhorse-1.0-video-edit

Global

5

5

Wanx Series

Singapore

Model

Service Deployment Scope

Throttling conditions (triggered when any value is exceeded)

Requests per second (RPS)

Concurrent requests in progress

wan3.0-video-prime

International

5

5

wan3.0-video

International

5

5

wan2.7-t2v-2026-06-12

International

5

5

wan2.7-t2v-2026-04-25

International

5

5

wan2.7-t2v

International

5

5

wan2.6-t2v

International

5

5

wan2.5-t2v-preview

International

5

5

wan2.2-t2v-plus

International

2

2

wan2.1-t2v-turbo

International

2

2

wan2.1-t2v-plus

International

2

2

wan2.7-i2v-2026-04-25

International

5

5

wan2.7-i2v

International

5

5

wan2.6-i2v-flash

International

5

5

wan2.6-i2v

International

5

5

wan2.5-i2v-preview

International

5

5

wan2.2-i2v-flash

International

2

2

wan2.1-i2v-plus

International

2

2

wan2.1-i2v-turbo

International

2

2

wan2.2-i2v-plus

International

2

2

wan2.2-kf2v-flash

International

2

2

wan2.1-kf2v-plus

International

1

2

wan2.1-vace-plus

International

2

2

wan2.7-videoedit

International

5

5

wan2.7-r2v

International

5

5

wan2.6-r2v-flash

International

5

5

wan2.6-r2v

International

5

5

wan2.2-animate-move

International

5

1

wan2.2-animate-mix

International

5

1

US (Virginia)

Model

Service Deployment Scope

Rate limiting conditions (triggered when any value is exceeded)

Calls per second (RPS)

Number of concurrent tasks being processed (concurrency)

wan3.0-video-prime

Global

5

5

wan3.0-video

Global

5

5

wan2.7-r2v-2026-06-12

Global

5

5

wan2.6-t2v

Global

5

5

wan2.6-i2v

Global

5

5

wan2.6-r2v

Global

5

5

wan2.6-t2v

United States

5

5

wan2.6-i2v

United States

5

5

North China 2 (Beijing)

Model

Rate limit conditions (triggered when exceeding any of the values)

Requests per second (RPS)

Number of tasks being processed concurrently (concurrency)

wan3.0-video-prime

5

5

wan3.0-video

5

5

wan2.7-r2v-2026-06-12

5

5

wan2.7-t2v-2026-06-12

5

5

wan2.7-t2v-2026-04-25

5

5

wan2.7-t2v

5

5

wan2.6-t2v

5

5

wan2.5-t2v-preview

5

5

wan2.2-t2v-plus

2

2

wanx2.1-t2v-turbo

2

2

wanx2.1-t2v-plus

2

2

wan2.7-i2v-2026-04-25

5

5

wan2.7-i2v

5

5

wan2.6-i2v-flash

5

5

wan2.6-i2v

5

5

wan2.5-i2v-preview

5

5

wan2.2-i2v-plus

2

2

wanx2.1-i2v-turbo

2

2

wanx2.1-i2v-plus

2

2

wan2.2-kf2v-flash

2

2

wanx2.1-kf2v-plus

2

2

wanx2.1-vace-plus

2

2

wan2.7-videoedit

5

5

wan2.7-r2v

5

5

wan2.6-r2v-flash

5

5

wan2.6-r2v

5

5

wan2.2-s2v-detect

5

Synchronous interface has no limit

wan2.2-s2v

5

1

wan2.2-animate-move

5

1

wan2.2-animate-mix

5

1

Germany (Frankfurt)

Model

Service Deployment Scope

Throttling conditions (triggered when any value is exceeded)

Calls per second (RPS)

Number of tasks being processed concurrently (concurrency)

wan3.0-video-prime

Global

5

5

wan3.0-video

Global

5

5

wan2.6-t2v

Global

5

5

wan2.6-i2v

Global

5

5

wan2.6-r2v

Global

5

5

Japan (Tokyo)

Model

Service Deployment Scope

Rate limit conditions (triggered when any value is exceeded)

Requests per second (RPS)

Number of concurrent tasks being processed (concurrency)

wan3.0-video-prime

Global

5

5

wan3.0-video

Global

5

5

Hong Kong (China)

Model

Service Deployment Scope

Rate limit conditions (rate limit is triggered when any value is exceeded)

Requests per second (RPS)

Number of concurrent tasks being processed (concurrency)

wan3.0-video-prime

Global

5

5

wan3.0-video

Global

5

5

AnimateAnyone (Dancing Portrait)

North China 2 (Beijing)

Model

Requests per second (RPS)

Number of concurrent tasks being processed

animate-anyone-detect-gen2

5

No limit for synchronous API

animate-anyone-template-gen2

5

1

At any given time, only 1 job is actually running, and other jobs in the queue are waiting.

animate-anyone-gen2

5

1

At any given time, only 1 job is actually running, and other jobs in the queue are waiting.

EMO

North China 2 (Beijing)

Model

Requests per second (RPS)

Number of tasks being processed simultaneously

emo-detect-v1

5

No limit for synchronous API

emo-v1

5

1

At any given time, only 1 job is actually running, and other jobs in the queue are waiting.

LivePortrait

North China 2 (Beijing)

Model

Number of calls per second(RPS)

Number of tasks being processed concurrently

liveportrait-detect

5

No limit for synchronous interface

liveportrait

5

1

At any given moment, only 1 job is actually running, while other jobs in the queue are in a queued state.

Shengdongrenxiang VideoRetalk

North China 2 (Beijing)

Model

Number of calls per second(RPS)

Number of tasks being processed concurrently

videoretalk

1

1

At any given moment, only 1 job is actually running, while other jobs in the queue are in a queued state.

Emoji

North China 2 (Beijing)

Model

Requests per second (RPS)

Concurrent processing tasks

emoji-detect-v1

1

No limit on synchronous interfaces

emoji-v1

1

1

At any given moment, only 1 job is actually running, and other jobs in the queue are in a pending state.

Video Style Repainting

North China 2 (Beijing)

Model

Requests per second (RPS)

Concurrent processing tasks

video-style-transform

20

2

At the same time, only 1 job is actually running, and the jobs in other queues are in the queueing state.

Vidu Series

Singapore

ModelRate Limit Conditions (triggered when any value is exceeded)
Calls per second (RPS)Number of tasks being processed simultaneously (concurrency)

vidu/viduq3-mix_reference2video

5

5

Under a single Bailian API Key, the models of the Vidu video generation series share 5 concurrent slots. That is, the total number of tasks in running state across all models cannot exceed 5.

vidu/viduq3-ad_reference2video

5

vidu/viduq3-drama_reference2video

5

vidu/viduq2-pro-fast_img2video

5

Audio generation

China (Beijing)

ModelRequests per second (RPS)
qwen-audio-3.1-tts-next3

Music generation

North China 2 (Beijing)

Model

Calls per minute (RPM)

fun-music-preview

180

fun-music-v1

180

Speech Dialogue

Real-time Speech Dialogue

Singapore

ModelService Deployment ScopeRate Limit Conditions (Triggered When Any Value Is Exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls Per Minute (RPM)Tokens Consumed Per Minute (TPM)
Includes Input and Output Tokens

qwen-audio-3.1-realtime-plus

International

60

100,000

qwen-audio-3.0-realtime-plus

International

60

100,000

qwen-audio-3.0-realtime-flash

International

60

100,000

North China 2 (Beijing)

ModelRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes Input and Output tokens

qwen-audio-3.1-realtime-plus

60

100,000

qwen-audio-3.0-realtime-plus

60

100,000

qwen-audio-3.0-realtime-flash

60

100,000

Speech Synthesis

Qwen-Audio-TTS

Singapore

Model

Service Deployment Scope

Requests per second (RPS)

qwen-audio-3.0-tts-plus

International

3

qwen-audio-3.0-tts-flash

International

3

North China 2 (Beijing)

Model

Requests per second (RPS)

qwen-audio-3.0-tts-plus

3

qwen-audio-3.0-tts-flash

3

Qwen-TTS

Singapore

Qwen3-TTS-Instruct-Flash

Model

Service Deployment Scope

Requests per minute (RPM)

qwen3-tts-instruct-flash

International

180

qwen3-tts-instruct-flash-2026-01-26

International

180

Qwen3-TTS-VD

Model

Service Deployment Scope

Calls per minute (RPM)

qwen3-tts-vd-2026-01-26

International

180

Qwen3-TTS-VC

Model

Service Deployment Scope

Calls per minute (RPM)

qwen3-tts-vc-2026-01-22

International

180

Qwen3-TTS-Flash

Model

Service Deployment Scope

Requests Per Minute (RPM)

qwen3-tts-flash

International

180

qwen3-tts-flash-2025-11-27

International

180

qwen3-tts-flash-2025-09-18

International

10

North China 2 (Beijing)

Qwen3-TTS-Instruct-Flash

Model

Requests Per Minute (RPM)

qwen3-tts-instruct-flash

180

qwen3-tts-instruct-flash-2026-01-26

180

Qwen3-TTS-VD

Model

Requests Per Minute (RPM)

qwen3-tts-vd-2026-01-26

180

Qwen3-TTS-VC

Model

Requests per minute (RPM)

qwen3-tts-vc-2026-01-22

180

Qwen3-TTS-Flash

Model

Requests per minute (RPM)

qwen3-tts-flash

180

qwen3-tts-flash-2025-11-27

180

qwen3-tts-flash-2025-09-18

10

Qwen-TTS

ModelThrottling conditions (triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may throttle based on RPS (RPM/60) and TPS (TPM/60)
Requests per minute (RPM)Tokens consumed per minute (TPM)
Includes Input and Output tokens

qwen-tts

10

100,000

qwen-tts-latest

qwen-tts-2025-05-22

qwen-tts-2025-04-10

Qwen-TTS-Realtime

Singapore

Qwen3-TTS-Instruct-Flash-Realtime

Model

Service Deployment Scope

Calls per minute (RPM)

qwen3-tts-instruct-flash-realtime

International

180

qwen3-tts-instruct-flash-realtime-2026-01-22

International

180

Qwen3-TTS-VD-Realtime

Model

Service Deployment Scope

Calls per minute (RPM)

qwen3-tts-vd-realtime-2026-01-15

International

180

qwen3-tts-vd-realtime-2025-12-16

International

Qwen3-TTS-VC-Realtime

Model

Service Deployment Scope

Calls per minute (RPM)

qwen3-tts-vc-realtime-2026-01-15

International

180

qwen3-tts-vc-realtime-2025-11-27

International

Qwen3-TTS-Flash-Realtime

Model

Service Deployment Scope

Calls per minute (RPM)

qwen3-tts-flash-realtime

International

180

qwen3-tts-flash-realtime-2025-11-27

International

180

qwen3-tts-flash-realtime-2025-09-18

International

10

North China 2 (Beijing)

Qwen3-TTS-Instruct-Flash-Realtime

Model

Calls per minute (RPM)

qwen3-tts-instruct-flash-realtime

180

qwen3-tts-instruct-flash-realtime-2026-01-22

180

Qwen3-TTS-VD-Realtime

Model

Calls per minute (RPM)

qwen3-tts-vd-realtime-2026-01-15

180

qwen3-tts-vd-realtime-2025-12-16

Qwen3-TTS-VC-Realtime

Model

Calls per minute (RPM)

qwen3-tts-vc-realtime-2026-01-15

180

qwen3-tts-vc-realtime-2025-11-27

Qwen3-TTS-Flash-Realtime

Model

Calls per minute (RPM)

qwen3-tts-flash-realtime

180

qwen3-tts-flash-realtime-2025-11-27

180

qwen3-tts-flash-realtime-2025-09-18

10

Qwen-TTS-Realtime

ModelThrottling conditions (throttling is triggered when any value is exceeded)
The following are per-minute rate limiting conditions. The service may also limit by RPS (RPM/60) and TPS (TPM/60)
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output tokens

qwen-tts-realtime

10

100,000

qwen-tts-realtime-latest

qwen-tts-realtime-2025-07-15

Qwen-TTS Voice Cloning

Singapore

Model

Service Deployment Scope

Calls per minute (RPM)

qwen-voice-enrollment

International

180

North China 2 (Beijing)

Model

Calls per minute (RPM)

qwen-voice-enrollment

180

Qwen-TTS Voice Design

Singapore

Model

Service Deployment Scope

Calls per minute (RPM)

qwen-voice-design

International

180

North China 2 (Beijing)

Model

Calls per minute (RPM)

qwen-voice-design

180

CosyVoice

Singapore

Model

Service Deployment Scope

Requests per second (RPS)

cosyvoice-v3-plus

International

3

cosyvoice-v3-flash

International

North China 2 (Beijing)

Model

Requests per second (RPS)

cosyvoice-v3.5-plus

3

cosyvoice-v3.5-flash

cosyvoice-v3-plus

cosyvoice-v3-flash

cosyvoice-v2

Qwen-Audio-TTS/CosyVoice Voice Cloning/Design

Qwen-Audio-TTS/CosyVoice Voice Cloning/Design share one model and share the rate limit quota.

Singapore

Model

Service Deployment Scope

Requests per second (RPS)

voice-enrollment

International

10

North China 2 (Beijing)

Model

Requests per second (RPS)

voice-enrollment

10

Speech Recognition (speech to text) and Translation (speech converted into text in a specified language)

Qwen3-LiveTranslate-Flash

Singapore

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Requests per minute (RPM)Tokens consumed per minute (TPM)
Includes Input and Output Tokens

qwen3-livetranslate-flash

International

100

100,000

qwen3-livetranslate-flash-2025-12-01

International

100

100,000

North China 2 (Beijing)

ModelRate Limit Conditions (Triggered When Any Value Is Exceeded)
The following are rate limit conditions per minute. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls Per Minute (RPM)Tokens Per Minute (TPM)
Includes Input and Output Tokens

qwen3-livetranslate-flash

100

100,000

qwen3-livetranslate-flash-2025-12-01

Qwen-LiveTranslate-Flash-Realtime

Singapore

ModelService Deployment ScopeRate limit conditions (rate limit is triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes Input and Output Tokens

qwen3.8-livetranslate-flash-realtime

International

10

100,000

qwen3.5-livetranslate-flash-realtime

International

qwen3.5-livetranslate-flash-realtime-2026-05-19

International

qwen3-livetranslate-flash-realtime

International

qwen3-livetranslate-flash-realtime-2025-09-22

International

North China 2 (Beijing)

ModelRate limit conditions (rate limit triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes Input and Output Token

qwen3.8-livetranslate-flash-realtime

10

100,000

qwen3.5-livetranslate-flash-realtime

qwen3.5-livetranslate-flash-realtime-2026-05-19

qwen3-livetranslate-flash-realtime

qwen3-livetranslate-flash-realtime-2025-09-22

Qwen-Audio-3.1-ASR-Flash-Message

Singapore

Model nameDeployment scopeRequests per minute (RPM)
qwen-audio-3.1-asr-flash-messageInternational1200

China (Beijing)

Model nameRequests per minute (RPM)
qwen-audio-3.1-asr-flash-message1200

Qwen-Audio-3.x-ASR-Flash-Streaming

Singapore

Model name

Service deployment scope

Requests per second (RPS)

Requests per minute (RPM)

qwen-audio-3.1-asr-flash-streaming

International

—600

qwen-audio-3.0-asr-flash-streaming

International

20

—

China (Beijing)

Model name

Requests per second (RPS)

Requests per minute (RPM)

qwen-audio-3.1-asr-flash-streaming

—600

qwen-audio-3.0-asr-flash-streaming

20

—

Qwen-Audio-3.x-ASR-Flash-Filetrans

Singapore

Model name

Service deployment scope

Requests per minute (RPM)

qwen-audio-3.1-asr-flash-filetrans

International

600

qwen-audio-3.0-asr-flash-filetrans

International

600

China (Beijing)

Model name

Requests per minute (RPM)

qwen-audio-3.1-asr-flash-filetrans

600

qwen-audio-3.0-asr-flash-filetrans

600

Qwen-Audio-3.x-ASR-Flash

Singapore

Model name

Service deployment scope

Requests per minute (RPM)

qwen-audio-3.1-asr-flash

International

600

qwen-audio-3.0-asr-flash

International

600

China (Beijing)

Model name

Requests per minute (RPM)

qwen-audio-3.1-asr-flash

600

qwen-audio-3.0-asr-flash

600

Qwen-ASR

Singapore

Qwen3-ASR-Flash-Filetrans

Model

Service Deployment Scope

Calls per minute (RPM)

qwen3-asr-flash-filetrans

International

100

qwen3-asr-flash-filetrans-2025-11-17

International

Qwen3-ASR-Flash

Model

Service Deployment Scope

Calls per minute (RPM)

qwen3-asr-flash

International

100

qwen3-asr-flash-2026-02-10

International

qwen3-asr-flash-2025-09-08

International

US (Virginia)

Model

Service Deployment Scope

Calls per minute (RPM)

qwen3-asr-flash

United States

100

qwen3-asr-flash-2025-09-08

United States

North China 2 (Beijing)

Qwen3-ASR-Flash-Filetrans

Model

Calls per minute (RPM)

qwen3-asr-flash-filetrans

100

qwen3-asr-flash-filetrans-2025-11-17

Qwen3-ASR-Flash

Model

Calls per minute (RPM)

qwen3-asr-flash

100

qwen3-asr-flash-2026-02-10

qwen3-asr-flash-2025-09-08

Qwen-ASR-Realtime

Singapore

Model

Service Deployment Scope

Calls per second (RPS)

qwen3-asr-flash-realtime

International

20

qwen3-asr-flash-realtime-2026-02-10

International

qwen3-asr-flash-realtime-2025-10-27

International

North China 2 (Beijing)

Model

Requests per second (RPS)

qwen3-asr-flash-realtime

20

qwen3-asr-flash-realtime-2026-02-10

qwen3-asr-flash-realtime-2025-10-27

Paraformer

North China 2 (Beijing)

Model

Requests per second (RPS)

paraformer-realtime-v2

20

paraformer-realtime-8k-v2

Model

Requests per minute (RPM)

paraformer-v2

1,200

Model

Requests per second (RPS)

Concurrent processing tasks (concurrency)

paraformer-8k-v2

20

100

Fun-ASR

Singapore

Model

Service Deployment Scope

Calls per minute (RPM)

fun-asr

International

600

fun-asr-2025-11-07

International

600

fun-asr-2025-08-25

International

600

fun-asr-mtl

International

100

fun-asr-mtl-2025-08-25

International

100

fun-asr-flash-2026-06-15

International

600

North China 2 (Beijing)

Model

Requests Per Minute (RPM)

fun-asr

600

fun-asr-2025-11-07

fun-asr-2025-08-25

fun-asr-mtl

fun-asr-mtl-2025-08-25

fun-asr-flash-2026-06-15

Fun-ASR-Realtime

Singapore

Model

Service Deployment Scope

Requests Per Second (RPS)

fun-asr-realtime

International

20

fun-asr-realtime-2025-11-07

International

North China 2 (Beijing)

Model

Requests Per Second (RPS)

fun-asr-realtime

20

fun-asr-realtime-2026-02-28

fun-asr-realtime-2025-11-07

fun-asr-realtime-2025-09-15

fun-asr-flash-8k-realtime

fun-asr-flash-8k-realtime-2026-01-28

Text Embedding

Singapore

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

qwen3.7-text-embedding

International

24,000

1,000,000

text-embedding-v4

International

1,800

1,000,000

text-embedding-v3

International

6,000

24,000,000

North China 2 (Beijing)

ModelRate limit conditions (triggered when any value is exceeded)
Requests per second (RPS)Tokens consumed per minute (TPM)
Includes input and Output tokens

text-embedding-v4

UseBatch APIWhen calling the service, it is not subject to rate limit.

30

1,200,000

Hong Kong (China)

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60)
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and Output Token

text-embedding-v4

Hong Kong (China)

1,800

1,200,000

Multimodal Embedding

Singapore

ModelService Deployment ScopeRate limit conditions
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Input only Token

tongyi-embedding-vision-plus

International

600

200,000

tongyi-embedding-vision-flash

International

600

200,000

North China 2 (Beijing)

ModelRate limit conditions
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Input only Token

qwen3-vl-embedding

2,400

1,200,000

multimodal-embedding-v1

120

1,000,000

Reranking Model

Singapore

ModelService Deployment ScopeRate Limit Conditions
The following are per-minute rate limit conditions. The service may be limited by RPS (RPM/60) and TPS (TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Input Tokens Only

qwen3-rerank

International

5,400

5,000,000,000

North China 2 (Beijing)

ModelRate Limit Conditions
The following are rate limit conditions per minute. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Requests Per Minute (RPM)Tokens Per Minute (TPM)
Only Input Tokens

qwen3-rerank

5,400

5,000,000,000

qwen3-vl-rerank

600

9,000,000

gte-rerank-v2

5,040

4,980,000,000

Industry

Intent Understanding

North China 2 (Beijing)

ModelRate Limit Conditions (triggered when any value is exceeded)
The following are rate limit conditions per minute. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Requests per minute (RPM)Tokens consumed per minute (TPM)
Includes Input and Output Tokens

tongyi-intent-detect-v3

1,200

1,000,000

Role-Playing

Singapore

ModelService Deployment ScopeThrottling conditions (triggered when any value is exceeded)
The following are per-minute throttling conditions. The service may throttle based on RPS (RPM/60) and TPS (TPM/60).
Requests per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Token

qwen-plus-character

International

120

500,000

qwen-flash-character

International

120

500,000

qwen-plus-character-ja

International

120

500,000

US (Virginia)

ModelService Deployment ScopeRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and output Tokens

qwen-plus-character

Global

120

500,000

qwen-flash-character

Global

120

500,000

North China 2 (Beijing)

ModelRate limit conditions (triggered when any value is exceeded)
The following are per-minute rate limit conditions. The service may limit by RPS (RPM/60) and TPS (TPM/60).
Requests per minute (RPM)Tokens per minute (TPM)
Includes input and output Tokens

qwen-plus-character

120

500,000

qwen-flash-character

120

500,000

Decision model

Singapore

Model nameRate limits (triggered if any value is exceeded)
The limits below are specified per minute. The service may enforce limits based on records per second (RPS) (RPM/60) and Transactions Per Second (TPS) (TPM/60).
Requests Per Minute (RPM)Tokens Per Minute (TPM)
Includes input and output tokens

decision-model-preview

1,200

2,000,000

China (Beijing)

Model nameRate limits (triggered if any value is exceeded)
The limits below are specified per minute. The service may enforce limits based on records per second (RPS) (RPM/60) and Transactions Per Second (TPS) (TPM/60).
Requests Per Minute (RPM)Tokens Per Minute (TPM)
Includes input and output tokens

decision-model-preview

1,200

2,000,000

Offline Models

For more information, see Model Offline Mechanism.

Offline on Jan 30, 2026

CategoryModelRate limit conditions (triggered when any value is exceeded)
Calls per minute (RPM)Tokens consumed per minute (TPM)
Includes input and Output Tokens

Qwen Plus

qwen-plus-2024-11-27

0

0

qwen-plus-2024-11-25

qwen-plus-2024-09-19

qwen-plus-2024-08-06

Qwen Turbo

qwen-turbo-2024-09-19

Qwen-VL

qwen-vl-max-2024-10-30

qwen-vl-max-2024-08-09

qwen-vl-plus-2024-08-09

Offline on Aug 20, 2025

CategoryModelRate Limit Conditions (triggered when any value is exceeded)
Calls Per Minute (RPM)Tokens Per Minute (TPM)
Including Input and Output Tokens

Text Generation-Qwen

qwen2-72b-instruct

0

0

qwen2-57b-a14b-instruct

qwen2-7b-instruct

qwen1.5-110b-chat

qwen1.5-72b-chat

qwen1.5-32b-chat

qwen1.5-14b-chat

qwen1.5-7b-chat