# qwen-turbo

> The Turbo model of the Qwen3 series. It effectively integrates thinking mode and non-thinking mode, allowing for mode switching during conversations. Its reasoning capabilities rival those of QwQ-32B with a smaller parameter size, while its general capabilities significantly surpass those of Qwen2.5-Turbo, achieving the SOTA level in the same scale within the industry.This model version is functionally equivalent to the snapshot model qwen-turbo-2025-04-28.

## Inference Service Provider <span id="h-490b2834e6" />

The inference service provider for `qwen-turbo` is Alibaba Cloud Model Studio.

## Model Capabilities <span id="h-a8f9d449c1" />

<Tabs>
  <Tab title="China (Beijing)">
    <table><thead><tr><th>Capability</th><th>Support</th><th>Capability</th><th>Support</th></tr></thead><tbody><tr><td><p>Input Modality</p></td><td><p><strong>Text</strong></p></td><td><p>Output Modality</p></td><td><p><strong>Text</strong></p></td></tr><tr><td><p>Model Experience</p></td><td><p>Supported</p></td><td><p>Function Calling</p></td><td><p>Unsupported</p></td></tr><tr><td><p>Structured Outputs</p></td><td><p>Supported</p></td><td><p>Web Search</p></td><td><p>Unsupported</p></td></tr><tr><td><p>Prefix Completion</p></td><td><p>Unsupported</p></td><td><p>Context Caching</p></td><td><p>Supported</p></td></tr><tr><td><p>Batch Inference</p></td><td><p>Unsupported</p></td><td><p>Fine-tuning</p></td><td><p>Unsupported</p></td></tr></tbody></table>
  </Tab>

  <Tab title="Singapore">
    Scope: International

    <table><thead><tr><th>Capability</th><th>Support</th><th>Capability</th><th>Support</th></tr></thead><tbody><tr><td><p>Input Modality</p></td><td><p><strong>Text</strong></p></td><td><p>Output Modality</p></td><td><p><strong>Text</strong></p></td></tr><tr><td><p>Model Experience</p></td><td><p>Supported</p></td><td><p>Function Calling</p></td><td><p>Unsupported</p></td></tr><tr><td><p>Structured Outputs</p></td><td><p>Supported</p></td><td><p>Web Search</p></td><td><p>Unsupported</p></td></tr><tr><td><p>Prefix Completion</p></td><td><p>Supported</p></td><td><p>Context Caching</p></td><td><p>Supported</p></td></tr><tr><td><p>Batch Inference</p></td><td><p>Supported</p></td><td><p>Fine-tuning</p></td><td><p>Unsupported</p></td></tr></tbody></table>
  </Tab>
</Tabs>

## Context Limits <span id="h-705e09ab60" />

<table><thead><tr><th>Parameter</th><th>Value</th><th>Parameter</th><th>Value</th></tr></thead><tbody><tr><td><p>Max Input Length</p></td><td><p>98304</p></td><td><p>Max Output Length</p></td><td><p>16384</p></td></tr><tr><td><p>Context Window</p></td><td><p>131072</p></td><td /><td /></tr></tbody></table>

## Pricing <span id="h-72afbe6c78" />

This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit [Model Studio Console](https://modelstudio.console.alibabacloud.com/ap-southeast-1/model/market) for promotional offers.

<Tabs>
  <Tab title="China (Beijing)">
    <table><thead><tr><th>Billing Item</th><th>Price (USD)</th><th>Unit</th></tr></thead><tbody><tr><td><p>Input</p></td><td><p>0.044</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Output</p></td><td><p>0.087</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Input(Thinking)</p></td><td><p>0.044</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Output(Thinking)</p></td><td><p>0.431</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Input(Implicit Cache)</p></td><td><p>0.009</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Input(Thinking Implicit Cache)</p></td><td><p>0.009</p></td><td><p>Per 1M tokens</p></td></tr></tbody></table>
  </Tab>

  <Tab title="Singapore">
    Scope: International

    <table><thead><tr><th>Billing Item</th><th>Price (USD)</th><th>Unit</th></tr></thead><tbody><tr><td><p>Input</p></td><td><p>0.05</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Output</p></td><td><p>0.2</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Input(Thinking)</p></td><td><p>0.05</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Output(Thinking)</p></td><td><p>0.5</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Input(Implicit Cache)</p></td><td><p>0.01</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Input(Thinking Implicit Cache)</p></td><td><p>0.01</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Thinking Output(Batch File)</p></td><td><p>0.25</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Input(Batch File)</p></td><td><p>0.025</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Output(Batch File)</p></td><td><p>0.1</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Input(Thinking Batch File)</p></td><td><p>0.025</p></td><td><p>Per 1M tokens</p></td></tr></tbody></table>
  </Tab>
</Tabs>

## Rate Limits <span id="h-13fb75232e" />

<Tabs>
  <Tab title="China (Beijing)">
    <table><thead><tr><th>Parameter</th><th>Value</th></tr></thead><tbody><tr><td><p>RPM (Requests Per Minute)</p></td><td><p>1200</p></td></tr><tr><td><p>TPM (Tokens Per Minute)</p></td><td><p>5,000,000</p></td></tr></tbody></table>
  </Tab>

  <Tab title="Singapore">
    Scope: International

    <table><thead><tr><th>Parameter</th><th>Value</th></tr></thead><tbody><tr><td><p>RPM (Requests Per Minute)</p></td><td><p>600</p></td></tr><tr><td><p>TPM (Tokens Per Minute)</p></td><td><p>5,000,000</p></td></tr></tbody></table>
  </Tab>
</Tabs>