# qwen-omni-turbo

> The new multimodal understanding and generation model, supporting text, image, speech, and video input as well as mixed input understanding. It can simultaneously stream text and speech, with significantly enhanced speed for multimodal content understanding. Additionally, it offers four natural dialogue tones.This model version is functionally equivalent to the snapshot model qwen-omni-turbo-2025-03-26.

## Inference Service Provider <span id="h-490b2834e6" />

The inference service provider for `qwen-omni-turbo` is Alibaba Cloud Model Studio.

## Model Capabilities <span id="h-a8f9d449c1" />

<Tabs>
  <Tab title="China (Beijing)">
    <table><thead><tr><th>Capability</th><th>Support</th><th>Capability</th><th>Support</th></tr></thead><tbody><tr><td><p>Input Modality</p></td><td><p><strong>Text</strong> <strong>Image</strong> <strong>Video</strong> <strong>Audio</strong></p></td><td><p>Output Modality</p></td><td><p><strong>Text</strong> <strong>Audio</strong></p></td></tr><tr><td><p>Model Experience</p></td><td><p>Supported</p></td><td><p>Function Calling</p></td><td><p>Unsupported</p></td></tr><tr><td><p>Structured Outputs</p></td><td><p>Unsupported</p></td><td><p>Web Search</p></td><td><p>Unsupported</p></td></tr><tr><td><p>Prefix Completion</p></td><td><p>Unsupported</p></td><td><p>Context Caching</p></td><td><p>Unsupported</p></td></tr><tr><td><p>Batch Inference</p></td><td><p>Supported</p></td><td><p>Fine-tuning</p></td><td><p>Unsupported</p></td></tr></tbody></table>
  </Tab>

  <Tab title="Singapore">
    Scope: International

    <table><thead><tr><th>Capability</th><th>Support</th><th>Capability</th><th>Support</th></tr></thead><tbody><tr><td><p>Input Modality</p></td><td><p><strong>Text</strong> <strong>Image</strong> <strong>Video</strong> <strong>Audio</strong></p></td><td><p>Output Modality</p></td><td><p><strong>Text</strong> <strong>Audio</strong></p></td></tr><tr><td><p>Model Experience</p></td><td><p>Supported</p></td><td><p>Function Calling</p></td><td><p>Unsupported</p></td></tr><tr><td><p>Structured Outputs</p></td><td><p>Unsupported</p></td><td><p>Web Search</p></td><td><p>Unsupported</p></td></tr><tr><td><p>Prefix Completion</p></td><td><p>Unsupported</p></td><td><p>Context Caching</p></td><td><p>Supported</p></td></tr><tr><td><p>Batch Inference</p></td><td><p>Unsupported</p></td><td><p>Fine-tuning</p></td><td><p>Unsupported</p></td></tr></tbody></table>
  </Tab>
</Tabs>

## Context Limits <span id="h-705e09ab60" />

<table><thead><tr><th>Parameter</th><th>Value</th><th>Parameter</th><th>Value</th></tr></thead><tbody><tr><td><p>Max Input Length</p></td><td><p>30720</p></td><td><p>Max Output Length</p></td><td><p>2048</p></td></tr><tr><td><p>Context Window</p></td><td><p>32768</p></td><td><p /></td><td><p /></td></tr></tbody></table>

## Pricing <span id="h-34214685cf" />

This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit [Model Studio Console](https://modelstudio.console.alibabacloud.com/ap-southeast-1/model/market) for promotional offers.

<Tabs>
  <Tab title="China (Beijing)">
    <table><thead><tr><th>Billing Item</th><th>Price (USD)</th><th>Unit</th></tr></thead><tbody><tr><td><p>Input: Text</p></td><td><p>0.058</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Input: Audio</p></td><td><p>3.584</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Input: Vision</p></td><td><p>0.216</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Output: Text (When input contains only text)</p></td><td><p>0.23</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Output: Text (When input contains images/audio/video)</p></td><td><p>0.646</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Output: Text\&Audio (Output text is not charged)</p></td><td><p>7.168</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Input: Vision(Implicit Cache)</p></td><td><p>0.044</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Input: Text(Implicit Cache)</p></td><td><p>0.012</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Input: Audio(Implicit Cache)</p></td><td><p>0.717</p></td><td><p>Per 1M tokens</p></td></tr></tbody></table>
  </Tab>

  <Tab title="Singapore">
    Scope: International

    <table><thead><tr><th>Billing Item</th><th>Price (USD)</th><th>Unit</th></tr></thead><tbody><tr><td><p>Input: Text</p></td><td><p>0.07</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Input: Audio</p></td><td><p>4.44</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Input: Vision</p></td><td><p>0.21</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Output: Text (When input contains only text)</p></td><td><p>0.27</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Output: Text (When input contains images/audio/video)</p></td><td><p>0.63</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Output: Text\&Audio (Output text is not charged)</p></td><td><p>8.89</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Input: Vision(Implicit Cache)</p></td><td><p>0.04</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Input: Text(Implicit Cache)</p></td><td><p>0.015</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Input: Audio(Implicit Cache)</p></td><td><p>0.89</p></td><td><p>Per 1M tokens</p></td></tr></tbody></table>
  </Tab>
</Tabs>

## Rate Limits <span id="h-66bb936f07" />

<Tabs>
  <Tab title="China (Beijing)">
    <table><thead><tr><th>Parameter</th><th>Value</th></tr></thead><tbody><tr><td><p>RPM (Requests Per Minute)</p></td><td><p>60</p></td></tr><tr><td><p>TPM (Tokens Per Minute)</p></td><td><p>100,000</p></td></tr></tbody></table>
  </Tab>

  <Tab title="Singapore">
    Scope: International

    <table><thead><tr><th>Parameter</th><th>Value</th></tr></thead><tbody><tr><td><p>RPM (Requests Per Minute)</p></td><td><p>60</p></td></tr><tr><td><p>TPM (Tokens Per Minute)</p></td><td><p>100,000</p></td></tr></tbody></table>
  </Tab>
</Tabs>

## Dynamic Updates <span id="h-0b2405555b" />

### qwen-omni-turbo-latest <span id="h-dd2aed55a9" />

The new multimodal understanding and generation model, supporting text, image, speech, and video input as well as mixed input understanding. It can simultaneously stream text and speech, with significantly enhanced speed for multimodal content understanding. Additionally, it offers four natural dialogue tones. This model is a dynamically updated version.

#### Inference Service Provider <span id="h-4d6d1e0a8e" />

The inference service provider for `qwen-omni-turbo-latest` is Alibaba Cloud Model Studio.

#### Model Capabilities <span id="h-60c2b46379" />

<table><thead><tr><th>Capability</th><th>Support</th><th>Capability</th><th>Support</th></tr></thead><tbody><tr><td><p>Input Modality</p></td><td><p><strong>Text</strong> <strong>Image</strong> <strong>Video</strong> <strong>Audio</strong></p></td><td><p>Output Modality</p></td><td><p><strong>Text</strong> <strong>Audio</strong></p></td></tr><tr><td><p>Model Experience</p></td><td><p>Supported</p></td><td><p>Function Calling</p></td><td><p>Unsupported</p></td></tr><tr><td><p>Structured Outputs</p></td><td><p>Unsupported</p></td><td><p>Web Search</p></td><td><p>Unsupported</p></td></tr><tr><td><p>Prefix Completion</p></td><td><p>Unsupported</p></td><td><p>Context Caching</p></td><td><p>Unsupported</p></td></tr><tr><td><p>Batch Inference</p></td><td><p>Unsupported</p></td><td><p>Fine-tuning</p></td><td><p>Unsupported</p></td></tr></tbody></table>

#### Context Limits <span id="h-38531b9f25" />

<table><thead><tr><th>Parameter</th><th>Value</th><th>Parameter</th><th>Value</th></tr></thead><tbody><tr><td><p>Max Input Length</p></td><td><p>30720</p></td><td><p>Max Output Length</p></td><td><p>2048</p></td></tr><tr><td><p>Context Window</p></td><td><p>32768</p></td><td><p /></td><td><p /></td></tr></tbody></table>

#### Pricing <span id="h-40bb02cd0b" />

This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit [Model Studio Console](https://modelstudio.console.alibabacloud.com/ap-southeast-1/model/market) for promotional offers.

<Tabs>
  <Tab title="China (Beijing)">
    <table><thead><tr><th>Billing Item</th><th>Price (USD)</th><th>Unit</th></tr></thead><tbody><tr><td><p>Input: Text</p></td><td><p>0.058</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Input: Audio</p></td><td><p>3.584</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Input: Vision</p></td><td><p>0.216</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Output: Text (When input contains only text)</p></td><td><p>0.23</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Output: Text (When input contains images/audio/video)</p></td><td><p>0.646</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Output: Text\&Audio (Output text is not charged)</p></td><td><p>7.168</p></td><td><p>Per 1M tokens</p></td></tr></tbody></table>
  </Tab>

  <Tab title="Singapore">
    Scope: International

    <table><thead><tr><th>Billing Item</th><th>Price (USD)</th><th>Unit</th></tr></thead><tbody><tr><td><p>Input: Text</p></td><td><p>0.07</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Input: Audio</p></td><td><p>4.44</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Input: Vision</p></td><td><p>0.21</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Output: Text (When input contains only text)</p></td><td><p>0.27</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Output: Text (When input contains images/audio/video)</p></td><td><p>0.63</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Output: Text\&Audio (Output text is not charged)</p></td><td><p>8.89</p></td><td><p>Per 1M tokens</p></td></tr></tbody></table>
  </Tab>
</Tabs>

#### Rate Limits <span id="h-7c0ea751c3" />

<Tabs>
  <Tab title="China (Beijing)">
    <table><thead><tr><th>Parameter</th><th>Value</th></tr></thead><tbody><tr><td><p>RPM (Requests Per Minute)</p></td><td><p>60</p></td></tr><tr><td><p>TPM (Tokens Per Minute)</p></td><td><p>100,000</p></td></tr></tbody></table>
  </Tab>

  <Tab title="Singapore">
    Scope: International

    <table><thead><tr><th>Parameter</th><th>Value</th></tr></thead><tbody><tr><td><p>RPM (Requests Per Minute)</p></td><td><p>60</p></td></tr><tr><td><p>TPM (Tokens Per Minute)</p></td><td><p>100,000</p></td></tr></tbody></table>
  </Tab>
</Tabs>

## Snapshot Versions <span id="h-e8dce298bb" />

### qwen-omni-turbo-2025-03-26 <span id="h-5e903adcc3" />

The new multimodal understanding and generation model, supporting text, image, speech, and video input as well as mixed input understanding. It can simultaneously stream text and speech, with significantly enhanced speed for multimodal content understanding. Additionally, it offers four natural dialogue tones. This model is the snapshot from March 26, 2025, with significant improvements on visual capabilities over the snapshot of January 19, 2025.

#### Inference Service Provider <span id="h-8f58f70145" />

The inference service provider for `qwen-omni-turbo-2025-03-26` is Alibaba Cloud Model Studio.

#### Model Capabilities <span id="h-3d1b9c070f" />

<table><thead><tr><th>Capability</th><th>Support</th><th>Capability</th><th>Support</th></tr></thead><tbody><tr><td><p>Input Modality</p></td><td><p><strong>Text</strong> <strong>Image</strong> <strong>Video</strong> <strong>Audio</strong></p></td><td><p>Output Modality</p></td><td><p><strong>Text</strong> <strong>Audio</strong></p></td></tr><tr><td><p>Model Experience</p></td><td><p>Supported</p></td><td><p>Function Calling</p></td><td><p>Unsupported</p></td></tr><tr><td><p>Structured Outputs</p></td><td><p>Unsupported</p></td><td><p>Web Search</p></td><td><p>Unsupported</p></td></tr><tr><td><p>Prefix Completion</p></td><td><p>Unsupported</p></td><td><p>Context Caching</p></td><td><p>Unsupported</p></td></tr><tr><td><p>Batch Inference</p></td><td><p>Unsupported</p></td><td><p>Fine-tuning</p></td><td><p>Unsupported</p></td></tr></tbody></table>

#### Context Limits <span id="h-3ddb7ff741" />

<table><thead><tr><th>Parameter</th><th>Value</th><th>Parameter</th><th>Value</th></tr></thead><tbody><tr><td><p>Max Input Length</p></td><td><p>30720</p></td><td><p>Max Output Length</p></td><td><p>2048</p></td></tr><tr><td><p>Context Window</p></td><td><p>32768</p></td><td><p /></td><td><p /></td></tr></tbody></table>

#### Pricing <span id="h-a118b50b7a" />

This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit [Model Studio Console](https://modelstudio.console.alibabacloud.com/ap-southeast-1/model/market) for promotional offers.

<Tabs>
  <Tab title="China (Beijing)">
    <table><thead><tr><th>Billing Item</th><th>Price (USD)</th><th>Unit</th></tr></thead><tbody><tr><td><p>Input: Text</p></td><td><p>0.058</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Input: Audio</p></td><td><p>3.584</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Input: Vision</p></td><td><p>0.216</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Output: Text (When input contains only text)</p></td><td><p>0.23</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Output: Text (When input contains images/audio/video)</p></td><td><p>0.646</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Output: Text\&Audio (Output text is not charged)</p></td><td><p>7.168</p></td><td><p>Per 1M tokens</p></td></tr></tbody></table>
  </Tab>

  <Tab title="Singapore">
    Scope: International

    <table><thead><tr><th>Billing Item</th><th>Price (USD)</th><th>Unit</th></tr></thead><tbody><tr><td><p>Input: Text</p></td><td><p>0.07</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Input: Audio</p></td><td><p>4.44</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Input: Vision</p></td><td><p>0.21</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Output: Text (When input contains only text)</p></td><td><p>0.27</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Output: Text (When input contains images/audio/video)</p></td><td><p>0.63</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Output: Text\&Audio (Output text is not charged)</p></td><td><p>8.89</p></td><td><p>Per 1M tokens</p></td></tr></tbody></table>
  </Tab>
</Tabs>

#### Rate Limits <span id="h-849db9af93" />

<Tabs>
  <Tab title="China (Beijing)">
    <table><thead><tr><th>Parameter</th><th>Value</th></tr></thead><tbody><tr><td><p>RPM (Requests Per Minute)</p></td><td><p>60</p></td></tr><tr><td><p>TPM (Tokens Per Minute)</p></td><td><p>100,000</p></td></tr></tbody></table>
  </Tab>

  <Tab title="Singapore">
    Scope: International

    <table><thead><tr><th>Parameter</th><th>Value</th></tr></thead><tbody><tr><td><p>RPM (Requests Per Minute)</p></td><td><p>60</p></td></tr><tr><td><p>TPM (Tokens Per Minute)</p></td><td><p>100,000</p></td></tr></tbody></table>
  </Tab>
</Tabs>

### qwen-omni-turbo-2025-01-19 <span id="h-06ce7e20ef" />

Qwen All-Modal Understanding and Generation Large Model supports text, image, speech, video input understanding and mixed input comprehension, features simultaneous streaming generation of text and speech, significantly enhanced multi-modal content understanding speed, provides 4 natural conversational voices. This version is a snapshot from January 19, 2025, and will be maintained until approximately one month before the next snapshot release.

#### Inference Service Provider <span id="h-42c55eeff8" />

The inference service provider for `qwen-omni-turbo-2025-01-19` is Alibaba Cloud Model Studio.

#### Model Capabilities <span id="h-0efa27c8de" />

<table><thead><tr><th>Capability</th><th>Support</th><th>Capability</th><th>Support</th></tr></thead><tbody><tr><td><p>Input Modality</p></td><td><p><strong>Text</strong> <strong>Image</strong> <strong>Video</strong> <strong>Audio</strong></p></td><td><p>Output Modality</p></td><td><p><strong>Text</strong> <strong>Audio</strong></p></td></tr><tr><td><p>Model Experience</p></td><td><p>Supported</p></td><td><p>Function Calling</p></td><td><p>Unsupported</p></td></tr><tr><td><p>Structured Outputs</p></td><td><p>Unsupported</p></td><td><p>Web Search</p></td><td><p>Unsupported</p></td></tr><tr><td><p>Prefix Completion</p></td><td><p>Unsupported</p></td><td><p>Context Caching</p></td><td><p>Unsupported</p></td></tr><tr><td><p>Batch Inference</p></td><td><p>Unsupported</p></td><td><p>Fine-tuning</p></td><td><p>Unsupported</p></td></tr></tbody></table>

#### Context Limits <span id="h-57a1f04d10" />

<table><thead><tr><th>Parameter</th><th>Value</th><th>Parameter</th><th>Value</th></tr></thead><tbody><tr><td><p>Max Input Length</p></td><td><p>30720</p></td><td><p>Max Output Length</p></td><td><p>2048</p></td></tr><tr><td><p>Context Window</p></td><td><p>32768</p></td><td><p /></td><td><p /></td></tr></tbody></table>

#### Pricing <span id="h-4e8db46641" />

This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit [Model Studio Console](https://modelstudio.console.alibabacloud.com/ap-southeast-1/model/market) for promotional offers.

<Tabs>
  <Tab title="China (Beijing)">
    <table><thead><tr><th>Billing Item</th><th>Price (USD)</th><th>Unit</th></tr></thead><tbody><tr><td><p>Input: Text</p></td><td><p>0.058</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Input: Audio</p></td><td><p>3.584</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Input: Vision</p></td><td><p>0.216</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Output: Text (When input contains only text)</p></td><td><p>0.23</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Output: Text (When input contains images/audio/video)</p></td><td><p>0.646</p></td><td><p>Per 1M tokens</p></td></tr><tr><td><p>Output: Text\&Audio (Output text is not charged)</p></td><td><p>7.168</p></td><td><p>Per 1M tokens</p></td></tr></tbody></table>
  </Tab>
</Tabs>

#### Rate Limits <span id="h-05a87a8bce" />

<Tabs>
  <Tab title="China (Beijing)">
    <table><thead><tr><th>Parameter</th><th>Value</th></tr></thead><tbody><tr><td><p>RPM (Requests Per Minute)</p></td><td><p>60</p></td></tr><tr><td><p>TPM (Tokens Per Minute)</p></td><td><p>100,000</p></td></tr></tbody></table>
  </Tab>
</Tabs>