A model for audio and video understanding and content analysis, with text, image, audio, and video input and text output.
Inference service provider
The inference service provider for qwen3.8-omni-flash is Alibaba Cloud Model Studio.
Model capabilities
Supported regions: China (Beijing), Singapore, China (Hong Kong), Japan (Tokyo), Germany (Frankfurt), and US (Virginia). Use an API key from the selected region.
| Capability | Support | Capability | Support |
|---|---|---|---|
| Input modalities | Text, images, audio, and video | Output modalities | Text |
| Function Calling | Supports custom tool calling | Thinking | Enabled by default, with adjustable reasoning effort. See thinking |
| Web search | Supports web search; the built-in Responses tool is web_search | Context caching | Supports automatic implicit caching and Responses Session caching |
| Audio input languages | 113 languages and dialects, the same as Qwen3.5-Omni. See the full list in Model selection | Multichannel audio | Supports spatial audio input; use use_multichannel to enable it in Chat Completions |
Use Chat Completions or Responses. See the non-real-time guide for examples.
Context limits
| Parameter | Value | Parameter | Value |
|---|---|---|---|
| Context window | 1M tokens | Maximum input length (non-thinking mode) | 991808 tokens |
| Maximum input length (thinking mode) | 983616 tokens | Maximum output length | 131072 tokens |
Model pricing
See Model pricing for inference prices.
Rate limits
See Rate limits for model request limits.