All Products
Search
Document Center

Alibaba Cloud Model Studio:qwen3.8-omni-flash

Last Updated:Sep 18, 2026

A model for audio and video understanding and content analysis, with text, image, audio, and video input and text output.

Inference service provider

The inference service provider for qwen3.8-omni-flash is Alibaba Cloud Model Studio.

Model capabilities

Supported regions: China (Beijing), Singapore, China (Hong Kong), Japan (Tokyo), Germany (Frankfurt), and US (Virginia). Use an API key from the selected region.

CapabilitySupportCapabilitySupport
Input modalitiesText, images, audio, and videoOutput modalitiesText
Function CallingSupports custom tool callingThinkingEnabled by default, with adjustable reasoning effort. See thinking
Web searchSupports web search; the built-in Responses tool is web_searchContext cachingSupports automatic implicit caching and Responses Session caching
Audio input languages113 languages and dialects, the same as Qwen3.5-Omni. See the full list in Model selectionMultichannel audioSupports spatial audio input; use use_multichannel to enable it in Chat Completions

Use Chat Completions or Responses. See the non-real-time guide for examples.

Context limits

ParameterValueParameterValue
Context window1M tokensMaximum input length (non-thinking mode)991808 tokens
Maximum input length (thinking mode)983616 tokensMaximum output length131072 tokens

Model pricing

See Model pricing for inference prices.

Rate limits

See Rate limits for model request limits.