A real-time duplex speech model with text and audio input and output, function calling, web search, and voice cloning.
The inference service provider is Alibaba Cloud Model Studio. The default voice is longanqian_v3.1. For integration, supported voices, and parameters, see the real-time speech guide.
Enable web search with enable_search. Web search and Function Calling cannot be enabled together.
Context limits
| Parameter | Limit (tokens) |
|---|---|
| Context window | 262,144 |
| Max input length | 245,760 |
| Max output length | 16,384 |
Pricing and rate limits
Singapore
Prices are per million tokens.
| Model ID | Text input | Audio input | Text output | Audio output |
|---|---|---|---|---|
qwen-audio-3.1-realtime-plus | $0.8 | $6.4 | $6.4 | $24 |
Rate limits: 60 RPM and 100,000 TPM.
China (Beijing)
Prices are per million tokens.
| Model ID | Text input | Audio input | Text output | Audio output |
|---|---|---|---|---|
qwen-audio-3.1-realtime-plus | $0.688 | $5.501 | $5.501 | $20.628 |
Rate limits: 60 RPM and 100,000 TPM.