Qwen-Audio-3.1-TTS-Next is an AudioGen model designed for unified audio generation. Going beyond conventional TTS systems that focus solely on speech synthesis, it can generate complete audio content in a single pass from inputs such as text, timestamps, and reference audio, seamlessly combining speech, sound effects, ambient sounds, and more. The model supports a wide range of tasks, including multilingual TTS, single-speaker speech, multi-speaker dialogue, podcasts, cinematic soundscapes, ambient audio, and sound effects. Its key strengths include more natural and expressive speech, consistent voice identity, flexible timing control, and high-quality soundscape generation. These capabilities make it well suited for content creation, short-form video, podcasts, film and television production, game audio, and other professional audio creation scenarios.
Inference service provider
Alibaba Cloud Model Studio provides inference services for this model.
Model capabilities
| Capability | Support | Capability | Support |
|---|---|---|---|
| Input modalities | Text and reference audio | Output modality | Audio |
| Languages | Chinese and English | Output modes | Non-streaming |
| Output formats | WAV, MP3, and PCM | Reference audio | URL or Base64 |
Context limits
| Parameter | Value |
|---|---|
| Maximum input length | 3,000 characters |
| Maximum output duration per request | Podcasts: 240 seconds (4 minutes); other scenarios: 120 seconds |
| Reference audio count | Up to 3 clips |
| Duration per reference clip | Up to 30 seconds |
| Size per reference clip | Up to 10 MB |
| Reference audio formats | WAV, MP3, and OGG Opus. Raw PCM is not supported. |
Model pricing
The following list price excludes promotions. For promotional offers, see the Model Studio console.
China (Beijing)
| Billing item | Unit price (USD/million tokens) |
|---|---|
| Input | 0.848 |
| Output | 1.696 |
Input and output tokens are billed separately. For details, see Model pricing.
Rate limits
China (Beijing)
| Metric | Value |
|---|---|
| Requests per second (RPS) | 3 |
API access
Use the HTTPS API to receive complete audio files. For examples and prompt guidance, see Audio generation. For parameters and sample code, see Audio generation API reference.