All Products
Search
Document Center

Alibaba Cloud Model Studio:Non-real-time speech synthesis

Last Updated:Sep 28, 2026

Non-real-time speech synthesis converts text to speech through the HTTP API. It is designed for latency-tolerant scenarios such as audiobook production, online education voiceovers, and content creation, and supports a wide range of voices, multiple languages, voice cloning, and voice design.

Overview

Convert complete text into audio files through the HTTP API. Two output modes are available: non-streaming and streaming.

  • Non-streaming returns an audio file URL valid for 24 hours; streaming returns audio data in chunks.
  • Multiple languages are supported, including Chinese dialects.
  • Supports Voice cloning and Voice Design for creating custom voices.
  • Supports Instruction control to control speech expressiveness through natural language instructions.

For low-latency streaming scenarios, see Real-time speech synthesis. For model selection recommendations, see Speech synthesis.

Audio synthesized on the voice design page in the Model Studio console can only be previewed online and cannot be downloaded as an audio file. To download the audio, call the API or the SDK. In non-streaming mode, the response returns an audio file URL valid for 24 hours.

Prerequisites

Before you begin, complete the following preparations:

Quick start

The following examples synthesize speech with qwen-audio-3.0-tts-flash and the system voice longanhuan_v3.6. For other available voices, see the voice list.

These examples use the Singapore region. Replace {WorkspaceId} with your workspace ID in this region, and set the DASHSCOPE_API_KEY environment variable to an API key from the same region.

Non-streaming output

In non-streaming mode, output.audio.url contains the download URL for the synthesized audio.

curl -X POST "https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1/services/audio/tts/SpeechSynthesizer" \
  -H "Authorization: Bearer $DASHSCOPE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen-audio-3.0-tts-flash",
    "input": {
        "text": "Today is a wonderful day.",
        "voice": "longanhuan_v3.6",
        "format": "wav",
        "sample_rate": 24000
    }
}'

Streaming output

Add the X-DashScope-SSE: enable request header to enable streaming output. The service returns audio chunks in SSE events. Each output.audio.data field contains Base64-encoded audio data.

curl --no-buffer -X POST "https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1/services/audio/tts/SpeechSynthesizer" \
  -H "Authorization: Bearer $DASHSCOPE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "X-DashScope-SSE: enable" \
  -d '{
    "model": "qwen-audio-3.0-tts-flash",
    "input": {
        "text": "Today is a wonderful day.",
        "voice": "longanhuan_v3.6",
        "format": "wav",
        "sample_rate": 24000
    }
}'

Advanced features

Instruction control

Instruction specifications by model:

Qwen-TTS

Supported models: Only Qwen3-TTS-Instruct-Flash series models are supported.

Usage: Pass the instruction content through the instructions parameter.

Supported languages for instruction text: Only Chinese and English are supported.

Instruction text length limit: Up to 1,600 tokens.

Dialects

This section describes how to generate speech in Chinese dialects (such as Henan dialect and Sichuan dialect). The configuration method varies by model and voice type.

Qwen-TTS

  • System voices: Use system voices that support dialects. See Qwen-TTS voice list.
  • Voice cloning voices: Dialects are not supported.
  • Voice design voices: Dialects are not supported.

Supported dialects: See the "Supported languages" section for each model in Qwen3-TTS.

Supported models and regions

Singapore

To call the following models, use an API key for the Singapore region:

  • Qwen-TTS:

    • Qwen3-TTS-Instruct-Flash: qwen3-tts-instruct-flash (stable version, currently equivalent to qwen3-tts-instruct-flash-2026-01-26), qwen3-tts-instruct-flash-2026-01-26 (latest snapshot)
    • Qwen3-TTS-VD: qwen3-tts-vd-2026-01-26 (latest snapshot)
    • Qwen3-TTS-VC: qwen3-tts-vc-2026-01-22 (latest snapshot)
    • Qwen3-TTS-Flash: qwen3-tts-flash (stable version, currently equivalent to qwen3-tts-flash-2025-11-27), qwen3-tts-flash-2025-11-27, qwen3-tts-flash-2025-09-18

China (Beijing)

To call the following models, use an API key for the Beijing region:

  • Qwen-TTS:

    • Qwen3-TTS-Instruct-Flash: qwen3-tts-instruct-flash (stable version, currently equivalent to qwen3-tts-instruct-flash-2026-01-26), qwen3-tts-instruct-flash-2026-01-26 (latest snapshot)
    • Qwen3-TTS-VD: qwen3-tts-vd-2026-01-26 (latest snapshot)
    • Qwen3-TTS-VC: qwen3-tts-vc-2026-01-22 (latest snapshot)
    • Qwen3-TTS-Flash: qwen3-tts-flash (stable version, currently equivalent to qwen3-tts-flash-2025-11-27), qwen3-tts-flash-2025-11-27, qwen3-tts-flash-2025-09-18
    • Qwen-TTS: qwen-tts (stable version, currently equivalent to qwen-tts-2025-04-10), qwen-tts-latest (latest version, currently equivalent to qwen-tts-2025-05-22), qwen-tts-2025-05-22 (snapshot), qwen-tts-2025-04-10 (snapshot)

Supported system voices

Different models support different voices. Set the voice request parameter to a value from the voice parameter column in the following tables.

API reference

FAQ

Q: How long is the audio file URL valid?

A: The audio file URL is valid for 24 hours after generation. After the URL expires, call the API again to obtain a new URL.