Choose the right model for speech synthesis, voice cloning, and voice design.
Migrating from proprietary models?
If you're currently using ElevenLabs, OpenAI, or Google for speech synthesis, use the following table to find the equivalent Model Studio model.
|
Use case |
Proprietary model |
Model Studio equivalent |
|
Built-in voices / standard synthesis |
OpenAI gpt-4o-tts, Google Chirp 3 |
|
|
Custom voice / voice cloning |
ElevenLabs Multilingual v3 |
|
Standard TTS or custom voice?
Speech synthesis models convert text into natural-sounding speech. Start by deciding whether built-in voices or custom voices fit your needs:
|
Standard TTS |
Custom voice |
|
|
Voice source |
Preset voice library — select and use immediately |
Clone from an audio sample, or create with a text description |
|
Setup effort |
No extra configuration — pick a model and voice to start |
Requires an audio sample or text description to create a voice |
|
Use cases |
Customer service bots, audiobooks, news broadcasts, e-commerce livestreams |
Brand-specific voices, virtual anchors, game character dubbing |
|
Recommended models |
|
|
-
Use standard TTS when the preset voice library meets your needs and you want quick setup with no extra configuration.
-
Use a custom voice when you need a brand-specific voice, want to replicate a specific person's voice, or need to create an entirely new character voice.
Voice cloning or voice design?
If you choose a custom voice, two creation methods are available:
|
Voice cloning |
Voice design |
|
|
Input |
An audio sample of the target speaker |
A text description of the desired voice (e.g., "warm low-pitched female voice") |
|
Result |
Synthesized speech closely resembles the original speaker |
A brand-new voice generated from scratch based on the description |
|
Use cases |
Brand spokesperson/anchor voice reuse, virtual anchors, personalized voice assistants |
Brand voice design (no recordings available), game/animation character dubbing, creative content production |
|
Recommended models |
|
|
|
Voice management service |
|
|
-
Use voice cloning when you have a recording of the target speaker and want to reproduce that voice in synthesized speech.
-
Use voice design when you don't have a recording and want to create a new voice from a text description.
WebSocket or HTTP?
-
WebSocket: Bidirectional streaming — supports streaming input and output, delivering audio as it's synthesized to minimize latency. Ideal for customer service bots, voice assistants, and call centers.
-
HTTP: Send complete text, receive audio in a streaming response (chunked output). Ideal for audiobooks, audio content production, and other high-volume batch scenarios.
Qwen-Audio-TTS/CosyVoice models use the same model name for both WebSocket and HTTP access. For Qwen models, the model name indicates the API: models with the -realtime suffix use WebSocket, while those without it use HTTP.
Access Qwen-Audio-TTS/CosyVoice and Qwen WebSocket models through the DashScope SDK (Java, Python). Other models require direct WebSocket or HTTP calls.
For WebSocket access, see Real-time speech synthesis. For HTTP access, see Non-real-time speech synthesis.
Instruction control
Describe the desired expression style in natural language to control speed, emotion, and style per request. Examples include "speak gently at a slower pace" or "use an excited broadcast style." Use instruction control for emotional content production, professional broadcasting, audiobooks, and other scenarios that require rich expressiveness.
Models that support instruction control: Qwen-Audio-TTS series (qwen-audio-3.0-tts-plus, qwen-audio-3.0-tts-flash), CosyVoice series (cosyvoice-v3.5-plus, cosyvoice-v3.5-flash, cosyvoice-v3-flash), and Qwen-TTS series (qwen3-tts-instruct-flash-realtime, qwen3-tts-instruct-flash). For details, see Real-time speech synthesis > Instruction control.
Recommended models
The following table lists the best model for each scenario. For full details, visit the Model Gallery.
|
Model ID |
Series |
API |
Voice cloning |
Voice design |
Instruction control |
|
|
Qwen-Audio-TTS |
WebSocket |
|
|
|
|
|
CosyVoice |
WebSocket |
|
|
|
|
|
CosyVoice |
WebSocket |
|
|
|
All models
Qwen-Audio-TTS
|
Model ID |
API |
Voice cloning |
Voice design |
Instruction control |
|
|
WebSocket |
|
|
|
|
|
WebSocket |
|
|
|
Supported languages:
-
System voices (varies by voice): Chinese (Mandarin), English
-
Cloned voices (dialects are configured through instruction control): Chinese (Mandarin, Cantonese, Chongqing, Northeastern, Gansu, Guizhou, Zhejiang, Hebei, Henan, Hubei, Hunan, Jiangxi, Ningbo, Ningxia, Qingdao, Shaanxi, Shanxi, Shandong, Shanghainese, Sichuan, Yunnan), English, Japanese, Korean, German, French, Italian, Russian, Portuguese, Thai, Indonesian, Malay, and Vietnamese
CosyVoice
Some CosyVoice models support SSML (Speech Synthesis Markup Language) markup and LaTeX formula reading.
|
Model ID |
API |
Voice cloning |
Voice design |
Instruction control |
|
|
WebSocket |
|
|
|
|
|
WebSocket |
|
|
|
|
|
WebSocket |
|
|
|
|
|
WebSocket |
|
|
|
|
|
WebSocket |
|
|
|
Supported languages (by version):
-
cosyvoice-v3.5-plus and cosyvoice-v3.5-flash (no system voices):
-
Cloned voices (dialects are configured through instruction control): Chinese (Mandarin, Cantonese, Northeastern, Gansu, Guizhou, Henan, Hubei, Jiangxi, Hokkien, Ningxia, Shanxi, Shaanxi, Shandong, Shanghainese, Sichuan, Tianjin, Yunnan), English, French, German, Japanese, Korean, Russian, Portuguese, Thai, Indonesian, and Vietnamese
-
Designed voices: Chinese (Mandarin), English
-
-
cosyvoice-v3-plus:
-
System voices (varies by voice): Chinese (Mandarin), English
-
Cloned voices (dialects are configured through instruction control): Chinese (Mandarin, Cantonese, Northeastern, Gansu, Guizhou, Henan, Hubei, Jiangxi, Hokkien, Ningxia, Shanxi, Shaanxi, Shandong, Shanghainese, Sichuan, Tianjin, Yunnan), English, French, German, Japanese, Korean, and Russian
-
Designed voices: Chinese (Mandarin), English
-
-
cosyvoice-v3-flash:
-
System voices (varies by voice; some dialects are directly supported by system voices, others through instruction control): Chinese (Mandarin, Cantonese, Northeastern, Henan, Hunan, Shaanxi, Shandong, Sichuan, Anhui, Hokkien), English
-
Cloned voices (dialects are configured through instruction control): Chinese (Mandarin, Cantonese, Northeastern, Gansu, Guizhou, Henan, Hubei, Jiangxi, Hokkien, Ningxia, Shanxi, Shaanxi, Shandong, Shanghainese, Sichuan, Tianjin, Yunnan), English, French, German, Japanese, Korean, Russian, Portuguese, Thai, Indonesian, and Vietnamese
-
Designed voices: Chinese (Mandarin), English
-
-
cosyvoice-v2 (voice design not supported):
-
System voices (varies by voice): Chinese (Mandarin, Cantonese, Northeastern, Hokkien, Shaanxi), English, Japanese, and Korean
-
Cloned voices: Chinese (Mandarin), English
-
Qwen3-TTS
|
Model ID |
API |
Voice cloning |
Voice design |
Instruction control |
|
|
HTTP |
|
|
|
|
|
HTTP |
|
|
|
|
|
HTTP |
|
|
|
|
|
WebSocket |
|
|
|
|
|
WebSocket |
|
|
|
|
|
WebSocket |
|
|
|
|
|
HTTP |
|
|
|
|
|
HTTP |
|
|
|
|
|
WebSocket |
|
|
|
|
|
WebSocket |
|
|
|
|
|
HTTP |
|
|
|
|
|
WebSocket |
|
|
|
|
|
WebSocket |
|
|
|
|
|
HTTP |
|
|
|
|
|
WebSocket |
|
|
|
|
|
WebSocket |
|
|
|
Supported languages (by version):
-
Qwen3-TTS-Flash series (system voices) (
qwen3-tts-flash,qwen3-tts-flash-2025-11-27,qwen3-tts-flash-2025-09-18,qwen3-tts-flash-realtime,qwen3-tts-flash-realtime-2025-11-27,qwen3-tts-flash-realtime-2025-09-18): Chinese (Mandarin, Beijing, Shanghainese, Sichuan, Nanjing, Shaanxi, Hokkien, Tianjin, Cantonese — varies by voice), English, German, Italian, Portuguese, Spanish, Japanese, Korean, French, and Russian -
Qwen3-TTS-Instruct-Flash series (system voices) (
qwen3-tts-instruct-flash,qwen3-tts-instruct-flash-2026-01-26,qwen3-tts-instruct-flash-realtime,qwen3-tts-instruct-flash-realtime-2026-01-22): Chinese (Mandarin), English, German, Italian, Portuguese, Spanish, Japanese, Korean, French, and Russian -
Qwen3-TTS-VC series (voice cloning) (
qwen3-tts-vc-2026-01-22,qwen3-tts-vc-realtime-2026-01-15,qwen3-tts-vc-realtime-2025-11-27): Chinese (Mandarin), English, German, Italian, Portuguese, Spanish, Japanese, Korean, French, and Russian -
Qwen3-TTS-VD series (voice design) (
qwen3-tts-vd-2026-01-26,qwen3-tts-vd-realtime-2026-01-15,qwen3-tts-vd-realtime-2025-12-16): Chinese (Mandarin), English, German, Italian, Portuguese, Spanish, Japanese, Korean, French, and Russian
Qwen-TTS (legacy, token-based billing)
The following legacy Qwen-TTS models use token-based billing. If you've migrated to Qwen3-TTS, use the models recommended earlier in this topic.
|
Model ID |
API |
Description |
|
|
HTTP |
Non-streaming synthesis, token-based billing |
|
|
HTTP |
Non-streaming synthesis, token-based billing |
|
|
HTTP |
Snapshot version, token-based billing |
|
|
HTTP |
Snapshot version, token-based billing |
|
|
WebSocket |
Streaming synthesis, token-based billing |
|
|
WebSocket |
Streaming synthesis, token-based billing |
|
|
WebSocket |
Snapshot version, streaming synthesis, token-based billing |
Supported languages (by version):
-
Qwen-TTS series (system voices) (
qwen-tts,qwen-tts-latest,qwen-tts-2025-05-22,qwen-tts-2025-04-10): Chinese (Mandarin, Beijing, Shanghainese, Sichuan — varies by voice), English, German, Italian, Portuguese, Spanish, Japanese, Korean, French, and Russian -
Qwen-TTS-Realtime series (system voices) (
qwen-tts-realtime,qwen-tts-realtime-latest,qwen-tts-realtime-2025-07-15): Chinese (Mandarin), English, German, Italian, Portuguese, Spanish, Japanese, Korean, French, and Russian