text_type string(required)
Set to PlainText.
voice string(required)
The voice used for speech synthesis.
- System voices: See Qwen-Audio-TTS voice list
- Cloned voices: Custom voices created through voice cloning
- Custom voices: Custom voices created through voice design
format string(optional)
The audio encoding format.
Valid values:
- pcm
- wav
- mp3 (default)
- opus
sample_rate integer(optional)
The audio sample rate in Hz.
Valid values: 8000, 16000, 22050 (default), 24000, 44100, 48000.
volume integer(optional)
The volume level.
Default value: 50.
Valid values: [0, 100].
rate float(optional)
The speech rate.
Default value: 1.0.
Valid values: [0.5, 2.0].
pitch float(optional)
The pitch.
Default value: 1.0.
Valid values: [0.5, 2.0].
bit_rate integer(optional)
The audio bit rate in kbps. When the audio format is mp3 or opus, use bit_rate to adjust the bit rate.
Default value: 32.
Valid values: [6, 510].
enable_ssml boolean(optional)
Specifies whether to enable SSML.
Default value: false.
When set to true, only one continue-task command is allowed.
For the SSML usage restrictions (supported models, voices, and APIs), see Limitations.
word_timestamp_enabled boolean(optional)
Specifies whether to enable word-level timestamps.
Default value: false.
Available only in streaming output mode. Cloned voices are supported. For supported system voices, see Qwen-Audio-TTS voice list.
seed integer(optional)
A random seed for controlling variation in the synthesis output. When the model version, text, voice, and other parameters are unchanged, using the same seed produces identical results.
Default value: 0.
Valid values: [0, 65535].
language_hints array[string](optional)
Important
- This parameter is an array, but the current version only processes the first element. Pass a single value.
- This parameter specifies the target language for speech synthesis. It's unrelated to the language of the audio sample used in voice cloning. To set the source language for a cloning task, see the voice cloning API reference.
Specifies the target language for speech synthesis to improve output quality.
When digit pronunciation, abbreviation expansion, symbol reading, or minority-language synthesis doesn't meet expectations, use this parameter. For example:
- Unexpected digit pronunciation: "hello, this is 110" is read as "hello, this is one zero" instead of the expected Chinese pronunciation
- Inaccurate symbol pronunciation: "@" is read as the Chinese equivalent instead of "at"
- Poor minor language synthesis quality with unnatural results
Valid values
- zh: Chinese
- en: English
- fr: French
- de: German
- ja: Japanese
- ko: Korean
- ru: Russian
- pt: Portuguese
- th: Thai
- id: Indonesian
- vi: Vietnamese
- es: Spanish
- it: Italian
- ms: Malaysian
- fil: Filipino
- ar: Arabic
instruction string(optional)
Sets an instruction to control dialect, emotion, or voice character during synthesis. For detailed usage, see Instruction control.
enable_aigc_tag boolean(optional)
Specifies whether to embed an AIGC watermark in the generated audio. When set to true, the watermark is embedded in audio files of supported formats (wav/mp3/opus).
Default value: false.
aigc_propagator string(optional)
Sets the ContentPropagator field in the AIGC watermark, identifying the content propagator. Takes effect only when enable_aigc_tag is true.
Default value: Alibaba Cloud UID.
aigc_propagate_id string(optional)
Sets the PropagateID field in the AIGC watermark, uniquely identifying a specific propagation action. Takes effect only when enable_aigc_tag is true.
Default value: The request ID of the current speech synthesis request.
hot_fix object(optional)
Configures pronunciation corrections and text replacements applied before synthesis.
Parameters:
- pronunciation: Custom pronunciation. Specifies pinyin annotations for words to correct inaccurate default pronunciations.
- replace: Text replacement. Replaces specified words with target text before synthesis. The replaced text is used as the actual synthesis input.
Example:
"hot_fix": {
"pronunciation": [
{"weather": "tian1 qi4"}
],
"replace": [
{"today": "gold day"}
]
}