text string(required)
The text to synthesize.
SSML and LaTeX format inputs are supported. Replace the text content with the corresponding format.
voice string(required)
The voice for synthesis.
Valid values:
format string(optional)
The audio format.
Default value: mp3.
Valid values:
sample_rate integer(optional)
The audio sample rate in Hz.
Valid values: 8000, 16000, 22050 (default), 24000, 44100, 48000.
volume integer(optional)
The volume level.
Default value: 50.
Valid values: [0, 100].
rate float(optional)
The speech rate.
Default value: 1.0.
Valid values: [0.5, 2.0].
bit_rate integer(optional)
The audio bitrate in kbps.
Default value: 32.
Valid values: [6, 510].
ImportantThis parameter is supported only when format is set to opus.
pitch float(optional)
The pitch.
Default value: 1.0.
Valid values: [0.5, 2.0].
enable_ssml boolean(optional)
Specifies whether to enable SSML. For the SSML usage restrictions (supported models, voices, and APIs), see Limitations.
word_timestamp_enabled boolean(optional)
Specifies whether to enable word-level timestamps.
Default value: false.
Available only in streaming output mode. Supported voices: cloned voices of qwen-audio-3.0-tts-plus, qwen-audio-3.1-tts-flash, qwen-audio-3.0-tts-flash, cosyvoice-v3.5-plus, cosyvoice-v3.5-flash, cosyvoice-v3-flash, cosyvoice-v3-plus, and cosyvoice-v2, and system voices marked as supported in Qwen-Audio-TTS voice list, CosyVoice Voice list. Cloned voices of other models do not support this feature.
seed integer(optional)
The random seed for synthesis. Different seeds produce different outputs. With identical model version, text, voice, and other parameters, the same seed reproduces the same output.
Default value: 0.
Valid values: [0, 65535].
language_hints array[string](optional)
Important
- This parameter is an array, but the current version processes only the first element. Pass a single value.
- This parameter specifies the target language for speech synthesis and is unrelated to the language of the sample audio used in voice cloning. To set the source language for a voice cloning task, see the Voice Cloning API reference.
Specifies the target language for speech synthesis to improve output quality.
Use this parameter when numbers, abbreviations, or symbols are not pronounced as expected, or when synthesis quality for secondary languages is poor. Examples:
- A number isn't read as expected: "hello, this is 110" is read as "hello, this is one one zero" instead of "hello, this is yao-yao-ling"
- A symbol isn't read correctly: "@" is read as "ai-te" instead of "at"
- Minor language synthesis sounds unnatural
Valid values:
- zh: Chinese
- en: English
- fr: French
- de: German
- ja: Japanese
- ko: Korean
- ru: Russian
- pt: Portuguese
- th: Thai
- id: Indonesian
- vi: Vietnamese
- es: Spanish
- it: Italian
- ms: Malaysian
- fil: Filipino
- ar: Arabic
instruction string(optional)
An instruction that controls the synthesis behavior, such as dialect, emotion, or role.
For usage details, see Non-real-time speech synthesis.
enable_aigc_tag boolean(optional)
Specifies whether to embed an AIGC watermark in the generated audio. When set to true, the watermark is embedded in audio files of supported formats (wav/mp3/opus).
Default value: false.
Only qwen-audio-3.0-tts-plus, qwen-audio-3.1-tts-flash, qwen-audio-3.0-tts-flash, cosyvoice-v3-flash, cosyvoice-v3-plus, and cosyvoice-v2 support this feature.
aigc_propagator string(optional)
Sets the ContentPropagator field in the AIGC watermark, identifying the content propagator. Takes effect only when enable_aigc_tag is true.
Default value: Alibaba Cloud UID.
Only qwen-audio-3.0-tts-plus, qwen-audio-3.1-tts-flash, qwen-audio-3.0-tts-flash, cosyvoice-v3-flash, cosyvoice-v3-plus, and cosyvoice-v2 support this feature.
aigc_propagate_id string(optional)
Sets the PropagateID field in the AIGC watermark, uniquely identifying a specific propagation action. Takes effect only when enable_aigc_tag is true.
Default value: The request ID of the current speech synthesis request.
Only qwen-audio-3.0-tts-plus, qwen-audio-3.1-tts-flash, qwen-audio-3.0-tts-flash, cosyvoice-v3-flash, cosyvoice-v3-plus, and cosyvoice-v2 support this feature.
hot_fix object(optional)
Text hot-fix configuration. Customizes the pronunciation of specified words or substitutes text before speech synthesis.
Not supported by cosyvoice-v2.
Parameters:
- pronunciation: Custom pronunciation. Provides pinyin for given words to correct cases where the default pronunciation is inaccurate.
- replace: Text substitution. Replaces specified words with target text before speech synthesis; the substituted text is what gets synthesized.
Example:
"hot_fix": {
"pronunciation": [
{"weather": "tian1 qi4"}
],
"replace": [
{"today": "gold day"}
]
}
enable_markdown_filter boolean(optional)
ImportantOnly cloned voices of cosyvoice-v3-flash support this feature.
Specifies whether to enable Markdown filtering. When enabled, the system automatically strips Markdown markup symbols from the input text before synthesis, preventing them from being read aloud.
Default value: false.
Valid values:
- true: Enable Markdown filtering
- false: Disable Markdown filtering