Non-real-time speech synthesis

Updated at:

Non-real-time speech synthesis converts text to speech through the HTTP API. It is designed for latency-tolerant scenarios such as audiobook production, online education voiceovers, and content creation, and supports a wide range of voices, multiple languages, voice cloning, and voice design.

Overview

Convert complete text into audio files through the HTTP API. Two output modes are available: non-streaming and streaming.

  • Non-streaming returns an audio file URL valid for 24 hours; streaming returns audio data in chunks.
  • Multiple languages are supported, including Chinese dialects.
  • Supports Voice cloning and Voice Design for creating custom voices.
  • Supports Instruction control to control speech expressiveness through natural language instructions.
  • Supports Emotion and rich language tags to embed tags in text for controlling emotional expression or inserting sound effects

NoteThe API endpoints differ by model series: Qwen-Audio-TTS and CosyVoice use https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/api/v1/services/audio/tts/SpeechSynthesizer; Qwen-TTS uses https://dashscope.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation; MiniMax uses https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation. The endpoints are not interchangeable. Use the endpoint shown in the examples of each model series.

For low-latency streaming scenarios, see Real-time speech synthesis. For model selection recommendations, see Speech synthesis.

Audio synthesized on the voice design page in the Model Studio console can only be previewed online and cannot be downloaded as an audio file. To download the audio, call the API or the SDK. In non-streaming mode, the response returns an audio file URL valid for 24 hours.

Prerequisites

Before you begin, complete the following preparations:

Quick start

The following tabs demonstrate speech synthesis for each model series. For more language examples and detailed parameter descriptions, see API reference.

Qwen-Audio-TTS

The following examples demonstrate how to synthesize speech with Qwen-Audio-TTS models.

ImportantQwen-Audio-TTS non-real-time speech synthesis is available only in the China (Beijing) region.

Non-streaming output

In non-streaming mode, the response contains a URL to the synthesized audio file. The URL is valid for 24 hours.

curl -X POST https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/api/v1/services/audio/tts/SpeechSynthesizer \
-H "Authorization: Bearer $DASHSCOPE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
    "model": "qwen-audio-3.0-tts-flash",
    "input": {
      "text": "There is a large garden behind my house.",
      "voice": "longanhuan_v3.6",
      "format": "wav",
      "sample_rate": 24000
    }
}'

Streaming output

Add the X-DashScope-SSE: enable header to enable streaming output. The server returns audio data incrementally using Server-Sent Events (SSE).

curl -X POST https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/api/v1/services/audio/tts/SpeechSynthesizer \
-H "Authorization: Bearer $DASHSCOPE_API_KEY" \
-H "Content-Type: application/json" \
-H "X-DashScope-SSE: enable" \
-d '{
    "model": "qwen-audio-3.0-tts-flash",
    "input": {
      "text": "There is a large garden behind my house.",
      "voice": "longanhuan_v3.6",
      "format": "wav",
      "sample_rate": 24000
    }
}'

CosyVoice

The following examples demonstrate how to synthesize speech with CosyVoice models.

ImportantCosyVoice non-real-time speech synthesis is available only in the China (Beijing) region.

Non-streaming output

In non-streaming mode, the response contains a URL to the synthesized audio file. The URL is valid for 24 hours.

curl -X POST https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/api/v1/services/audio/tts/SpeechSynthesizer \
-H "Authorization: Bearer $DASHSCOPE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
    "model": "cosyvoice-v3-flash",
    "input": {
      "text": "There is a large garden behind my house.",
      "voice": "longanyang",
      "format": "wav",
      "sample_rate": 24000
    }
}'

Streaming output

Add the X-DashScope-SSE: enable header to enable streaming output. The server returns audio data incrementally using Server-Sent Events (SSE).

curl -X POST https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/api/v1/services/audio/tts/SpeechSynthesizer \
-H "Authorization: Bearer $DASHSCOPE_API_KEY" \
-H "Content-Type: application/json" \
-H "X-DashScope-SSE: enable" \
-d '{
    "model": "cosyvoice-v3-flash",
    "input": {
      "text": "There is a large garden behind my house.",
      "voice": "longanyang",
      "format": "wav",
      "sample_rate": 24000
    }
}'

MiniMax

MiniMax supports emotion control, speed adjustment, and pitch adjustment.

ImportantMiniMax non-real-time speech synthesis is available only in the China (Beijing) region.

Non-streaming output

In non-streaming mode, the response contains the complete synthesized audio.

curl -X POST "https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation" \
-H "Authorization: Bearer $DASHSCOPE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
  "model": "MiniMax/speech-2.8-hd",
  "input": {
    "text": "The weather is great today, perfect for a walk.",
    "voice_setting": {
      "voice_id": "male-qn-qingse",
      "speed": 1,
      "vol": 1,
      "pitch": 0,
      "emotion": "happy"
    },
    "audio_setting": {
      "sample_rate": 32000,
      "bitrate": 128000,
      "format": "mp3",
      "channel": 1
    }
  }
}'

Streaming output

Add the X-DashScope-SSE: enable header to enable streaming output.

# Get an API Key: https://help.aliyun.com/zh/model-studio/get-api-key

curl -X POST "https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation" \
-H "Authorization: Bearer $DASHSCOPE_API_KEY" \
-H "Content-Type: application/json" \
-H "X-DashScope-SSE: enable" \
-d '{
  "model": "MiniMax/speech-2.8-hd",
  "input": {
    "text": "The weather is great today, perfect for a walk.",
    "voice_setting": {
      "voice_id": "male-qn-qingse",
      "speed": 1,
      "vol": 1,
      "pitch": 0,
      "emotion": "happy"
    },
    "audio_setting": {
      "sample_rate": 32000,
      "bitrate": 128000,
      "format": "mp3",
      "channel": 1
    }
  }
}'

Advanced features

Instruction control

Instruction specifications by model:

NoteThe instruction parameter name differs by model series: CosyVoice uses instruction, and Qwen-TTS uses instructions. Update the parameter name when migrating across model series.

Qwen-Audio-TTS

Supported models: qwen-audio-3.1-tts-flash, qwen-audio-3.0-tts-plus, qwen-audio-3.0-tts-flash

System voices and voice cloning voices: accept any instruction.

CosyVoice

Supported models: cosyvoice-v3.5-plus, cosyvoice-v3.5-flash, cosyvoice-v3-plus, cosyvoice-v3-flash

Instruction format requirements vary by model:

  • cosyvoice-v3.5-plus, cosyvoice-v3.5-flash:

    • Voice cloning/design voices: accept any instruction.
    • System voices: v3.5 doesn't support system voices.
  • cosyvoice-v3-plus:

    • Voice cloning/design voices: don't support instruction control.
    • System voices: instructions must use a fixed format and content. See CosyVoice Voice list.
  • cosyvoice-v3-flash:

    • Voice cloning/design voices: accept any instruction.
    • System voices: instructions must use a fixed format and content. See CosyVoice Voice list.

Usage: Specify instruction content through the instruction parameter.

Supported languages for instruction text:

  • cosyvoice-v3.5-plus, cosyvoice-v3.5-flash:

    • Voice cloning/design voices: Chinese, English, French, German, Japanese, Korean, Russian, Portuguese, Thai, Indonesian, and Vietnamese.
    • System voices: v3.5 doesn't support system voices.
  • cosyvoice-v3-plus:

    • Voice cloning/design voices: Chinese, English, French, German, Japanese, Korean, and Russian.
    • System voices: instructions must use a fixed format and content. See CosyVoice Voice list.
  • cosyvoice-v3-flash:

    • Voice cloning/design voices: Chinese, English, French, German, Japanese, Korean, and Russian.
    • System voices: Chinese only.

Instruction text length limit: 100 characters maximum. Chinese characters (including simplified/traditional Chinese, Japanese kanji, and Korean hanja) count as 2 characters each. All other characters (such as punctuation, letters, digits, Japanese kana, and Korean hangul) count as 1 character each.

Qwen-TTS

Supported models: Only Qwen3-TTS-Instruct-Flash series models are supported.

Usage: Pass the instruction content through the instructions parameter.

Supported languages for instruction text: Only Chinese and English are supported.

Instruction text length limit: Up to 1,600 tokens.

Dialects

This section describes how to generate speech in Chinese dialects (such as Henan dialect and Sichuan dialect). The configuration method varies by model and voice type.

Qwen-Audio-TTS

  • System voices: Choose one of the following voice types:

    • System voices that support dialects natively — no additional configuration required.
    • Voices that support Instruction control and allow dialect specification — specify the dialect through instruction text.
  • Voice cloning voices: Configure through the Instruction control feature. For example, set the instruction text to 请用河南话表达.

  • Voice design voices: Dialects are not supported.

Supported dialects: See the "Supported languages" section for each model in Qwen-Audio-TTS.

CosyVoice

  • System voices: Select one of the following voice types from the CosyVoice Voice list:

    • System voices that support dialects (for example, longshange_v3) — no additional configuration required.
    • Voices that support Instruction control and allow dialect specification (for example, longanhuan_v3) — specify the dialect through instruction text.
  • Voice cloning voices: Configure through the Instruction control feature. For example, set the instruction text to 请用河南话表达.

  • Voice design voices: Dialects are not supported.

Supported dialects: See the "Supported languages" section for each model in CosyVoice.

Example: Using cosyvoice-v3-flash with the longanhuan_v3 voice and the instruction text "请用河南话表达。" to generate speech in Henan dialect.

curl -X POST https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/api/v1/services/audio/tts/SpeechSynthesizer \
-H "Authorization: Bearer $DASHSCOPE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
    "model": "cosyvoice-v3-flash",
    "input": {
      "text": "叫你去买盐,你买回来一袋面,这不是弄啥嘞吗!",
      "voice": "longanhuan_v3",
      "format": "wav",
      "sample_rate": 24000,
      "instruction": "请用河南话表达。"
    }
}'

NoteThe instruction parameter name here is specific to CosyVoice. For Qwen-TTS, the parameter name is instructions. Do not mix them up.

Qwen-TTS

  • System voices: Use system voices that support dialects. See Qwen-TTS voice list.
  • Voice cloning voices: Dialects are not supported.
  • Voice design voices: Dialects are not supported.

Supported dialects: See the "Supported languages" section for each model in Qwen3-TTS.

Emotion and rich language tags

Qwen-Audio-TTS series models support embedding emotion and rich language tags directly in the text to synthesize (the text parameter). These tags control emotional expression or insert vocal effects (such as laughter and sighs) at specified positions, producing more expressive speech without configuring complex audio parameters.

ImportantSupported models: Only qwen-audio-3.1-tts-flash, qwen-audio-3.0-tts-plus and qwen-audio-3.0-tts-flash.

Control tags

Control tags set the emotion or style of the speech. Place a tag in the text to affect all subsequent text until the next control tag appears or the sentence is automatically segmented due to length.

Tag

Description

[sad]

Sad

[amazed]

Amazed

[deep and loud shouting]

Deep, loud shouting

[trembling]

Trembling

[angry]

Angry

[excited]

Excited

[sarcastic]

Sarcastic

[curious]

Curious

[like dracula]

Dracula style (deep, eerie)

[bored]

Bored

[tired]

Tired

[scornful]

Scornful

[shouting]

Shouting

[asmr]

ASMR soft whisper

[panicked]

Panicked

[mischievously]

Mischievous

[empathetic]

Empathetic

[whispers]

Whisper

[reluctantly]

Reluctant

[crying]

Crying

[serious]

Serious

[very slowly]

Very slow speech

[very fast]

Very fast speech

Rich language tags

Rich language tags insert a vocal effect at the current position in the text without affecting the emotional style of surrounding text.

Tag

Description

[gasp]

Gasp

[sighing]

Sigh

[clears throat]

Throat clearing

[giggles]

Giggle

[laughing]

Laughter

[cough]

Cough

[snorts]

Snort

Usage examples

The following example shows how to combine control tags and rich language tags in the text parameter:

[excited]What a beautiful day today![laughing]Let's go out and have fun together!

In this text, [excited] is a control tag that applies an excited emotion to all subsequent text. [laughing] is a rich language tag that inserts a laugh at that position before continuing to synthesize the remaining text.

You can also switch between different emotions within the same text:

[serious]Please pay attention to the safety precautions.[excited]Alright, let's get started now!

Here, [serious] sets the first sentence to a serious tone, and [excited] switches to an excited tone starting from the second sentence.

Text preprocessing recommendations

When cosyvoice-v3-flash synthesizes text that contains number segments separated by a middle dot (·), a segment may be skipped or read aloud twice. For example, consecutive room numbers separated by middle dots may be pronounced incorrectly.

To avoid this issue, replace each middle dot (·) in the input text with a Chinese comma (,) before you submit the text for synthesis:

  • Original: 主楼五楼·501房是PU·502房是OOO
  • After preprocessing: 主楼五楼501房是PU,502房是OOO

This is a known limitation at the model level. Until the model is optimized, preprocess the input text on the application side as described above.

Supported models and regions

China (Beijing)

To call the following models, use an API key for the Beijing region:

  • Qwen-Audio-TTS: qwen-audio-3.0-tts-plus, qwen-audio-3.1-tts-flash, qwen-audio-3.0-tts-flash

  • CosyVoice: cosyvoice-v3.5-plus, cosyvoice-v3.5-flash, cosyvoice-v3-plus, cosyvoice-v3-flash, cosyvoice-v2

  • Qwen-TTS:

    • Qwen3-TTS-Instruct-Flash: qwen3-tts-instruct-flash (stable version, currently equivalent to qwen3-tts-instruct-flash-2026-01-26), qwen3-tts-instruct-flash-2026-01-26 (latest snapshot)
    • Qwen3-TTS-VD: qwen3-tts-vd-2026-01-26 (latest snapshot)
    • Qwen3-TTS-VC: qwen3-tts-vc-2026-01-22 (latest snapshot)
    • Qwen3-TTS-Flash: qwen3-tts-flash (stable version, currently equivalent to qwen3-tts-flash-2025-11-27), qwen3-tts-flash-2025-11-27, qwen3-tts-flash-2025-09-18
    • Qwen-TTS: qwen-tts (stable version, currently equivalent to qwen-tts-2025-04-10), qwen-tts-latest (latest version, currently equivalent to qwen-tts-2025-05-22), qwen-tts-2025-05-22 (snapshot), qwen-tts-2025-04-10 (snapshot)
  • MiniMax: MiniMax/speech-2.8-hd, MiniMax/speech-02-hd, MiniMax/speech-2.8-turbo, MiniMax/speech-02-turbo

Singapore

To call the following models, use an API key for the Singapore region:

  • Qwen-TTS:

    • Qwen3-TTS-Instruct-Flash: qwen3-tts-instruct-flash (stable version, currently equivalent to qwen3-tts-instruct-flash-2026-01-26), qwen3-tts-instruct-flash-2026-01-26 (latest snapshot)
    • Qwen3-TTS-VD: qwen3-tts-vd-2026-01-26 (latest snapshot)
    • Qwen3-TTS-VC: qwen3-tts-vc-2026-01-22 (latest snapshot)
    • Qwen3-TTS-Flash: qwen3-tts-flash (stable version, currently equivalent to qwen3-tts-flash-2025-11-27), qwen3-tts-flash-2025-11-27, qwen3-tts-flash-2025-09-18

Supported system voices

Different models support different voices. Set the voice request parameter to a value from the voice parameter column in the following tables.

API reference

FAQ

Q: How long is the audio file URL valid?

A: The audio file URL is valid for 24 hours after generation. After the URL expires, call the API again to obtain a new URL.