API reference

Updated at:

The long-text-to-speech feature converts extra-long text, such as text with thousands of words, into binary speech data.

Note

The Speech Synthesis product page provides samples for most voices. If the voice you want is not available on the product page, you can call the API to listen to a sample. For more information about the API, see Java SDK and C++ SDK.

Billing and concurrency limits

Features

  • Supports PCM, WAV, and MP3 encoding formats for output data.

  • Supports adjustments for speech rate, pitch, and volume.

  • Supports male and female voices.

  • Long-text-to-speech has unique advantages over standard speech synthesis:

    • Supports longer text input. You can synthesize up to 10,000 characters in a single request. Each Chinese character, English letter, punctuation mark, or space counts as one character.

    • Exclusive voices: Provides high-quality, exclusive voices tailored for specific scenarios, such as reading novels, news, and video dubbing.

  • Supports multi-emotional voices. For more information, see the <emotion> tag in Markup language. Tags do not count as characters.

Important

To use the long-text-to-speech feature, update your SDK to the latest version.

Voice list

Note

To listen to more voice samples, go to the Speech Synthesis product page. The product page provides samples for most voices. If the voice you want is not available on the product page, you can call the API to listen to a sample. For more information about the API, see Java SDK and C++ SDK.

Name

voice parameter value

Type

Scenarios

Supported languages

Supported sample rates (Hz)

Supports word/sentence-level timestamps

Supports r-sound

Voice quality

Abin

abin

Cantonese-accented Mandarin

Conversational digital human

Supports Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz/48 kHz

No

No

Standard Edition

Zhixiaobai

zhixiaobai

Mandarin female voice

Conversational digital human

Supports Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz/48 kHz

No

Yes

Standard Edition

Zhixiaoxia

zhixiaoxia

Mandarin female voice

Conversational digital human

Supports Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz/48 kHz

No

Yes

Standard Edition

Zhixiaomei

zhixiaomei

Mandarin female voice

Livestreaming digital human

Supports Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

Yes

Standard Edition

Zhigui

zhigui

Mandarin female voice

Livestreaming digital human

Supports Chinese and Chinese-English mixed scenarios

8K/16K

Yes

Yes

Standard Edition

Zhishuo

zhishuo

Mandarin male voice

Customer service digital human

Supports Chinese and Chinese-English mixed scenarios

8 K/16 K

Yes

Yes

Standard Edition

Aixia

aixia

Mandarin female voice

Customer service digital human

Supports Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

Yes

Standard Edition

Cally

cally

American English female voice

Spoken English conversational digital human

Supports English-only scenarios

8K or 16K

Yes

Yes

Standard Edition

Zhifeng_Multi-emotional

zhifeng_emo

Multi-emotional male voice

General

Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

Yes

Standard Edition

Zhibing_Multi-emotional

zhibing_emo

Multi-emotional male voice

General

Chinese-only scenarios

8 kHz/16 kHz/24 kHz

Yes

Yes

Standard Edition

Zhimiao_Multi-emotional

zhimiao_emo

Multi-emotional female voice

Chinese-English scenarios

Chinese and English scenarios

8K/16K

Yes

Yes

Standard Edition

Zhimi_Multi-emotional

zhimi_emo

Multi-emotional female voice

General

Chinese and Chinese-English mixed scenarios

8K/16K

Yes

No

Standard Edition

Zhiyan_Multi-emotional

zhiyan_emo

Multi-emotional female voice

General

Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

No

Standard Edition

Zhibei_Multi-emotional

zhibei_emo

Multi-emotional child voice

General

Chinese and Chinese-English mixed scenarios

8 K/16 K

Yes

No

Standard Edition

Zhitian_Multi-emotional

zhitian_emo

Multi-emotional female voice

General

Chinese and Chinese-English mixed scenarios

8K/16K

Yes

No

Standard Edition

Xiaoyun

xiaoyun

Standard female voice

General

Chinese and Chinese-English mixed scenarios

8K/16K

No

No

Lite Edition

Xiaogang

xiaogang

Standard male voice

General

Chinese and Chinese-English mixed scenarios

8K or 16K

No

No

Lite Edition

Ruoxi

ruoxi

Gentle female voice

General

Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz

No

No

Standard Edition

Siqi

siqi

Gentle female voice

General

Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

Sijia

sijia

Standard female voice

General

Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz

No

No

Standard Edition

Sicheng

sicheng

Standard male voice

General

Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

Aiqi

aiqi

Gentle female voice

General

Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

No

Standard Edition

Aijia

aijia

Standard female voice

General

Chinese and Chinese-English mixed scenarios

8K/16K

Yes

No

Standard Edition

Aicheng

aicheng

Standard male voice

General

Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

No

Standard Edition

Aida

aida

Standard male voice

General

Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

No

Standard Edition

Ninger

ninger

Standard female voice

General

Chinese-only scenarios

8 kHz/16 kHz/24 kHz

No

No

Standard Edition

Ruilin

ruilin

Standard female voice

General

Chinese-only scenarios

8 kHz/16 kHz/24 kHz

No

No

Standard Edition

Siyue

siyue

Gentle female voice

Customer service

Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

Aiya

aiya

Stern female voice

Customer service

Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

No

Standard Edition

Aimei

aimei

Sweet female voice

Customer service

Chinese and Chinese-English mixed scenarios

8 K/16 K

Yes

No

Standard Edition

Aiyu

aiyu

Natural female voice

Customer service

Chinese and Chinese-English mixed scenarios

8 K or 16 K

Yes

No

Standard Edition

Aiyue

aiyue

Gentle female voice

Customer service

Chinese and Chinese-English mixed scenarios

8K/16K

Yes

No

Standard Edition

Aijing

aijing

Stern female voice

Customer service

Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

No

Standard Edition

Xiaomei

xiaomei

Sweet female voice

Customer service

Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz

No

No

Standard Edition

Aina

aina

Zhejiang-accented Mandarin female voice

Customer service

Chinese-only scenarios

8K or 16K

Yes

No

Standard Edition

Yina

yina

Zhejiang-accented Mandarin female voice

Customer service

Chinese-only scenarios

8 kHz/16 kHz/24 kHz

No

No

Standard Edition

Sijing

sijing

Stern female voice

Customer service

Chinese-only scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

Sitong

sitong

Child voice

Child voice scenarios

Chinese-only scenarios

8 kHz/16 kHz/24 kHz

No

No

Standard Edition

Xiaobei

xiaobei

Lolita female voice

Child voice scenarios

Chinese-only scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

Aitong

aitong

Child voice

Child voice scenarios

Chinese-only scenarios

8K or 16K

Yes

No

Standard Edition

Aiwei

aiwei

Lolita female voice

Child voice scenarios

Chinese-only scenarios

8K/16K

Yes

No

Standard Edition

Aibao

aibao

Lolita female voice

Child voice scenarios

Chinese-only scenarios

8K/16K

Yes

No

Standard Edition

Harry

harry

British English male voice

English scenarios

English scenarios

8K or 16K

No

No

Standard Edition

Abby

abby

American English female voice

English scenarios

English scenarios

8K or 16K

Yes

No

Standard Edition

Andy

andy

American English male voice

English scenarios

English scenarios

8K or 16K

Yes

No

Standard Edition

Eric

eric

British English male voice

English scenarios

English scenarios

8K or 16K

Yes

No

Standard Edition

Emily

emily

British English female voice

English scenarios

English scenarios

8K or 16K

Yes

No

Standard Edition

Luna

luna

British English female voice

English scenarios

English scenarios

8K/16K

Yes

No

Standard Edition

Luca

luca

British English male voice

English scenarios

English scenarios

8K/16K

Yes

No

Standard Edition

Wendy

wendy

British English female voice

English scenarios

English scenarios

8 K, 16 K, or 24 K

No

No

Standard Edition

William

william

British English male voice

English scenarios

English scenarios

8 kHz/16 kHz/24 kHz

No

No

Standard Edition

Olivia

olivia

British English female voice

English scenarios

English scenarios

8 kHz/16 kHz/24 kHz

No

No

Standard Edition

Shanshan

shanshan

Cantonese female voice

Dialect scenarios

Standard Cantonese (Simplified Chinese) and Cantonese-English mixed scenarios

8 kHz/16 kHz/24 kHz

No

No

Standard Edition

Aiyuan

aiyuan

Trusted Advisor

Literary scenarios

Chinese and Chinese-English mixed scenarios

8K/16K

Yes

Yes

Premium Edition

Aiying

aiying

Cute child voice

Literary scenarios

Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz

Yes

Yes

Premium Edition

Aixiang

aixiang

Magnetic male voice

Literary scenarios

Chinese and Chinese-English mixed scenarios

8 K or 16 K

Yes

Yes

Premium Edition

Aimo

aimo

Emotional male voice

Literary scenarios

Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

Yes

Premium Edition

Aiye

aiye

Young male voice

Literary scenarios

Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

Yes

Premium Edition

Aiting

aiting

Radio female voice

Literary scenarios

Chinese and Chinese-English mixed scenarios

8K/16K

Yes

Yes

Premium Edition

Aifan

aifan

Emotional female voice

Literary scenarios

Chinese and Chinese-English mixed scenarios

8 K/16 K

Yes

Yes

Premium Edition

Lydia

lydia

Bilingual English-Chinese female voice

English scenarios

English and English-Chinese mixed scenarios

8 kHz/16 kHz

Yes

No

Standard Edition

Xiaoyue

chuangirl

Sichuan dialect female voice

Dialect scenarios

Chinese and Chinese-English mixed scenarios

8 K/16 K

No

No

Standard Edition

Aishuo

aishuo

Natural male voice

Customer service

Chinese and Chinese-English mixed scenarios

8K/16K

Yes

No

Standard Edition

Qingqing

qingqing

Taiwanese Mandarin female voice

Dialect scenarios

Chinese-only scenarios

8 kHz/16 kHz

No

No

Standard Edition

Cuijie

cuijie

Northeastern dialect female voice

Dialect scenarios

Chinese-only scenarios

8 K/16 K

Yes

Yes

Standard Edition

Xiaoze

xiaoze

Hunanese heavy-accent male voice

Dialect scenarios

Chinese-only scenarios

8K or 16K

No

No

Standard Edition

Ainan

ainan

Advertisement male voice

Literary scenarios

Chinese and Chinese-English mixed scenarios

8K/16K

Yes

Yes

Premium Edition

Aihao

aihao

News male voice

Literary scenarios

Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

Yes

Premium Edition

Aiming

aiming

Humorous male voice

Literary scenarios

Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

Yes

Premium Edition

Aixiao

aixiao

News female voice

Literary scenarios

Chinese and Chinese-English mixed scenarios

8K/16K

Yes

Yes

Premium Edition

Aichu

aichu

Food documentary male voice

Literary scenarios

Chinese and Chinese-English mixed scenarios

8K/16K

Yes

Yes

Premium Edition

Aiqian

aiqian

News female voice

Literary scenarios

Chinese and Chinese-English mixed scenarios

8 K/16 K

Yes

Yes

Premium Edition

Tomoka

tomoka

Japanese female voice

Multilingual scenarios

Japanese-only scenarios

8K / 16K

Yes

No

Standard Edition

Tomoya

tomoya

Japanese male voice

Multilingual scenarios

Japanese-only scenarios

8K / 16K

Yes

No

Standard Edition

Annie

annie

American English female voice

English scenarios

English-only scenarios

8K or 16K

Yes

No

Standard Edition

Aishu

aishu

News male voice

Literary scenarios

Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

Yes

Premium Edition

Airu

airu

News female voice

Literary scenarios

Chinese and Chinese-English mixed scenarios

8K/16K

Yes

Yes

Premium Edition

Jiajia

jiajia

Cantonese female voice

Dialect scenarios

Standard Cantonese (Simplified Chinese) and Cantonese-English mixed scenarios

8 K/16 K

Yes

No

Standard Edition

Indah

indah

Indonesian female voice

Multilingual scenarios

Indonesian-only scenarios

8K or 16K

No

No

Standard Edition

Peach

taozi

Cantonese female voice

Dialect scenarios

Supports Standard Cantonese (Simplified Chinese) and Cantonese-English mixed scenarios

8K/16K

Yes

No

Standard Edition

Sales associate

guijie

Friendly female voice

General

Supports Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz

Yes

Yes

Standard Edition

Stella

stella

Intellectual female voice

General

Supports Chinese and Chinese-English mixed scenarios

8K/16K

Yes

Yes

Standard Edition

Stanley

stanley

Calm male voice

General

Supports Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

Yes

Standard Edition

Kenny

kenny

Calm male voice

General

Supports Chinese and Chinese-English mixed scenarios

8K/16K

Yes

Yes

Standard Edition

Rosa

rosa

Natural female voice

General

Supports Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

Yes

Standard Edition

Farah

farah

Malay female voice

Multilingual scenarios

Supports Malay-only scenarios

8K/16K

No

No

Standard Edition

Mashu

mashu

Children's drama male voice

General

General

8K/16K

Yes

No

Standard Edition

Zhiqi

zhiqi

Gentle female voice

Ultra-high definition scenarios

Supports Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz/48 kHz

Yes

No

Premium Edition

Zhichu

zhichu

Food documentary male voice

Ultra-high definition scenarios

Supports Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz/48 kHz

Yes

Yes

Premium Edition

Xiaoxian

xiaoxian

Friendly female voice

Livestreaming scenarios

Supports Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz

Yes

Yes

Standard Edition

Yuer

yuer

Children's drama female voice

General

Supports Chinese-only scenarios

8K or 16K

Yes

No

Standard Edition

Maoxiaomei

maoxiaomei

Energetic female voice

Livestreaming scenarios

Supports Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

Yes

Standard Edition

Zhixiang

zhixiang

Magnetic male voice

Ultra-high definition scenarios

Supports Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz/48 kHz

Yes

No

Premium Edition

Zhijia

zhijia

Standard female voice

Ultra-high definition scenarios

Supports Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz/48 kHz

Yes

No

Premium Edition

Zhinan

zhinan

Advertisement male voice

Ultra-high definition scenarios

Supports Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz/48 kHz

Yes

No

Premium Edition

Zhiqian

zhiqian

News female voice

Ultra-high definition scenarios

Supports Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz/48 kHz

Yes

No

Premium Edition

Zhiru

zhiru

News female voice

Ultra-high definition scenarios

Supports Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz/48 kHz

Yes

No

Premium Edition

Zhide

zhide

News male voice

Ultra-high definition scenarios

Supports Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz/48 kHz

Yes

No

Premium Edition

Zhifei

zhifei

Passionate commentary

Ultra-high definition scenarios

Supports Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

No

Premium Edition

Aifei

aifei

Passionate commentary

Livestreaming scenarios

Supports Chinese and Chinese-English mixed scenarios

8 K or 16 K

Yes

Yes

Standard Edition

Subgroup

yaqun

Store broadcast

Livestreaming scenarios

Supports Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

Yes

Standard Edition

Qiaowei

qiaowei

Store broadcast

Livestreaming scenarios

Supports Chinese and Chinese-English mixed scenarios

8K/16K

Yes

Yes

Standard Edition

Dahu

dahu

Northeastern dialect male voice

Dialect scenarios

Supports Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

Yes

Standard Edition

ava

ava

American English female voice

English scenarios

Supports English-only scenarios

8 K/16 K

Yes

No

Standard Edition

Zhilun

zhilun

Suspense commentary

Ultra-high definition scenarios

Supports Chinese and Chinese-English mixed scenarios

8K/16K

Yes

No

Premium Edition

Ailun

ailun

Suspense commentary

Livestreaming scenarios

Supports Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz

Yes

Yes

Standard Edition

Jielidou

jielidou

Healing a Child's Voice

Child voice scenarios

Supports Chinese-only scenarios

8K or 16K

Yes

Yes

Standard Edition

Zhiwei

zhiwei

Sweet Female Voice

Ultra-high definition scenarios

Supports Chinese-only scenarios

8 kHz/16 kHz/24 kHz/48 kHz

Yes

No

Premium Edition

Laotie

laotie

Northeastern buddy

Livestreaming scenarios

Supports Chinese-only scenarios

8K/16K

Yes

Yes

Standard Edition

Younger sister

laomei

Female Hawker

Livestreaming scenarios

Supports Chinese-only scenarios

8K or 16K

Yes

Yes

Standard Edition

Aikan

aikan

Tianjin dialect male voice

Dialect scenarios

Supports Chinese-only scenarios

8K/16K

Yes

Yes

Standard Edition

Tala

tala

Filipino female voice

Multilingual scenarios

Supports Filipino-only scenarios

8K or 16K

No

No

Standard Edition

Zhitian

zhitian

Sweet female voice

General

Supports Chinese and Chinese-English mixed scenarios

8K/16K

Yes

No

Premium Edition

Zhiqing

zhiqing

Taiwan (China) dialect female voice

Dialect scenarios

Supports Chinese-only scenarios

8K or 16K

Yes

No

Premium Edition

Tien

tien

Vietnamese female voice

Multilingual scenarios

Supports Vietnamese-only scenarios

8K/16K

No

No

Standard Edition

Becca

becca

American English customer service female voice

American English

Supports English-only scenarios

8K or 16K

No

No

Standard Edition

Kyong

Kyong

Korean female voice

Korean scenarios

Korean

8 kHz/16 kHz

No

No

Standard Edition

masha

masha

Russian female voice

Russian scenarios

Russian

8K/16K

No

No

Standard Edition

camila

camila

Spanish female voice

Spanish scenarios

Spanish

8 kHz/16 kHz

No

No

Standard Edition

perla

perla

Italian female voice

Italian scenarios

Italian

8 kHz/16 kHz

No

No

Standard Edition

Zhimao

zhimao

Mandarin female voice

Livestreaming

Chinese

8 kHz/16 kHz

Yes

No

Standard Edition

Zhiyuan

zhiyuan

Mandarin female voice

General

Chinese

8 kHz/16 kHz

Yes

No

Standard Edition

Zhiya

zhiya

Mandarin female voice

Customer service

Chinese

8 kHz/16 kHz

Yes

No

Standard Edition

Zhiyue

zhiyue

Mandarin female voice

General

Chinese

8 kHz/16 kHz

Yes

No

Standard Edition

Zhida

zhida

Mandarin male voice

General

Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz

Yes

No

Standard Edition

Zhi Sha

zhistella

Mandarin female voice

General

Chinese

8 kHz/16 kHz

Yes

No

Standard Edition

Kelly

kelly

Hong Kong Cantonese female voice

Dialect scenarios

Hong Kong Cantonese

8 kHz/16 kHz

Yes

No

Standard Edition

clara

clara

French female voice

General

French

8 kHz/16 kHz

No

No

Standard Edition

hanna

hanna

German female voice

General

German

8 kHz/16 kHz

No

No

Standard Edition

waan

waan

Thai female voice

General

Thai

8 kHz/16 kHz

No

No

Standard Edition

betty

betty

American English female voice

General

American English

8 kHz/16 kHz

Yes

No

Standard Edition

beth

beth

American English female voice

General

American English

8 kHz/16 kHz

Yes

No

Standard Edition

cindy

cindy

American English female voice

General

American English

8 kHz/16 kHz

Yes

No

Standard Edition

donna

donna

American English female voice

General

American English

8 kHz/16 kHz

Yes

No

Standard Edition

eva

eva

American English female voice

General

American English

8 kHz/16 kHz

Yes

No

Standard Edition

brian

brian

American English male voice

General

American English

8 kHz/16 kHz

Yes

No

Standard Edition

david

david

American English male voice

General

American English

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

abby_ecmix

abby_ecmix

American English female voice

General

English and English-Chinese mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

annie_ecmix

annie_ecmix

American English female voice

General

English and English-Chinese mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

andy_ecmix

andy_ecmix

American English male voice

General

English and English-Chinese mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

ava_ecmix

ava_ecmix

American English female voice

General

English and English-Chinese mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

betty_ecmix

betty_ecmix

American English female voice

General

English and English-Chinese mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

beth_ecmix

beth_ecmix

American English female voice

General

English and English-Chinese mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

brian_ecmix

brian_ecmix

American English male voice

General

English and English-Chinese mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

cindy_ecmix

cindy_ecmix

American English female voice

General

English and English-Chinese mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

cally_ecmix

cally_ecmix

American English female voice

General

English and English-Chinese mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

donna_ecmix

donna_ecmix

American English female voice

General

English and English-Chinese mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

david_ecmix

david_ecmix

American English male voice

General

English and English-Chinese mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

eva_ecmix

eva_ecmix

American English female voice

General

English and English-Chinese mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

Support for multi-emotional voices

Only multi-emotional voice models support emotion selection. The supported emotions are listed in the following table. The supported emotion categories vary by voice. The main categories include neutral, happy, angry, sad, fear, hate, surprise, arousal, serious, disgust, jealousy, embarrassed, frustrated, affectionate, gentle, newscast, customer service, and living.

Voice name

voice parameter value

Emotion category

Zhifeng_Multi-emotional

zhifeng_emo

angry, fear, happy, neutral, sad, surprise

Zhibing_Multi-emotional

zhibing_emo

angry, fear, happy, neutral, sad, surprise

Zhimiao_Multi-emotional

zhimiao_emo

serious, sad, disgust, jealousy, embarrassed, happy, fear, surprise, neutral, frustrated, affectionate, gentle, angry, newscast, customer-service, story, living

Zhimi_Multi-emotional

zhimi_emo

angry, fear, happy, hate, neutral, sad, surprise

Zhiyan_Multi-emotional

zhiyan_emo

neutral, happy, angry, sad, fear, hate, surprise, arousal

Zhibei_Multi-emotional

zhibei_emo

neutral, happy, angry, sad, fear, hate, surprise

Zhitian_Multi-emotional

zhitian_emo

neutral, happy, angry, sad, fear, hate, surprise

Call instructions

  • The input text must be UTF-8 encoded.

  • Long text-to-speech and standard speech synthesis share many similarities.

Intelligent access to the nearest region

Speech Synthesis supports intelligent access to the nearest region using the domain name nls-gateway.aliyuncs.com.

Connect to the nearest region for optimal performance. The system automatically resolves requests to the server in the nearest region based on the client's geographical location. For example, a request from the China (Beijing) region is automatically routed to a server in the China (Beijing) region. This provides the same result as specifying the nls-gateway-cn-beijing.aliyuncs.com domain name.

Service endpoints

Access type

Description

URL

Public access (defaults to the China (Shanghai) region)

All servers can use the public access URL. The SDK is configured with the public access URL by default.

  • China (Shanghai): wss://nls-gateway-cn-shanghai.aliyuncs.com/ws/v1

  • China (Beijing): wss://nls-gateway-cn-beijing.aliyuncs.com/ws/v1

  • China (Shenzhen): wss://nls-gateway-cn-shenzhen.aliyuncs.com/ws/v1

ECS private network access

If you use an Alibaba Cloud ECS instance in the China (Shanghai), China (Beijing), or China (Shenzhen) region, you can use the private network access URL. ECS instances in the classic network cannot access AnyTunnel and therefore cannot access the Voice Service over the private network. To use AnyTunnel, create a Virtual Private Cloud (VPC) and access the service from within the VPC.

Note

  • Private network access does not incur data transfer costs for your ECS instance.

  • For more information about ECS network types, see Network types.

  • China (Shanghai): ws://nls-gateway-cn-shanghai-internal.aliyuncs.com:80/ws/v1

  • China (Beijing): ws://nls-gateway-cn-beijing-internal.aliyuncs.com:80/ws/v1

  • China (Shenzhen): ws://nls-gateway-cn-shenzhen-internal.aliyuncs.com:80/ws/v1

Interaction flow

Note
  • The preceding figure does not show the interaction flow for the RESTful API. For the RESTful API interaction flowchart, see RESTful API.

  • In addition to the audio stream, the server response header includes the `task_id` parameter. This parameter is the unique ID of the request.

  • To play the audio stream that is returned by the server in real time, use an audio player that supports stream playback, such as ffmpeg, PyAudio (Python), AudioFormat (Java), or MediaSource (JavaScript).

  1. Authentication

    The client uses a token for authentication when establishing a WebSocket connection with the server. For more information about obtaining a token, see Obtain a token.

  2. Start synthesis

    The client sends a speech synthesis request and configures parameters in the request message. You can set each parameter using the corresponding `set` method of the `SpeechSynthesizer` object in the SDK. The parameters are described in the following table.

    Parameter

    Type

    Required

    Description

    appkey

    String

    Yes

    The AppKey of the project that you created in the console.

    text

    String

    Yes

    The text to synthesize. The text must be UTF-8 encoded. Add a space between English words.

    Note

    To use a multi-emotional voice, add the `ssml-emotion` tag to the text. For more information, see <emotion>.

    Only voices that support multiple emotions can use the <emotion> tag. Otherwise, the `Illegal ssml text` error is reported.

    voice

    String

    No

    The voice. The default value is xiaoyun.

    format

    String

    No

    The audio encoding format. Valid values: PCM, WAV, and MP3. Default value: pcm.

    sample_rate

    Integer

    No

    The audio sample rate. Default value: 16000 Hz.

    volume

    Integer

    No

    The volume. Valid values: 0 to 100. Default value: 50.

    speech_rate

    Integer

    No

    The speech rate. Valid values: -500 to 500. Default value: 0.

    The range [-500, 0, 500] corresponds to a speed multiplier of [0.5, 1.0, 2.0].

    • -500 indicates 0.5 times the default speed.

    • 0 indicates the default speed (1.0x). The default speed varies slightly for each voice but is approximately four characters per second.

    • 500 indicates 2.0 times the default speed.

    The calculation method is as follows:

    • 0.8x speed: (1 - 1/0.8) / 0.002 = -125

    • 1.2x speed: (1 - 1/1.2) / 0.001 = 166

    Important

    • Use a coefficient of 0.002 for speeds less than 1.0x.

    • Use a coefficient of 0.001 for speeds greater than 1.0x.

    The actual algorithm result is an approximate value.

    pitch_rate

    Integer

    No

    The pitch. Valid values: -500 to 500. Default value: 0.

    enable_subtitle

    Boolean

    No

    Enables word-level timestamps. For more information about how to use this feature, see Timestamps.

  3. Receive synthesized data

    The server returns the synthesized speech as binary data. The SDK receives and processes the binary data.

  4. End synthesis

    After the speech is synthesized, the server sends a `SynthesisCompleted` event notification. The following code provides an example.

    {
        "header":{
            "namespace":"SpeechLongSynthesizer",
            "name":"SynthesisCompleted",
            "status":20000000,
            "message_id":"396c80b3abf84082a48cb9e5c424****",
            "task_id":"f5805be640364cdcafc8da63e512****",
            "status_text":"Gateway:SUCCESS:Success."
        }
    }
  5. Handle synthesis failures

    If the synthesis task fails due to incorrect parameters or other reasons, a `TaskFailed` notification is returned. The underlying connection is then closed. The following code provides an example.

    {
       "header":{
          "namespace":"Default",
          "name":"TaskFailed",
          "status":41020001,
          "message_id":"62c126f7d9b340deb82b5b7eaca0****",
          "task_id":"4552df26d1f547aab9a2c4a94678****",
          "status_text":"TTS:TtsClientError:[tts]Engine return error code: 418"
       }
    }

Service status codes

Each service response includes a `status` field, which is the service status code. The following sections describe the meanings of different status codes.

General-purpose error codes

Status code

Status message

Cause

Solution

40000000

The default client error code. This code corresponds to multiple error messages.

Invalid parameters or call logic was used.

Compare your code with the sample code in the official documentation to test and verify it.

40000001

The token 'xxx' has expired.

The token 'xxx' is invalid

Invalid parameters or call logic was used. This is a general-purpose client error code that usually indicates an incorrect token, such as an expired or invalid token.

Compare your code with the sample code in the official documentation to test and verify it.

40000002

Gateway:MESSAGE_INVALID:Can't process message in state'FAILED'!

The message is invalid or incorrect.

Compare your code with the sample code in the official documentation to test and verify it.

40000003

PARAMETER_INVALID

Failed to decode url params

The parameters passed by the user are incorrect. This error is common for RESTful API calls.

Compare your code with the sample code in the official documentation to test and verify it.

40000005

Gateway:TOO_MANY_REQUESTS:Too many requests!

Too many concurrent requests.

If you are using the Free Edition, you can upgrade to a commercial version to increase the concurrency.

If you are already using a commercial version, you can purchase a concurrency resource plan to increase your concurrency quota.

40000009

Invalid wav header!

The message header is invalid.

If you send a WAV audio file and set the format parameter to wav, check whether the WAV header of the audio file is correct. If the header is incorrect, the server may reject the request.

40000009

Too large wav header!

The WAV header of the transmitted audio is invalid.

You can send the audio stream in a format such as PCM or OPUS. If you use the WAV format, make sure that the WAV header of the audio file contains the correct data length.

40000010

Gateway:FREE_TRIAL_EXPIRED:The free trial has expired!

The trial period has ended, and the commercial version is not activated or your account has an overdue payment.

You can log on to the console to check the service activation status and your account balance.

40010001

Gateway:NAMESPACE_NOT_FOUND:RESTful url path illegal

The operation or parameter is not supported.

Check whether the parameters passed in the call are consistent with the requirements in the official documentation. You can compare them with the error message to identify and set the correct parameters.

For example, if you are using a curl command to make a RESTful API request, check whether the URL you constructed is valid.

40010003

Gateway:DIRECTIVE_INVALID:[xxx]

A general-purpose client-side error code.

This error indicates that the client passed an incorrect parameter or instruction. Detailed error messages are available for different operations. You can refer to the corresponding documentation to set the parameters correctly.

40010004

Gateway:CLIENT_DISCONNECT:Client disconnected before task finished!

The client actively terminated the connection before the request was processed.

None. Alternatively, you can close the connection after the server responds.

40010005

Gateway:TASK_STATE_ERROR:Got stop directive while task is stopping!

The client sent a message instruction that is not currently supported.

Compare your code with the sample code in the official documentation to test and verify it.

40020105

Meta:APPKEY_NOT_EXIST:Appkey not exist!

A non-existent Appkey was used.

Confirm whether a non-existent Appkey was used. You can log on to the console and view the project configuration to find the Appkey.

40020106

Meta:APPKEY_UID_MISMATCH:Appkey and user mismatch!

The Appkey and token passed in the call were not created by the same Alibaba Cloud account UID. This causes a mismatch.

Check whether you are using resources from two different accounts. Do not use an Appkey from Account A with a token generated from Account B.

403

Forbidden

The token is invalid. For example, the token does not exist or has expired.

Set a valid token. Tokens have an expiration period. You must obtain a new token before the current one expires.

41000003

MetaInfo doesn't have end point info

Failed to retrieve the routing information for this Appkey.

Check whether you are using resources from two different accounts. Do not use an Appkey from Account A with a token generated from Account B.

41010101

UNSUPPORTED_SAMPLE_RATE

The sample rate is not supported.

Real-time speech recognition currently supports only audio with a sample rate of 8000 Hz or 16000 Hz.

41040201

Realtime:GET_CLIENT_DATA_TIMEOUT:Client data does not send continuously!

Failed to retrieve data from the client due to a timeout.

When you call real-time speech recognition, the client must send data at a real-time rate and close the connection promptly after the data is sent.

50000000

GRPC_ERROR:Grpc error!

An exception caused by factors such as machine load or network issues. This error usually occurs randomly.

You can retry the call to resolve the issue.

50000001

GRPC_ERROR:Grpc error!

An exception caused by factors such as machine load or network issues. This error usually occurs randomly.

You can retry the call to resolve the issue.

52010001

GRPC_ERROR:Grpc error!

An exception caused by factors such as machine load or network issues. This error usually occurs randomly.

You can retry the call to resolve the issue.

Speech synthesis/Long text-to-speech error codes

Status code

Status message

Cause

Solution

21050000

SUCCESS

Success.

None.

21050001

RUNNING

The asynchronous long text-to-speech task is running.

You can send a GET request later to query the detection result.

21050002

QUEUEING

The asynchronous long text-to-speech task is in the queue.

Send a GET request later to query the detection result.

40000001

Gateway:ACCESS_DENIED:No privilege to this voice!

An invalid voice name was set.

Set a valid voice name based on the official documentation.

40000004

Gateway:IDLE_TIMEOUT:Websocket session is idle for too long time,the last directive is 'StartSynthesis'!

After a connection is established, no data is sent for an extended period. The server returns this error message after 10 seconds.

Close the connection promptly after the request is processed. This error may also occur when the server is under high momentary pressure and cannot return data in time. In this case, you can retry the call.

40010003

Gateway:DIRECTIVE_INVALID:No text specified!

No valid text was set for synthesis.

Set the text to be synthesized based on the sample code in the official documentation.

41020001

Speech synthesis call client-side error

Multiple error messages may be displayed. Resolve each error based on the corresponding message.

  • If the error message is Engine return error code: 424., it indicates that the background music or concatenated recording is in an invalid format. Set the background music to a supported format as described in the documentation.

  • If the error message is Engine return error code:418, it indicates that an unsupported voice name was passed.

  • If the error message is Engine return error code: 413, it indicates that the SSML format is invalid.

  • If the error message is Request json illegal,failed to parse request., it indicates that the JSON format is invalid.

  • If the error message is SSML text length should be less than 300., it indicates that the synthesis text is too long. We recommend that you use the long text-to-speech API.

51020001

TTS:TtsServerError

An exception occurred due to factors such as machine load or network issues. This error is usually intermittent.

Retry the call.