API reference
The long-text-to-speech feature converts extra-long text, such as text with thousands of words, into binary speech data.
The Speech Synthesis product page provides samples for most voices. If the voice you want is not available on the product page, you can call the API to listen to a sample. For more information about the API, see Java SDK and C++ SDK.
Billing and concurrency limits
Real-time long-text-to-speech is available only in the commercial version and does not support a free trial. For more information, see Trial and commercial versions. To use this feature, you must activate the commercial version. For more information, see Upgrade from the trial version to the commercial version.
For more information about billing methods, see Billing methods.
For more information about concurrency limits, see Concurrency and QPS limits.
Features
Supports PCM, WAV, and MP3 encoding formats for output data.
Supports adjustments for speech rate, pitch, and volume.
Supports male and female voices.
Long-text-to-speech has unique advantages over standard speech synthesis:
Supports longer text input. You can synthesize up to 10,000 characters in a single request. Each Chinese character, English letter, punctuation mark, or space counts as one character.
Exclusive voices: Provides high-quality, exclusive voices tailored for specific scenarios, such as reading novels, news, and video dubbing.
Supports multi-emotional voices. For more information, see the <emotion> tag in Markup language. Tags do not count as characters.
To use the long-text-to-speech feature, update your SDK to the latest version.
Voice list
To listen to more voice samples, go to the Speech Synthesis product page. The product page provides samples for most voices. If the voice you want is not available on the product page, you can call the API to listen to a sample. For more information about the API, see Java SDK and C++ SDK.
Name | voice parameter value | Type | Scenarios | Supported languages | Supported sample rates (Hz) | Supports word/sentence-level timestamps | Supports r-sound | Voice quality |
Abin | abin | Cantonese-accented Mandarin | Conversational digital human | Supports Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz/48 kHz | No | No | Standard Edition |
Zhixiaobai | zhixiaobai | Mandarin female voice | Conversational digital human | Supports Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz/48 kHz | No | Yes | Standard Edition |
Zhixiaoxia | zhixiaoxia | Mandarin female voice | Conversational digital human | Supports Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz/48 kHz | No | Yes | Standard Edition |
Zhixiaomei | zhixiaomei | Mandarin female voice | Livestreaming digital human | Supports Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | Yes | Standard Edition |
Zhigui | zhigui | Mandarin female voice | Livestreaming digital human | Supports Chinese and Chinese-English mixed scenarios | 8K/16K | Yes | Yes | Standard Edition |
Zhishuo | zhishuo | Mandarin male voice | Customer service digital human | Supports Chinese and Chinese-English mixed scenarios | 8 K/16 K | Yes | Yes | Standard Edition |
Aixia | aixia | Mandarin female voice | Customer service digital human | Supports Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | Yes | Standard Edition |
Cally | cally | American English female voice | Spoken English conversational digital human | Supports English-only scenarios | 8K or 16K | Yes | Yes | Standard Edition |
Zhifeng_Multi-emotional | zhifeng_emo | Multi-emotional male voice | General | Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | Yes | Standard Edition |
Zhibing_Multi-emotional | zhibing_emo | Multi-emotional male voice | General | Chinese-only scenarios | 8 kHz/16 kHz/24 kHz | Yes | Yes | Standard Edition |
Zhimiao_Multi-emotional | zhimiao_emo | Multi-emotional female voice | Chinese-English scenarios | Chinese and English scenarios | 8K/16K | Yes | Yes | Standard Edition |
Zhimi_Multi-emotional | zhimi_emo | Multi-emotional female voice | General | Chinese and Chinese-English mixed scenarios | 8K/16K | Yes | No | Standard Edition |
Zhiyan_Multi-emotional | zhiyan_emo | Multi-emotional female voice | General | Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | No | Standard Edition |
Zhibei_Multi-emotional | zhibei_emo | Multi-emotional child voice | General | Chinese and Chinese-English mixed scenarios | 8 K/16 K | Yes | No | Standard Edition |
Zhitian_Multi-emotional | zhitian_emo | Multi-emotional female voice | General | Chinese and Chinese-English mixed scenarios | 8K/16K | Yes | No | Standard Edition |
Xiaoyun | xiaoyun | Standard female voice | General | Chinese and Chinese-English mixed scenarios | 8K/16K | No | No | Lite Edition |
Xiaogang | xiaogang | Standard male voice | General | Chinese and Chinese-English mixed scenarios | 8K or 16K | No | No | Lite Edition |
Ruoxi | ruoxi | Gentle female voice | General | Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz | No | No | Standard Edition |
Siqi | siqi | Gentle female voice | General | Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
Sijia | sijia | Standard female voice | General | Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz | No | No | Standard Edition |
Sicheng | sicheng | Standard male voice | General | Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
Aiqi | aiqi | Gentle female voice | General | Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | No | Standard Edition |
Aijia | aijia | Standard female voice | General | Chinese and Chinese-English mixed scenarios | 8K/16K | Yes | No | Standard Edition |
Aicheng | aicheng | Standard male voice | General | Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | No | Standard Edition |
Aida | aida | Standard male voice | General | Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | No | Standard Edition |
Ninger | ninger | Standard female voice | General | Chinese-only scenarios | 8 kHz/16 kHz/24 kHz | No | No | Standard Edition |
Ruilin | ruilin | Standard female voice | General | Chinese-only scenarios | 8 kHz/16 kHz/24 kHz | No | No | Standard Edition |
Siyue | siyue | Gentle female voice | Customer service | Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
Aiya | aiya | Stern female voice | Customer service | Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | No | Standard Edition |
Aimei | aimei | Sweet female voice | Customer service | Chinese and Chinese-English mixed scenarios | 8 K/16 K | Yes | No | Standard Edition |
Aiyu | aiyu | Natural female voice | Customer service | Chinese and Chinese-English mixed scenarios | 8 K or 16 K | Yes | No | Standard Edition |
Aiyue | aiyue | Gentle female voice | Customer service | Chinese and Chinese-English mixed scenarios | 8K/16K | Yes | No | Standard Edition |
Aijing | aijing | Stern female voice | Customer service | Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | No | Standard Edition |
Xiaomei | xiaomei | Sweet female voice | Customer service | Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz | No | No | Standard Edition |
Aina | aina | Zhejiang-accented Mandarin female voice | Customer service | Chinese-only scenarios | 8K or 16K | Yes | No | Standard Edition |
Yina | yina | Zhejiang-accented Mandarin female voice | Customer service | Chinese-only scenarios | 8 kHz/16 kHz/24 kHz | No | No | Standard Edition |
Sijing | sijing | Stern female voice | Customer service | Chinese-only scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
Sitong | sitong | Child voice | Child voice scenarios | Chinese-only scenarios | 8 kHz/16 kHz/24 kHz | No | No | Standard Edition |
Xiaobei | xiaobei | Lolita female voice | Child voice scenarios | Chinese-only scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
Aitong | aitong | Child voice | Child voice scenarios | Chinese-only scenarios | 8K or 16K | Yes | No | Standard Edition |
Aiwei | aiwei | Lolita female voice | Child voice scenarios | Chinese-only scenarios | 8K/16K | Yes | No | Standard Edition |
Aibao | aibao | Lolita female voice | Child voice scenarios | Chinese-only scenarios | 8K/16K | Yes | No | Standard Edition |
Harry | harry | British English male voice | English scenarios | English scenarios | 8K or 16K | No | No | Standard Edition |
Abby | abby | American English female voice | English scenarios | English scenarios | 8K or 16K | Yes | No | Standard Edition |
Andy | andy | American English male voice | English scenarios | English scenarios | 8K or 16K | Yes | No | Standard Edition |
Eric | eric | British English male voice | English scenarios | English scenarios | 8K or 16K | Yes | No | Standard Edition |
Emily | emily | British English female voice | English scenarios | English scenarios | 8K or 16K | Yes | No | Standard Edition |
Luna | luna | British English female voice | English scenarios | English scenarios | 8K/16K | Yes | No | Standard Edition |
Luca | luca | British English male voice | English scenarios | English scenarios | 8K/16K | Yes | No | Standard Edition |
Wendy | wendy | British English female voice | English scenarios | English scenarios | 8 K, 16 K, or 24 K | No | No | Standard Edition |
William | william | British English male voice | English scenarios | English scenarios | 8 kHz/16 kHz/24 kHz | No | No | Standard Edition |
Olivia | olivia | British English female voice | English scenarios | English scenarios | 8 kHz/16 kHz/24 kHz | No | No | Standard Edition |
Shanshan | shanshan | Cantonese female voice | Dialect scenarios | Standard Cantonese (Simplified Chinese) and Cantonese-English mixed scenarios | 8 kHz/16 kHz/24 kHz | No | No | Standard Edition |
Aiyuan | aiyuan | Trusted Advisor | Literary scenarios | Chinese and Chinese-English mixed scenarios | 8K/16K | Yes | Yes | Premium Edition |
Aiying | aiying | Cute child voice | Literary scenarios | Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz | Yes | Yes | Premium Edition |
Aixiang | aixiang | Magnetic male voice | Literary scenarios | Chinese and Chinese-English mixed scenarios | 8 K or 16 K | Yes | Yes | Premium Edition |
Aimo | aimo | Emotional male voice | Literary scenarios | Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | Yes | Premium Edition |
Aiye | aiye | Young male voice | Literary scenarios | Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | Yes | Premium Edition |
Aiting | aiting | Radio female voice | Literary scenarios | Chinese and Chinese-English mixed scenarios | 8K/16K | Yes | Yes | Premium Edition |
Aifan | aifan | Emotional female voice | Literary scenarios | Chinese and Chinese-English mixed scenarios | 8 K/16 K | Yes | Yes | Premium Edition |
Lydia | lydia | Bilingual English-Chinese female voice | English scenarios | English and English-Chinese mixed scenarios | 8 kHz/16 kHz | Yes | No | Standard Edition |
Xiaoyue | chuangirl | Sichuan dialect female voice | Dialect scenarios | Chinese and Chinese-English mixed scenarios | 8 K/16 K | No | No | Standard Edition |
Aishuo | aishuo | Natural male voice | Customer service | Chinese and Chinese-English mixed scenarios | 8K/16K | Yes | No | Standard Edition |
Qingqing | qingqing | Taiwanese Mandarin female voice | Dialect scenarios | Chinese-only scenarios | 8 kHz/16 kHz | No | No | Standard Edition |
Cuijie | cuijie | Northeastern dialect female voice | Dialect scenarios | Chinese-only scenarios | 8 K/16 K | Yes | Yes | Standard Edition |
Xiaoze | xiaoze | Hunanese heavy-accent male voice | Dialect scenarios | Chinese-only scenarios | 8K or 16K | No | No | Standard Edition |
Ainan | ainan | Advertisement male voice | Literary scenarios | Chinese and Chinese-English mixed scenarios | 8K/16K | Yes | Yes | Premium Edition |
Aihao | aihao | News male voice | Literary scenarios | Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | Yes | Premium Edition |
Aiming | aiming | Humorous male voice | Literary scenarios | Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | Yes | Premium Edition |
Aixiao | aixiao | News female voice | Literary scenarios | Chinese and Chinese-English mixed scenarios | 8K/16K | Yes | Yes | Premium Edition |
Aichu | aichu | Food documentary male voice | Literary scenarios | Chinese and Chinese-English mixed scenarios | 8K/16K | Yes | Yes | Premium Edition |
Aiqian | aiqian | News female voice | Literary scenarios | Chinese and Chinese-English mixed scenarios | 8 K/16 K | Yes | Yes | Premium Edition |
Tomoka | tomoka | Japanese female voice | Multilingual scenarios | Japanese-only scenarios | 8K / 16K | Yes | No | Standard Edition |
Tomoya | tomoya | Japanese male voice | Multilingual scenarios | Japanese-only scenarios | 8K / 16K | Yes | No | Standard Edition |
Annie | annie | American English female voice | English scenarios | English-only scenarios | 8K or 16K | Yes | No | Standard Edition |
Aishu | aishu | News male voice | Literary scenarios | Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | Yes | Premium Edition |
Airu | airu | News female voice | Literary scenarios | Chinese and Chinese-English mixed scenarios | 8K/16K | Yes | Yes | Premium Edition |
Jiajia | jiajia | Cantonese female voice | Dialect scenarios | Standard Cantonese (Simplified Chinese) and Cantonese-English mixed scenarios | 8 K/16 K | Yes | No | Standard Edition |
Indah | indah | Indonesian female voice | Multilingual scenarios | Indonesian-only scenarios | 8K or 16K | No | No | Standard Edition |
Peach | taozi | Cantonese female voice | Dialect scenarios | Supports Standard Cantonese (Simplified Chinese) and Cantonese-English mixed scenarios | 8K/16K | Yes | No | Standard Edition |
Sales associate | guijie | Friendly female voice | General | Supports Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz | Yes | Yes | Standard Edition |
Stella | stella | Intellectual female voice | General | Supports Chinese and Chinese-English mixed scenarios | 8K/16K | Yes | Yes | Standard Edition |
Stanley | stanley | Calm male voice | General | Supports Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | Yes | Standard Edition |
Kenny | kenny | Calm male voice | General | Supports Chinese and Chinese-English mixed scenarios | 8K/16K | Yes | Yes | Standard Edition |
Rosa | rosa | Natural female voice | General | Supports Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | Yes | Standard Edition |
Farah | farah | Malay female voice | Multilingual scenarios | Supports Malay-only scenarios | 8K/16K | No | No | Standard Edition |
Mashu | mashu | Children's drama male voice | General | General | 8K/16K | Yes | No | Standard Edition |
Zhiqi | zhiqi | Gentle female voice | Ultra-high definition scenarios | Supports Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz/48 kHz | Yes | No | Premium Edition |
Zhichu | zhichu | Food documentary male voice | Ultra-high definition scenarios | Supports Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz/48 kHz | Yes | Yes | Premium Edition |
Xiaoxian | xiaoxian | Friendly female voice | Livestreaming scenarios | Supports Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz | Yes | Yes | Standard Edition |
Yuer | yuer | Children's drama female voice | General | Supports Chinese-only scenarios | 8K or 16K | Yes | No | Standard Edition |
Maoxiaomei | maoxiaomei | Energetic female voice | Livestreaming scenarios | Supports Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | Yes | Standard Edition |
Zhixiang | zhixiang | Magnetic male voice | Ultra-high definition scenarios | Supports Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz/48 kHz | Yes | No | Premium Edition |
Zhijia | zhijia | Standard female voice | Ultra-high definition scenarios | Supports Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz/48 kHz | Yes | No | Premium Edition |
Zhinan | zhinan | Advertisement male voice | Ultra-high definition scenarios | Supports Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz/48 kHz | Yes | No | Premium Edition |
Zhiqian | zhiqian | News female voice | Ultra-high definition scenarios | Supports Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz/48 kHz | Yes | No | Premium Edition |
Zhiru | zhiru | News female voice | Ultra-high definition scenarios | Supports Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz/48 kHz | Yes | No | Premium Edition |
Zhide | zhide | News male voice | Ultra-high definition scenarios | Supports Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz/48 kHz | Yes | No | Premium Edition |
Zhifei | zhifei | Passionate commentary | Ultra-high definition scenarios | Supports Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | No | Premium Edition |
Aifei | aifei | Passionate commentary | Livestreaming scenarios | Supports Chinese and Chinese-English mixed scenarios | 8 K or 16 K | Yes | Yes | Standard Edition |
Subgroup | yaqun | Store broadcast | Livestreaming scenarios | Supports Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | Yes | Standard Edition |
Qiaowei | qiaowei | Store broadcast | Livestreaming scenarios | Supports Chinese and Chinese-English mixed scenarios | 8K/16K | Yes | Yes | Standard Edition |
Dahu | dahu | Northeastern dialect male voice | Dialect scenarios | Supports Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | Yes | Standard Edition |
ava | ava | American English female voice | English scenarios | Supports English-only scenarios | 8 K/16 K | Yes | No | Standard Edition |
Zhilun | zhilun | Suspense commentary | Ultra-high definition scenarios | Supports Chinese and Chinese-English mixed scenarios | 8K/16K | Yes | No | Premium Edition |
Ailun | ailun | Suspense commentary | Livestreaming scenarios | Supports Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz | Yes | Yes | Standard Edition |
Jielidou | jielidou | Healing a Child's Voice | Child voice scenarios | Supports Chinese-only scenarios | 8K or 16K | Yes | Yes | Standard Edition |
Zhiwei | zhiwei | Sweet Female Voice | Ultra-high definition scenarios | Supports Chinese-only scenarios | 8 kHz/16 kHz/24 kHz/48 kHz | Yes | No | Premium Edition |
Laotie | laotie | Northeastern buddy | Livestreaming scenarios | Supports Chinese-only scenarios | 8K/16K | Yes | Yes | Standard Edition |
Younger sister | laomei | Female Hawker | Livestreaming scenarios | Supports Chinese-only scenarios | 8K or 16K | Yes | Yes | Standard Edition |
Aikan | aikan | Tianjin dialect male voice | Dialect scenarios | Supports Chinese-only scenarios | 8K/16K | Yes | Yes | Standard Edition |
Tala | tala | Filipino female voice | Multilingual scenarios | Supports Filipino-only scenarios | 8K or 16K | No | No | Standard Edition |
Zhitian | zhitian | Sweet female voice | General | Supports Chinese and Chinese-English mixed scenarios | 8K/16K | Yes | No | Premium Edition |
Zhiqing | zhiqing | Taiwan (China) dialect female voice | Dialect scenarios | Supports Chinese-only scenarios | 8K or 16K | Yes | No | Premium Edition |
Tien | tien | Vietnamese female voice | Multilingual scenarios | Supports Vietnamese-only scenarios | 8K/16K | No | No | Standard Edition |
Becca | becca | American English customer service female voice | American English | Supports English-only scenarios | 8K or 16K | No | No | Standard Edition |
Kyong | Kyong | Korean female voice | Korean scenarios | Korean | 8 kHz/16 kHz | No | No | Standard Edition |
masha | masha | Russian female voice | Russian scenarios | Russian | 8K/16K | No | No | Standard Edition |
camila | camila | Spanish female voice | Spanish scenarios | Spanish | 8 kHz/16 kHz | No | No | Standard Edition |
perla | perla | Italian female voice | Italian scenarios | Italian | 8 kHz/16 kHz | No | No | Standard Edition |
Zhimao | zhimao | Mandarin female voice | Livestreaming | Chinese | 8 kHz/16 kHz | Yes | No | Standard Edition |
Zhiyuan | zhiyuan | Mandarin female voice | General | Chinese | 8 kHz/16 kHz | Yes | No | Standard Edition |
Zhiya | zhiya | Mandarin female voice | Customer service | Chinese | 8 kHz/16 kHz | Yes | No | Standard Edition |
Zhiyue | zhiyue | Mandarin female voice | General | Chinese | 8 kHz/16 kHz | Yes | No | Standard Edition |
Zhida | zhida | Mandarin male voice | General | Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz | Yes | No | Standard Edition |
Zhi Sha | zhistella | Mandarin female voice | General | Chinese | 8 kHz/16 kHz | Yes | No | Standard Edition |
Kelly | kelly | Hong Kong Cantonese female voice | Dialect scenarios | Hong Kong Cantonese | 8 kHz/16 kHz | Yes | No | Standard Edition |
clara | clara | French female voice | General | French | 8 kHz/16 kHz | No | No | Standard Edition |
hanna | hanna | German female voice | General | German | 8 kHz/16 kHz | No | No | Standard Edition |
waan | waan | Thai female voice | General | Thai | 8 kHz/16 kHz | No | No | Standard Edition |
betty | betty | American English female voice | General | American English | 8 kHz/16 kHz | Yes | No | Standard Edition |
beth | beth | American English female voice | General | American English | 8 kHz/16 kHz | Yes | No | Standard Edition |
cindy | cindy | American English female voice | General | American English | 8 kHz/16 kHz | Yes | No | Standard Edition |
donna | donna | American English female voice | General | American English | 8 kHz/16 kHz | Yes | No | Standard Edition |
eva | eva | American English female voice | General | American English | 8 kHz/16 kHz | Yes | No | Standard Edition |
brian | brian | American English male voice | General | American English | 8 kHz/16 kHz | Yes | No | Standard Edition |
david | david | American English male voice | General | American English | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
abby_ecmix | abby_ecmix | American English female voice | General | English and English-Chinese mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
annie_ecmix | annie_ecmix | American English female voice | General | English and English-Chinese mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
andy_ecmix | andy_ecmix | American English male voice | General | English and English-Chinese mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
ava_ecmix | ava_ecmix | American English female voice | General | English and English-Chinese mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
betty_ecmix | betty_ecmix | American English female voice | General | English and English-Chinese mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
beth_ecmix | beth_ecmix | American English female voice | General | English and English-Chinese mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
brian_ecmix | brian_ecmix | American English male voice | General | English and English-Chinese mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
cindy_ecmix | cindy_ecmix | American English female voice | General | English and English-Chinese mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
cally_ecmix | cally_ecmix | American English female voice | General | English and English-Chinese mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
donna_ecmix | donna_ecmix | American English female voice | General | English and English-Chinese mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
david_ecmix | david_ecmix | American English male voice | General | English and English-Chinese mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
eva_ecmix | eva_ecmix | American English female voice | General | English and English-Chinese mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
Support for multi-emotional voices
Only multi-emotional voice models support emotion selection. The supported emotions are listed in the following table. The supported emotion categories vary by voice. The main categories include neutral, happy, angry, sad, fear, hate, surprise, arousal, serious, disgust, jealousy, embarrassed, frustrated, affectionate, gentle, newscast, customer service, and living.
Voice name | voice parameter value | Emotion category |
Zhifeng_Multi-emotional | zhifeng_emo | angry, fear, happy, neutral, sad, surprise |
Zhibing_Multi-emotional | zhibing_emo | angry, fear, happy, neutral, sad, surprise |
Zhimiao_Multi-emotional | zhimiao_emo | serious, sad, disgust, jealousy, embarrassed, happy, fear, surprise, neutral, frustrated, affectionate, gentle, angry, newscast, customer-service, story, living |
Zhimi_Multi-emotional | zhimi_emo | angry, fear, happy, hate, neutral, sad, surprise |
Zhiyan_Multi-emotional | zhiyan_emo | neutral, happy, angry, sad, fear, hate, surprise, arousal |
Zhibei_Multi-emotional | zhibei_emo | neutral, happy, angry, sad, fear, hate, surprise |
Zhitian_Multi-emotional | zhitian_emo | neutral, happy, angry, sad, fear, hate, surprise |
Call instructions
The input text must be
UTF-8encoded.Long text-to-speech and standard speech synthesis share many similarities.
Intelligent access to the nearest region
Speech Synthesis supports intelligent access to the nearest region using the domain name nls-gateway.aliyuncs.com.
Connect to the nearest region for optimal performance. The system automatically resolves requests to the server in the nearest region based on the client's geographical location. For example, a request from the China (Beijing) region is automatically routed to a server in the China (Beijing) region. This provides the same result as specifying the nls-gateway-cn-beijing.aliyuncs.com domain name.
Service endpoints
Access type | Description | URL |
Public access (defaults to the China (Shanghai) region) | All servers can use the public access URL. The SDK is configured with the public access URL by default. |
|
ECS private network access | If you use an Alibaba Cloud ECS instance in the China (Shanghai), China (Beijing), or China (Shenzhen) region, you can use the private network access URL. ECS instances in the classic network cannot access AnyTunnel and therefore cannot access the Voice Service over the private network. To use AnyTunnel, create a Virtual Private Cloud (VPC) and access the service from within the VPC. Note
|
|
Interaction flow
The preceding figure does not show the interaction flow for the RESTful API. For the RESTful API interaction flowchart, see RESTful API.
In addition to the audio stream, the server response header includes the `task_id` parameter. This parameter is the unique ID of the request.
To play the audio stream that is returned by the server in real time, use an audio player that supports stream playback, such as ffmpeg, PyAudio (Python), AudioFormat (Java), or MediaSource (JavaScript).
Authentication
The client uses a token for authentication when establishing a WebSocket connection with the server. For more information about obtaining a token, see Obtain a token.
Start synthesis
The client sends a speech synthesis request and configures parameters in the request message. You can set each parameter using the corresponding `set` method of the `SpeechSynthesizer` object in the SDK. The parameters are described in the following table.
Parameter
Type
Required
Description
appkey
String
Yes
The AppKey of the project that you created in the console.
text
String
Yes
The text to synthesize. The text must be
UTF-8encoded. Add a space between English words.NoteTo use a multi-emotional voice, add the `ssml-emotion` tag to the text. For more information, see <emotion>.
Only voices that support multiple emotions can use the <emotion> tag. Otherwise, the `Illegal ssml text` error is reported.
voice
String
No
The voice. The default value is xiaoyun.
format
String
No
The audio encoding format. Valid values: PCM, WAV, and MP3. Default value:
pcm.sample_rate
Integer
No
The audio sample rate. Default value: 16000 Hz.
volume
Integer
No
The volume. Valid values: 0 to 100. Default value: 50.
speech_rate
Integer
No
The speech rate. Valid values: -500 to 500. Default value: 0.
The range [-500, 0, 500] corresponds to a speed multiplier of [0.5, 1.0, 2.0].
-500 indicates 0.5 times the default speed.
0 indicates the default speed (1.0x). The default speed varies slightly for each voice but is approximately four characters per second.
500 indicates 2.0 times the default speed.
The calculation method is as follows:
0.8x speed: (1 - 1/0.8) / 0.002 = -125
1.2x speed: (1 - 1/1.2) / 0.001 = 166
ImportantUse a coefficient of 0.002 for speeds less than 1.0x.
Use a coefficient of 0.001 for speeds greater than 1.0x.
The actual algorithm result is an approximate value.
pitch_rate
Integer
No
The pitch. Valid values: -500 to 500. Default value: 0.
enable_subtitle
Boolean
No
Enables word-level timestamps. For more information about how to use this feature, see Timestamps.
Receive synthesized data
The server returns the synthesized speech as binary data. The SDK receives and processes the binary data.
End synthesis
After the speech is synthesized, the server sends a `SynthesisCompleted` event notification. The following code provides an example.
{ "header":{ "namespace":"SpeechLongSynthesizer", "name":"SynthesisCompleted", "status":20000000, "message_id":"396c80b3abf84082a48cb9e5c424****", "task_id":"f5805be640364cdcafc8da63e512****", "status_text":"Gateway:SUCCESS:Success." } }Handle synthesis failures
If the synthesis task fails due to incorrect parameters or other reasons, a `TaskFailed` notification is returned. The underlying connection is then closed. The following code provides an example.
{ "header":{ "namespace":"Default", "name":"TaskFailed", "status":41020001, "message_id":"62c126f7d9b340deb82b5b7eaca0****", "task_id":"4552df26d1f547aab9a2c4a94678****", "status_text":"TTS:TtsClientError:[tts]Engine return error code: 418" } }
Service status codes
Each service response includes a `status` field, which is the service status code. The following sections describe the meanings of different status codes.
General-purpose error codes
Status code | Status message | Cause | Solution |
40000000 | The default client error code. This code corresponds to multiple error messages. | Invalid parameters or call logic was used. | Compare your code with the sample code in the official documentation to test and verify it. |
40000001 | The token 'xxx' has expired. The token 'xxx' is invalid | Invalid parameters or call logic was used. This is a general-purpose client error code that usually indicates an incorrect token, such as an expired or invalid token. | Compare your code with the sample code in the official documentation to test and verify it. |
40000002 | Gateway:MESSAGE_INVALID:Can't process message in state'FAILED'! | The message is invalid or incorrect. | Compare your code with the sample code in the official documentation to test and verify it. |
40000003 | PARAMETER_INVALID Failed to decode url params | The parameters passed by the user are incorrect. This error is common for RESTful API calls. | Compare your code with the sample code in the official documentation to test and verify it. |
40000005 | Gateway:TOO_MANY_REQUESTS:Too many requests! | Too many concurrent requests. | If you are using the Free Edition, you can upgrade to a commercial version to increase the concurrency. If you are already using a commercial version, you can purchase a concurrency resource plan to increase your concurrency quota. |
40000009 | Invalid wav header! | The message header is invalid. | If you send a WAV audio file and set the |
40000009 | Too large wav header! | The WAV header of the transmitted audio is invalid. | You can send the audio stream in a format such as PCM or OPUS. If you use the WAV format, make sure that the WAV header of the audio file contains the correct data length. |
40000010 | Gateway:FREE_TRIAL_EXPIRED:The free trial has expired! | The trial period has ended, and the commercial version is not activated or your account has an overdue payment. | You can log on to the console to check the service activation status and your account balance. |
40010001 | Gateway:NAMESPACE_NOT_FOUND:RESTful url path illegal | The operation or parameter is not supported. | Check whether the parameters passed in the call are consistent with the requirements in the official documentation. You can compare them with the error message to identify and set the correct parameters. For example, if you are using a curl command to make a RESTful API request, check whether the URL you constructed is valid. |
40010003 | Gateway:DIRECTIVE_INVALID:[xxx] | A general-purpose client-side error code. | This error indicates that the client passed an incorrect parameter or instruction. Detailed error messages are available for different operations. You can refer to the corresponding documentation to set the parameters correctly. |
40010004 | Gateway:CLIENT_DISCONNECT:Client disconnected before task finished! | The client actively terminated the connection before the request was processed. | None. Alternatively, you can close the connection after the server responds. |
40010005 | Gateway:TASK_STATE_ERROR:Got stop directive while task is stopping! | The client sent a message instruction that is not currently supported. | Compare your code with the sample code in the official documentation to test and verify it. |
40020105 | Meta:APPKEY_NOT_EXIST:Appkey not exist! | A non-existent Appkey was used. | Confirm whether a non-existent Appkey was used. You can log on to the console and view the project configuration to find the Appkey. |
40020106 | Meta:APPKEY_UID_MISMATCH:Appkey and user mismatch! | The Appkey and token passed in the call were not created by the same Alibaba Cloud account UID. This causes a mismatch. | Check whether you are using resources from two different accounts. Do not use an Appkey from Account A with a token generated from Account B. |
403 | Forbidden | The token is invalid. For example, the token does not exist or has expired. | Set a valid token. Tokens have an expiration period. You must obtain a new token before the current one expires. |
41000003 | MetaInfo doesn't have end point info | Failed to retrieve the routing information for this Appkey. | Check whether you are using resources from two different accounts. Do not use an Appkey from Account A with a token generated from Account B. |
41010101 | UNSUPPORTED_SAMPLE_RATE | The sample rate is not supported. | Real-time speech recognition currently supports only audio with a sample rate of 8000 Hz or 16000 Hz. |
41040201 | Realtime:GET_CLIENT_DATA_TIMEOUT:Client data does not send continuously! | Failed to retrieve data from the client due to a timeout. | When you call real-time speech recognition, the client must send data at a real-time rate and close the connection promptly after the data is sent. |
50000000 | GRPC_ERROR:Grpc error! | An exception caused by factors such as machine load or network issues. This error usually occurs randomly. | You can retry the call to resolve the issue. |
50000001 | GRPC_ERROR:Grpc error! | An exception caused by factors such as machine load or network issues. This error usually occurs randomly. | You can retry the call to resolve the issue. |
52010001 | GRPC_ERROR:Grpc error! | An exception caused by factors such as machine load or network issues. This error usually occurs randomly. | You can retry the call to resolve the issue. |
Speech synthesis/Long text-to-speech error codes
Status code | Status message | Cause | Solution |
21050000 | SUCCESS | Success. | None. |
21050001 | RUNNING | The asynchronous long text-to-speech task is running. | You can send a GET request later to query the detection result. |
21050002 | QUEUEING | The asynchronous long text-to-speech task is in the queue. | Send a GET request later to query the detection result. |
40000001 | Gateway:ACCESS_DENIED:No privilege to this voice! | An invalid voice name was set. | Set a valid voice name based on the official documentation. |
40000004 | Gateway:IDLE_TIMEOUT:Websocket session is idle for too long time,the last directive is 'StartSynthesis'! | After a connection is established, no data is sent for an extended period. The server returns this error message after 10 seconds. | Close the connection promptly after the request is processed. This error may also occur when the server is under high momentary pressure and cannot return data in time. In this case, you can retry the call. |
40010003 | Gateway:DIRECTIVE_INVALID:No text specified! | No valid text was set for synthesis. | Set the text to be synthesized based on the sample code in the official documentation. |
41020001 | Speech synthesis call client-side error | Multiple error messages may be displayed. Resolve each error based on the corresponding message. |
|
51020001 | TTS:TtsServerError | An exception occurred due to factors such as machine load or network issues. This error is usually intermittent. | Retry the call. |