The long text-to-speech feature converts very long text, such as thousands or tens of thousands of characters, into binary audio data.
← Return to the Speech Synthesis product page
Billing and concurrency limits
Asynchronous long text-to-speech is available only in the commercial version and does not offer a free trial. For more information, see Free trial and commercial versions. To use this feature, you must upgrade to the commercial version. For more information, see Upgrade from the free trial to the commercial version.
For more information about billing methods, see Billing methods.
For more information about concurrency limits, see Concurrency and QPS.
New ultra-high-definition voices
The service now offers multiple ultra-high-definition (UHD) voices. They provide superior audio quality with a sample rate of up to 48 kHz for lossless, high-fidelity sound.
Listen to UHD voice samples:
To listen to more voice samples, go to the Speech Synthesis product page.
Features
Supports PCM, WAV, and MP3 encoding formats.
Supports adjustments for speech rate, pitch, and volume.
Supports male and female voices.
You can only retrieve synthesis results asynchronously.
The RESTful API supports sentence-level timestamps. For more information, see Timestamp feature.
Long text-to-speech offers unique advantages over standard speech synthesis:
Supports longer text input: Synthesize up to 100,000 characters at a time. One Chinese character, one English letter, one punctuation mark, or one space between words counts as a single character.
Fast synthesis speed: Synthesize 50,000 characters in as little as 10 minutes.
Reusable: The synthesized audio files can be cached on the client for repeated use.
Exclusive voices: Provides exclusive, high-quality voices tailored for specific scenarios, such as reading novels, news narration, and video dubbing.
After you submit a long text-to-speech request, the synthesis is completed within 3 hours. The audio file is stored on the server for 7 days.
Supports multi-emotion voices. For more information, see the <emotion> tag in Markup language. Tags are not counted as characters.
To use the long text-to-speech feature, you must update your software development kit (SDK) to the latest version.
Voice list
Name | voice parameter value | Type | Scenario | Supported languages | Supported sample rates (Hz) | Supports word/sentence-level timestamps | Supports r-sound | Voice quality |
Abin | abin | Guangdong Mandarin | Conversational digital human | Supports Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz/48 kHz | No | No | Standard Edition |
Zhixiaobai | zhixiaobai | Mandarin female voice | Conversational digital human | Supports Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz/48 kHz | No | Yes | Standard Edition |
Zhixiaoxia | zhixiaoxia | Mandarin female voice | Conversational digital human | Supports Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz/48 kHz | No | Yes | Standard Edition |
Zhixiaomei | zhixiaomei | Mandarin female voice | Livestreaming digital human | Supports Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | Yes | Standard Edition |
Zhigui | zhigui | Mandarin female voice | Livestreaming digital human | Supports Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz | Yes | Yes | Standard Edition |
Zhishuo | zhishuo | Mandarin male voice | Customer service digital human | Supports Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | Yes | Standard Edition |
Aixia | aixia | Mandarin female voice | Customer service digital human | Supports Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz | Yes | Yes | Standard Edition |
Cally | cally | American English female voice | Spoken English conversational digital human | Supports only English scenarios | 8K or 16K | Yes | Yes | Standard Edition |
Zhifeng_emo | zhifeng_emo | Multi-emotion male voice | General | Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | Yes | Standard Edition |
Zhibing_emo | zhibing_emo | Multi-emotion male voice | General | Chinese-only scenarios | 8 kHz/16 kHz/24 kHz | Yes | Yes | Standard Edition |
Zhimiao_emo | zhimiao_emo | Multi-emotion female voice | Chinese-English scenarios | Chinese and English scenarios | 8K/16K | Yes | Yes | Standard Edition |
Zhimi_emo | zhimi_emo | Multi-emotion female voice | General | Chinese and Chinese-English mixed scenarios | 8K/16K | Yes | No | Standard Edition |
Zhiyan_emo | zhiyan_emo | Multi-emotion female voice | General | Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | No | Standard Edition |
Zhibei_emo | zhibei_emo | Multi-emotion child voice | General | Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | No | Standard Edition |
Zhitian_emo | zhitian_emo | Multi-emotion female voice | General | Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | No | Standard Edition |
Xiaoyun | xiaoyun | Standard female voice | General | Chinese and Chinese-English mixed scenarios | 8K or 16K | No | No | Lite Edition |
Xiaogang | xiaogang | Standard male voice | General | Chinese and Chinese-English mixed scenarios | 8K or 16K | No | No | Lite Edition |
Ruoxi | ruoxi | Gentle female voice | General | Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz | No | No | Standard Edition |
Siqi | siqi | Gentle female voice | General | Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
Sijia | sijia | Standard female voice | General | Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz | No | No | Standard Edition |
Sicheng | sicheng | Standard male voice | General | Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
Aiqi | aiqi | Gentle female voice | General | Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | No | Standard Edition |
Aijia | aijia | Standard female voice | General | Chinese and Chinese-English mixed scenarios | 8K / 16K | Yes | No | Standard Edition |
Aicheng | aicheng | Standard male voice | General | Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | No | Standard Edition |
Aida | aida | Standard male voice | General | Chinese and Chinese-English mixed scenarios | 8 K/16 K | Yes | No | Standard Edition |
Ninger | ninger | Standard female voice | General | Chinese-only scenarios | 8 kHz/16 kHz/24 kHz | No | No | Standard Edition |
Ruilin | ruilin | Standard female voice | General | Chinese-only scenarios | 8 kHz/16 kHz/24 kHz | No | No | Standard Edition |
Siyue | siyue | Gentle female voice | Customer service | Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
Aiya | aiya | Stern female voice | Customer service | Chinese and Chinese-English mixed scenarios | 8K / 16K | Yes | No | Standard Edition |
Aimei | aimei | Sweet female voice | Customer service | Chinese and Chinese-English mixed scenarios | 8K/16K | Yes | No | Standard Edition |
Aiyu | aiyu | Natural female voice | Customer service | Chinese and Chinese-English mixed scenarios | 8K / 16K | Yes | No | Standard Edition |
Aiyue | aiyue | Gentle female voice | Customer service | Chinese and Chinese-English mixed scenarios | 8K/16K | Yes | No | Standard Edition |
Aijing | aijing | Stern female voice | Customer service | Chinese and Chinese-English mixed scenarios | 8 K / 16 K | Yes | No | Standard Edition |
Xiaomei | xiaomei | Sweet female voice | Customer service | Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz | No | No | Standard Edition |
Aina | aina | Zhejiang-accented Mandarin female voice | Customer service | Chinese-only scenarios | 8 K or 16 K | Yes | No | Standard Edition |
Yina | yina | Zhejiang-accented Mandarin female voice | Customer service | Chinese-only scenarios | 8 kHz/16 kHz/24 kHz | No | No | Standard Edition |
Sijing | sijing | Stern female voice | Customer service | Chinese-only scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
Sitong | sitong | Child voice | Child voice scenarios | Chinese-only scenarios | 8 kHz/16 kHz/24 kHz | No | No | Standard Edition |
Xiaobei | xiaobei | Young Girl's Voice | Child voice scenarios | Chinese-only scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
Aitong | aitong | Child voice | Child voice scenarios | Chinese-only scenarios | 8K or 16K | Yes | No | Standard Edition |
Aiwei | aiwei | Lolita female voice | Child voice scenarios | Chinese-only scenarios | 8K or 16K | Yes | No | Standard Edition |
Aibao | aibao | Sweet Girlish Voice | Child voice scenarios | Chinese-only scenarios | 8K or 16K | Yes | No | Standard Edition |
Harry | harry | British English male voice | English scenarios | English scenarios | 8K/16K | No | No | Standard Edition |
Abby | abby | American English female voice | English scenarios | English scenarios | 8K/16K | Yes | No | Standard Edition |
Andy | andy | American English male voice | English scenarios | English scenarios | 8K or 16K | Yes | No | Standard Edition |
Eric | eric | British English male voice | English scenarios | English scenarios | 8K/16K | Yes | No | Standard Edition |
Emily | emily | British English female voice | English scenarios | English scenarios | 8 K or 16 K | Yes | No | Standard Edition |
Luna | luna | British English female voice | English scenarios | English scenarios | 8 K/16 K | Yes | No | Standard Edition |
Luca | luca | British English male voice | English scenarios | English scenarios | 8 K or 16 K | Yes | No | Standard Edition |
Wendy | wendy | British English female voice | English scenarios | English scenarios | 8 kHz/16 kHz/24 kHz | No | No | Standard Edition |
William | william | British English male voice | English scenarios | English scenarios | 8 kHz/16 kHz/24 kHz | No | No | Standard Edition |
Olivia | olivia | British English female voice | English scenarios | English scenarios | 8 kHz/16 kHz/24 kHz | No | No | Standard Edition |
Shanshan | shanshan | Cantonese female voice | Dialect scenarios | Standard Cantonese (Simplified) and Cantonese-English mixed scenarios | 8 kHz/16 kHz/24 kHz | No | No | Standard Edition |
Aiyuan | aiyuan | Trusted Advisor | Literary scenarios | Chinese and Chinese-English mixed scenarios | 8 K/16 K | Yes | Yes | Premium Edition |
Aiying | aiying | Cute child voice | Literary scenarios | Chinese and Chinese-English mixed scenarios | 8 K/16 K | Yes | Yes | Premium Edition |
Aixiang | aixiang | Magnetic male voice | Literary scenarios | Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | Yes | Premium Edition |
Aimo | aimo | Emotional male voice | Literary scenarios | Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz | Yes | Yes | Premium Edition |
Aiye | aiye | Young male voice | Literary scenarios | Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | Yes | Premium Edition |
Aiting | aiting | Radio-style female voice | Literary scenarios | Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | Yes | Premium Edition |
Aifan | aifan | Emotional female voice | Literary scenarios | Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | Yes | Premium Edition |
Lydia | lydia | Bilingual (English-Chinese) female voice | English scenarios | English and English-Chinese mixed scenarios | 8 kHz/16 kHz | Yes | No | Standard Edition |
Xiaoyue | chuangirl | Sichuanese female voice | Dialect scenarios | Chinese and Chinese-English mixed scenarios | 8K or 16K | No | No | Standard Edition |
Aishuo | aishuo | Natural male voice | Customer service | Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | No | Standard Edition |
Qingqing | qingqing | Taiwanese Mandarin female voice | Dialect scenarios | Chinese-only scenarios | 8 kHz/16 kHz | No | No | Standard Edition |
Cuijie | cuijie | Northeastern Mandarin female voice | Dialect scenarios | Chinese-only scenarios | 8K or 16K | Yes | Yes | Standard Edition |
Xiaoze | xiaoze | Hunan-accented male voice | Dialect scenarios | Chinese-only scenarios | 8K/16K | No | No | Standard Edition |
Ainan | ainan | Advertisement-style male voice | Literary scenarios | Chinese and Chinese-English mixed scenarios | 8 K or 16 K | Yes | Yes | Premium Edition |
Aihao | aihao | News-style male voice | Literary scenarios | Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | Yes | Premium Edition |
Aiming | aiming | Humorous male voice | Literary scenarios | Chinese and Chinese-English mixed scenarios | 8K/16K | Yes | Yes | Premium Edition |
Aixiao | aixiao | News-style female voice | Literary scenarios | Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | Yes | Premium Edition |
Aichu | aichu | Food documentary-style male voice | Literary scenarios | Chinese and Chinese-English mixed scenarios | 8K/16K | Yes | Yes | Premium Edition |
Aiqian | aiqian | News-style female voice | Literary scenarios | Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | Yes | Premium Edition |
Tomoka | tomoka | Japanese female voice | Multi-language scenarios | Japanese-only scenarios | 8 K or 16 K | Yes | No | Standard Edition |
Tomoya | tomoya | Japanese male voice | Multi-language scenarios | Japanese-only scenarios | 8K or 16K | Yes | No | Standard Edition |
Annie | annie | American English female voice | English scenarios | Supports only English scenarios | 8K or 16K | Yes | No | Standard Edition |
Aishu | aishu | News-style male voice | Literary scenarios | Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | Yes | Premium Edition |
Airu | airu | Newscast female voice | Literary scenarios | Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | Yes | Premium Edition |
Jiajia | jiajia | Cantonese female voice | Dialect scenarios | Standard Cantonese (Simplified) and Cantonese-English mixed scenarios | 8K / 16K | Yes | No | Standard Edition |
Indah | indah | Indonesian female voice | Multi-language scenarios | Indonesian-only scenarios | 8 K/16 K | No | No | Standard Edition |
Peach | taozi | Cantonese female voice | Dialect scenarios | Supports Standard Cantonese (Simplified) and Cantonese-English mixed scenarios | 8K or 16K | Yes | No | Standard Edition |
Sales Associate | guijie | Friendly female voice | General | Supports Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz | Yes | Yes | Standard Edition |
Stella | stella | Intellectual female voice | General | Supports Chinese and Chinese-English mixed scenarios | 8K/16K | Yes | Yes | Standard Edition |
Stanley | stanley | Calm male voice | General | Supports Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | Yes | Standard Edition |
Kenny | kenny | Calm male voice | General | Supports Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | Yes | Standard Edition |
Rosa | rosa | Natural female voice | General | Supports Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | Yes | Standard Edition |
Farah | farah | Malay female voice | Multi-language scenarios | Supports only Malay scenarios | 8 K/16 K | No | No | Standard Edition |
Mashu | mashu | Children's drama male voice | General | General | 8K/16K | Yes | No | Standard Edition |
Zhiqi | zhiqi | Gentle female voice | UHD scenarios | Supports Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz/48 kHz | Yes | No | Premium Edition |
Zhichu | zhichu | Food documentary-style male voice | UHD scenarios | Supports Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz/48 kHz | Yes | Yes | Premium Edition |
Xiaoxian | xiaoxian | Friendly female voice | Livestreaming scenarios | Supports Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz | Yes | Yes | Standard Edition |
Yuer | yuer | Children's drama female voice | General | Supports only Chinese scenarios | 8K or 16K | Yes | No | Standard Edition |
Maoxiaomei | maoxiaomei | Energetic female voice | Livestreaming scenarios | Supports Chinese and Chinese-English mixed scenarios | 8 K/16 K | Yes | Yes | Standard Edition |
Zhixiang | zhixiang | Magnetic male voice | UHD scenarios | Supports Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz/48 kHz | Yes | No | Premium Edition |
Zhijia | zhijia | Standard female voice | UHD scenarios | Supports Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz/48 kHz | Yes | No | Premium Edition |
Zhinan | zhinan | Advertisement-style male voice | UHD scenarios | Supports Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz/48 kHz | Yes | No | Premium Edition |
Zhiqian | zhiqian | News-style female voice | UHD scenarios | Supports Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz/48 kHz | Yes | No | Premium Edition |
Zhiru | zhiru | Newscast female voice | UHD scenarios | Supports Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz/48 kHz | Yes | No | Premium Edition |
Zhide | zhide | Newscast male voice | UHD scenarios | Supports Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz/24 kHz/48 kHz | Yes | No | Premium Edition |
Zhifei | zhifei | Passionate narrator voice | UHD scenarios | Supports Chinese and Chinese-English mixed scenarios | 8K/16K | Yes | No | Premium Edition |
Aifei | aifei | Passionate narrator voice | Livestreaming scenarios | Supports Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | Yes | Standard Edition |
Subgroup | yaqun | Store broadcast voice | Livestreaming scenarios | Supports Chinese and Chinese-English mixed scenarios | 8K/16K | Yes | Yes | Standard Edition |
Qiaowei | qiaowei | Store broadcast voice | Livestreaming scenarios | Supports Chinese and Chinese-English mixed scenarios | 8K/16K | Yes | Yes | Standard Edition |
Dahu | dahu | Northeastern Mandarin male voice | Dialect scenarios | Supports Chinese and Chinese-English mixed scenarios | 8 K / 16 K | Yes | Yes | Standard Edition |
ava | ava | American English female voice | English scenarios | Supports only English scenarios | 8 K/16 K | Yes | No | Standard Edition |
Zhilun | zhilun | Suspense narrator voice | UHD scenarios | Supports Chinese and Chinese-English mixed scenarios | 8K or 16K | Yes | No | Premium Edition |
Ailun | ailun | Suspense narrator voice | Livestreaming scenarios | Supports Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz | Yes | Yes | Standard Edition |
Jielidou | jielidou | Healing a Child's Voice | Child voice scenarios | Supports only Chinese scenarios | 8K / 16K | Yes | Yes | Standard Edition |
Zhiwei | zhiwei | Childlike Female Voice | UHD scenarios | Supports only Chinese scenarios | 8 kHz/16 kHz/24 kHz/48 kHz | Yes | No | Premium Edition |
Laotie | laotie | Northeastern Buddy | Livestreaming scenarios | Supports only Chinese scenarios | 8K/16K | Yes | Yes | Standard Edition |
Younger sister | laomei | Hawking female voice | Livestreaming scenarios | Supports only Chinese scenarios | 8 K/16 K | Yes | Yes | Standard Edition |
Aikan | aikan | Tianjin-accented male voice | Dialect scenarios | Supports only Chinese scenarios | 8 K/16 K | Yes | Yes | Standard Edition |
Tala | tala | Filipino female voice | Multi-language scenarios | Supports only Filipino scenarios | 8K or 16K | No | No | Standard Edition |
Zhitian | zhitian | Sweet female voice | General | Supports Chinese and Chinese-English mixed scenarios | 8 K/16 K | Yes | No | Premium Edition |
Zhiqing | zhiqing | Taiwanese Mandarin female voice | Dialect scenarios | Supports only Chinese scenarios | 8K / 16K | Yes | No | Premium Edition |
Tien | tien | Vietnamese female voice | Multi-language scenarios | Supports only Vietnamese scenarios | 8 K/16 K | No | No | Standard Edition |
Becca | becca | American English customer service female voice | American English | Supports only English scenarios | 8K or 16K | No | No | Standard Edition |
Kyong | Kyong | Korean female voice | Korean scenarios | Korean | 8 kHz/16 kHz | No | No | Standard Edition |
masha | masha | Russian female voice | Russian scenarios | Russian | 8K/16K | No | No | Standard Edition |
camila | camila | Spanish female voice | Spanish scenarios | Spanish | 8 kHz/16 kHz | No | No | Standard Edition |
perla | perla | Italian female voice | Italian scenarios | Italian | 8 kHz/16 kHz | No | No | Standard Edition |
Zhimao | zhimao | Mandarin female voice | Livestreaming | Chinese | 8 kHz/16 kHz | Yes | No | Standard Edition |
Zhiyuan | zhiyuan | Mandarin female voice | General | Chinese | 8 kHz/16 kHz | Yes | No | Standard Edition |
Zhiya | zhiya | Mandarin female voice | Customer service | Chinese | 8 kHz/16 kHz | Yes | No | Standard Edition |
Zhiyue | zhiyue | Mandarin female voice | General | Chinese | 8 kHz/16 kHz | Yes | No | Standard Edition |
Zhida | zhida | Mandarin male voice | General | Chinese and Chinese-English mixed scenarios | 8 kHz/16 kHz | Yes | No | Standard Edition |
Zhisha | zhistella | Mandarin female voice | General | Chinese | 8 kHz/16 kHz | Yes | No | Standard Edition |
Kelly | kelly | Hong Kong Cantonese female voice | Dialect scenarios | Hong Kong Cantonese | 8 kHz/16 kHz | Yes | No | Standard Edition |
clara | clara | French female voice | General | French | 8 kHz/16 kHz | No | No | Standard Edition |
hanna | hanna | German female voice | General | German | 8 kHz/16 kHz | No | No | Standard Edition |
waan | waan | Thai female voice | General | Thai | 8 kHz/16 kHz | No | No | Standard Edition |
betty | betty | American English female voice | General | American English | 8 kHz/16 kHz | Yes | No | Standard Edition |
beth | beth | American English female voice | General | American English | 8 kHz/16 kHz | Yes | No | Standard Edition |
cindy | cindy | American English female voice | General | American English | 8 kHz/16 kHz | Yes | No | Standard Edition |
donna | donna | American English female voice | General | American English | 8 kHz/16 kHz | Yes | No | Standard Edition |
eva | eva | American English female voice | General | American English | 8 kHz/16 kHz | Yes | No | Standard Edition |
brian | brian | American English male voice | General | American English | 8 kHz/16 kHz | Yes | No | Standard Edition |
david | david | American English male voice | General | American English | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
abby_ecmix | abby_ecmix | American English female voice | General | English and English-Chinese mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
annie_ecmix | annie_ecmix | American English female voice | General | English and English-Chinese mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
andy_ecmix | andy_ecmix | American English male voice | General | English and English-Chinese mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
ava_ecmix | ava_ecmix | American English female voice | General | English and English-Chinese mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
betty_ecmix | betty_ecmix | American English female voice | General | English and English-Chinese mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
beth_ecmix | beth_ecmix | American English female voice | General | English and English-Chinese mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
brian_ecmix | brian_ecmix | American English male voice | General | English and English-Chinese mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
cindy_ecmix | cindy_ecmix | American English female voice | General | English and English-Chinese mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
cally_ecmix | cally_ecmix | American English female voice | General | English and English-Chinese mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
donna_ecmix | donna_ecmix | American English female voice | General | English and English-Chinese mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
david_ecmix | david_ecmix | American English male voice | General | English and English-Chinese mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
eva_ecmix | eva_ecmix | American English female voice | General | English and English-Chinese mixed scenarios | 8 kHz/16 kHz/24 kHz | Yes | No | Standard Edition |
Multi-emotion voice support
Only multi-emotion voice models support emotion selection. The supported emotions are listed in the following table. The supported emotion categories vary by voice and include neutral, happy, angry, sad, fear, hate, surprise, arousal, serious, disgust, jealousy, embarrassed, frustrated, affectionate, gentle, newscast, customer-service, story, and lively.
Voice name | voice parameter value | Emotion categorization |
Zhifeng_Multi-emotion | zhifeng_emo | angry, fear, happy, neutral, sad, surprise |
Zhibing_Multi-emotion | zhibing_emo | angry, fear, happy, neutral, sad, surprise |
Zhimiao_Multi-emotion | zhimiao_emo | serious, sad, disgust, jealousy, embarrassed, happy, fear, surprise, neutral, frustrated, affectionate, gentle, angry, newscast, customer-service, story, living |
Zhimi_Multi-emotion | zhimi_emo | angry, fear, happy, hate, neutral, sad, surprise |
Zhiyan_Multi-emotion | zhiyan_emo | neutral, happy, angry, sad, fear, hate, surprise, arousal |
Zhibei_Multi-emotion | zhibei_emo | neutral, happy, angry, sad, fear, hate, surprise |
Zhitian_Multi-emotion | zhitian_emo | neutral, happy, angry, sad, fear, hate, surprise |
Call instructions
The input text must be
UTF-8encoded.Long text-to-speech and standard speech synthesis share many similarities.
Service endpoints
Access type | Description | URL | Host |
Public network access | All servers can use the public endpoint URL. The SDK is configured with the public endpoint URL by default. | https://nls-gateway-cn-shanghai.aliyuncs.com/rest/v1/tts/async | nls-gateway-cn-shanghai.aliyuncs.com |
Internal access from Alibaba Cloud ECS in Shanghai | If you use an Alibaba Cloud ECS instance in the China (Shanghai) region, you can use the internal endpoint URL. ECS instances in the classic network cannot access AnyTunnel and therefore cannot access Voice Service over the internal network. To use AnyTunnel, create a VPC and access the service from within the VPC. Note
| http://nls-gateway-cn-shanghai-internal.aliyuncs.com/rest/v1/tts/async | nls-gateway-cn-shanghai-internal.aliyuncs.com |
Interaction flow
The client sends an HTTPS POST request containing text to the server, and the server returns a response. The client can then handle the response in two ways:
Poll the synthesis status until the task is complete.
Wait for the server to complete the synthesis and send a callback to the webhook address that you configured. Your client program can then perform subsequent actions.
Unlike the RESTful API for standard speech synthesis, the RESTful API for long text-to-speech does not directly return the synthesized audio data to the client. Instead, it returns an HTTP URL that you can use to download or play the audio file.
The server response header also includes the `task_id` parameter, which is the unique ID for the request.
Request parameters
The request parameters for asynchronous long text-to-speech are described in the following table.
When you send an HTTP/HTTPS request, include these parameters in the request body.
Name | Type | Required | Description |
appkey | String | Yes | The AppKey of your application. For more information, see Manage projects. |
token | String | No | The authentication token for the service. |
text | String | Yes | The text to synthesize. The text must be Note To use the multi-emotion feature of a voice, add the SSML emotion tag to the text. For more information, see <emotion>. Only voices that support multiple emotions can use the <emotion> tag. Otherwise, an `Illegal ssml text` error is reported. |
format | String | Yes | The audio encoding format. Valid values: |
sample_rate | Integer | Yes | The audio sample rate. Valid values: 16000 Hz and 8000 Hz. Default value: 16000 Hz. |
voice | String | No | The voice. Default value: xiaoyun. For more voices, see Voice list. |
volume | Integer | No | The volume. The value must be in the range of 0 to 100. Default value: 50. |
speech_rate | Integer | No | The speech rate. The value must be in the range of -500 to 500. Default value: 0. |
pitch_rate | Integer | No | The pitch. The value must be in the range of -500 to 500. Default value: 0. |
enable_subtitle | Boolean | No | Specifies whether to enable the sentence-level timestamp feature. Default value: false. |
enable_notify | Boolean | Yes | Specifies whether to enable the callback feature. Default value: false. |
notify_url | String | No | The webhook address for callbacks. This parameter is required if The URL must use the HTTP or HTTPS protocol. The host cannot be an IP address. |
Service status codes
Each service response includes a `status` field, which is the service status code. The following sections describe the meanings of different status codes.
General-purpose error codes
Status code | Status message | Cause | Solution |
40000000 | The default client error code. This code corresponds to multiple error messages. | Invalid parameters or call logic was used. | Compare your code with the sample code in the official documentation to test and verify it. |
40000001 | The token 'xxx' has expired. The token 'xxx' is invalid | Invalid parameters or call logic was used. This is a general-purpose client error code that usually indicates an incorrect token, such as an expired or invalid token. | Compare your code with the sample code in the official documentation to test and verify it. |
40000002 | Gateway:MESSAGE_INVALID:Can't process message in state'FAILED'! | The message is invalid or incorrect. | Compare your code with the sample code in the official documentation to test and verify it. |
40000003 | PARAMETER_INVALID Failed to decode url params | The parameters passed by the user are incorrect. This error is common for RESTful API calls. | Compare your code with the sample code in the official documentation to test and verify it. |
40000005 | Gateway:TOO_MANY_REQUESTS:Too many requests! | Too many concurrent requests. | If you are using the Free Edition, you can upgrade to a commercial version to increase the concurrency. If you are already using a commercial version, you can purchase a concurrency resource plan to increase your concurrency quota. |
40000009 | Invalid wav header! | The message header is invalid. | If you send a WAV audio file and set the |
40000009 | Too large wav header! | The WAV header of the transmitted audio is invalid. | You can send the audio stream in a format such as PCM or OPUS. If you use the WAV format, make sure that the WAV header of the audio file contains the correct data length. |
40000010 | Gateway:FREE_TRIAL_EXPIRED:The free trial has expired! | The trial period has ended, and the commercial version is not activated or your account has an overdue payment. | You can log on to the console to check the service activation status and your account balance. |
40010001 | Gateway:NAMESPACE_NOT_FOUND:RESTful url path illegal | The operation or parameter is not supported. | Check whether the parameters passed in the call are consistent with the requirements in the official documentation. You can compare them with the error message to identify and set the correct parameters. For example, if you are using a curl command to make a RESTful API request, check whether the URL you constructed is valid. |
40010003 | Gateway:DIRECTIVE_INVALID:[xxx] | A general-purpose client-side error code. | This error indicates that the client passed an incorrect parameter or instruction. Detailed error messages are available for different operations. You can refer to the corresponding documentation to set the parameters correctly. |
40010004 | Gateway:CLIENT_DISCONNECT:Client disconnected before task finished! | The client actively terminated the connection before the request was processed. | None. Alternatively, you can close the connection after the server responds. |
40010005 | Gateway:TASK_STATE_ERROR:Got stop directive while task is stopping! | The client sent a message instruction that is not currently supported. | Compare your code with the sample code in the official documentation to test and verify it. |
40020105 | Meta:APPKEY_NOT_EXIST:Appkey not exist! | A non-existent Appkey was used. | Confirm whether a non-existent Appkey was used. You can log on to the console and view the project configuration to find the Appkey. |
40020106 | Meta:APPKEY_UID_MISMATCH:Appkey and user mismatch! | The Appkey and token passed in the call were not created by the same Alibaba Cloud account UID. This causes a mismatch. | Check whether you are using resources from two different accounts. Do not use an Appkey from Account A with a token generated from Account B. |
403 | Forbidden | The token is invalid. For example, the token does not exist or has expired. | Set a valid token. Tokens have an expiration period. You must obtain a new token before the current one expires. |
41000003 | MetaInfo doesn't have end point info | Failed to retrieve the routing information for this Appkey. | Check whether you are using resources from two different accounts. Do not use an Appkey from Account A with a token generated from Account B. |
41010101 | UNSUPPORTED_SAMPLE_RATE | The sample rate is not supported. | Real-time speech recognition currently supports only audio with a sample rate of 8000 Hz or 16000 Hz. |
41040201 | Realtime:GET_CLIENT_DATA_TIMEOUT:Client data does not send continuously! | Failed to retrieve data from the client due to a timeout. | When you call real-time speech recognition, the client must send data at a real-time rate and close the connection promptly after the data is sent. |
50000000 | GRPC_ERROR:Grpc error! | An exception caused by factors such as machine load or network issues. This error usually occurs randomly. | You can retry the call to resolve the issue. |
50000001 | GRPC_ERROR:Grpc error! | An exception caused by factors such as machine load or network issues. This error usually occurs randomly. | You can retry the call to resolve the issue. |
52010001 | GRPC_ERROR:Grpc error! | An exception caused by factors such as machine load or network issues. This error usually occurs randomly. | You can retry the call to resolve the issue. |
Speech synthesis/Long text-to-speech error codes
Status code | Status message | Cause | Solution |
21050000 | SUCCESS | Success. | None. |
21050001 | RUNNING | The asynchronous long text-to-speech task is running. | You can send a GET request later to query the detection result. |
21050002 | QUEUEING | The asynchronous long text-to-speech task is in the queue. | Send a GET request later to query the detection result. |
40000001 | Gateway:ACCESS_DENIED:No privilege to this voice! | An invalid voice name was set. | Set a valid voice name based on the official documentation. |
40000004 | Gateway:IDLE_TIMEOUT:Websocket session is idle for too long time,the last directive is 'StartSynthesis'! | After a connection is established, no data is sent for an extended period. The server returns this error message after 10 seconds. | Close the connection promptly after the request is processed. This error may also occur when the server is under high momentary pressure and cannot return data in time. In this case, you can retry the call. |
40010003 | Gateway:DIRECTIVE_INVALID:No text specified! | No valid text was set for synthesis. | Set the text to be synthesized based on the sample code in the official documentation. |
41020001 | Speech synthesis call client-side error | Multiple error messages may be displayed. Resolve each error based on the corresponding message. |
|
51020001 | TTS:TtsServerError | An exception occurred due to factors such as machine load or network issues. This error is usually intermittent. | Retry the call. |