API reference

更新时间:
复制 MD 格式

The long text-to-speech feature converts very long text, such as thousands or tens of thousands of characters, into binary audio data.

← Return to the Speech Synthesis product page

Billing and concurrency limits

New ultra-high-definition voices

The service now offers multiple ultra-high-definition (UHD) voices. They provide superior audio quality with a sample rate of up to 48 kHz for lossless, high-fidelity sound.

Listen to UHD voice samples:

Zhiqi (zhiqi)

Zhichu (zhichu)

To listen to more voice samples, go to the Speech Synthesis product page.

Features

  • Supports PCM, WAV, and MP3 encoding formats.

  • Supports adjustments for speech rate, pitch, and volume.

  • Supports male and female voices.

  • You can only retrieve synthesis results asynchronously.

  • The RESTful API supports sentence-level timestamps. For more information, see Timestamp feature.

  • Long text-to-speech offers unique advantages over standard speech synthesis:

    • Supports longer text input: Synthesize up to 100,000 characters at a time. One Chinese character, one English letter, one punctuation mark, or one space between words counts as a single character.

    • Fast synthesis speed: Synthesize 50,000 characters in as little as 10 minutes.

    • Reusable: The synthesized audio files can be cached on the client for repeated use.

    • Exclusive voices: Provides exclusive, high-quality voices tailored for specific scenarios, such as reading novels, news narration, and video dubbing.

  • After you submit a long text-to-speech request, the synthesis is completed within 3 hours. The audio file is stored on the server for 7 days.

  • Supports multi-emotion voices. For more information, see the <emotion> tag in Markup language. Tags are not counted as characters.

Important

To use the long text-to-speech feature, you must update your software development kit (SDK) to the latest version.

Voice list

Name

voice parameter value

Type

Scenario

Supported languages

Supported sample rates (Hz)

Supports word/sentence-level timestamps

Supports r-sound

Voice quality

Abin

abin

Guangdong Mandarin

Conversational digital human

Supports Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz/48 kHz

No

No

Standard Edition

Zhixiaobai

zhixiaobai

Mandarin female voice

Conversational digital human

Supports Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz/48 kHz

No

Yes

Standard Edition

Zhixiaoxia

zhixiaoxia

Mandarin female voice

Conversational digital human

Supports Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz/48 kHz

No

Yes

Standard Edition

Zhixiaomei

zhixiaomei

Mandarin female voice

Livestreaming digital human

Supports Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

Yes

Standard Edition

Zhigui

zhigui

Mandarin female voice

Livestreaming digital human

Supports Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz

Yes

Yes

Standard Edition

Zhishuo

zhishuo

Mandarin male voice

Customer service digital human

Supports Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

Yes

Standard Edition

Aixia

aixia

Mandarin female voice

Customer service digital human

Supports Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz

Yes

Yes

Standard Edition

Cally

cally

American English female voice

Spoken English conversational digital human

Supports only English scenarios

8K or 16K

Yes

Yes

Standard Edition

Zhifeng_emo

zhifeng_emo

Multi-emotion male voice

General

Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

Yes

Standard Edition

Zhibing_emo

zhibing_emo

Multi-emotion male voice

General

Chinese-only scenarios

8 kHz/16 kHz/24 kHz

Yes

Yes

Standard Edition

Zhimiao_emo

zhimiao_emo

Multi-emotion female voice

Chinese-English scenarios

Chinese and English scenarios

8K/16K

Yes

Yes

Standard Edition

Zhimi_emo

zhimi_emo

Multi-emotion female voice

General

Chinese and Chinese-English mixed scenarios

8K/16K

Yes

No

Standard Edition

Zhiyan_emo

zhiyan_emo

Multi-emotion female voice

General

Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

No

Standard Edition

Zhibei_emo

zhibei_emo

Multi-emotion child voice

General

Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

No

Standard Edition

Zhitian_emo

zhitian_emo

Multi-emotion female voice

General

Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

No

Standard Edition

Xiaoyun

xiaoyun

Standard female voice

General

Chinese and Chinese-English mixed scenarios

8K or 16K

No

No

Lite Edition

Xiaogang

xiaogang

Standard male voice

General

Chinese and Chinese-English mixed scenarios

8K or 16K

No

No

Lite Edition

Ruoxi

ruoxi

Gentle female voice

General

Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz

No

No

Standard Edition

Siqi

siqi

Gentle female voice

General

Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

Sijia

sijia

Standard female voice

General

Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz

No

No

Standard Edition

Sicheng

sicheng

Standard male voice

General

Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

Aiqi

aiqi

Gentle female voice

General

Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

No

Standard Edition

Aijia

aijia

Standard female voice

General

Chinese and Chinese-English mixed scenarios

8K / 16K

Yes

No

Standard Edition

Aicheng

aicheng

Standard male voice

General

Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

No

Standard Edition

Aida

aida

Standard male voice

General

Chinese and Chinese-English mixed scenarios

8 K/16 K

Yes

No

Standard Edition

Ninger

ninger

Standard female voice

General

Chinese-only scenarios

8 kHz/16 kHz/24 kHz

No

No

Standard Edition

Ruilin

ruilin

Standard female voice

General

Chinese-only scenarios

8 kHz/16 kHz/24 kHz

No

No

Standard Edition

Siyue

siyue

Gentle female voice

Customer service

Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

Aiya

aiya

Stern female voice

Customer service

Chinese and Chinese-English mixed scenarios

8K / 16K

Yes

No

Standard Edition

Aimei

aimei

Sweet female voice

Customer service

Chinese and Chinese-English mixed scenarios

8K/16K

Yes

No

Standard Edition

Aiyu

aiyu

Natural female voice

Customer service

Chinese and Chinese-English mixed scenarios

8K / 16K

Yes

No

Standard Edition

Aiyue

aiyue

Gentle female voice

Customer service

Chinese and Chinese-English mixed scenarios

8K/16K

Yes

No

Standard Edition

Aijing

aijing

Stern female voice

Customer service

Chinese and Chinese-English mixed scenarios

8 K / 16 K

Yes

No

Standard Edition

Xiaomei

xiaomei

Sweet female voice

Customer service

Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz

No

No

Standard Edition

Aina

aina

Zhejiang-accented Mandarin female voice

Customer service

Chinese-only scenarios

8 K or 16 K

Yes

No

Standard Edition

Yina

yina

Zhejiang-accented Mandarin female voice

Customer service

Chinese-only scenarios

8 kHz/16 kHz/24 kHz

No

No

Standard Edition

Sijing

sijing

Stern female voice

Customer service

Chinese-only scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

Sitong

sitong

Child voice

Child voice scenarios

Chinese-only scenarios

8 kHz/16 kHz/24 kHz

No

No

Standard Edition

Xiaobei

xiaobei

Young Girl's Voice

Child voice scenarios

Chinese-only scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

Aitong

aitong

Child voice

Child voice scenarios

Chinese-only scenarios

8K or 16K

Yes

No

Standard Edition

Aiwei

aiwei

Lolita female voice

Child voice scenarios

Chinese-only scenarios

8K or 16K

Yes

No

Standard Edition

Aibao

aibao

Sweet Girlish Voice

Child voice scenarios

Chinese-only scenarios

8K or 16K

Yes

No

Standard Edition

Harry

harry

British English male voice

English scenarios

English scenarios

8K/16K

No

No

Standard Edition

Abby

abby

American English female voice

English scenarios

English scenarios

8K/16K

Yes

No

Standard Edition

Andy

andy

American English male voice

English scenarios

English scenarios

8K or 16K

Yes

No

Standard Edition

Eric

eric

British English male voice

English scenarios

English scenarios

8K/16K

Yes

No

Standard Edition

Emily

emily

British English female voice

English scenarios

English scenarios

8 K or 16 K

Yes

No

Standard Edition

Luna

luna

British English female voice

English scenarios

English scenarios

8 K/16 K

Yes

No

Standard Edition

Luca

luca

British English male voice

English scenarios

English scenarios

8 K or 16 K

Yes

No

Standard Edition

Wendy

wendy

British English female voice

English scenarios

English scenarios

8 kHz/16 kHz/24 kHz

No

No

Standard Edition

William

william

British English male voice

English scenarios

English scenarios

8 kHz/16 kHz/24 kHz

No

No

Standard Edition

Olivia

olivia

British English female voice

English scenarios

English scenarios

8 kHz/16 kHz/24 kHz

No

No

Standard Edition

Shanshan

shanshan

Cantonese female voice

Dialect scenarios

Standard Cantonese (Simplified) and Cantonese-English mixed scenarios

8 kHz/16 kHz/24 kHz

No

No

Standard Edition

Aiyuan

aiyuan

Trusted Advisor

Literary scenarios

Chinese and Chinese-English mixed scenarios

8 K/16 K

Yes

Yes

Premium Edition

Aiying

aiying

Cute child voice

Literary scenarios

Chinese and Chinese-English mixed scenarios

8 K/16 K

Yes

Yes

Premium Edition

Aixiang

aixiang

Magnetic male voice

Literary scenarios

Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

Yes

Premium Edition

Aimo

aimo

Emotional male voice

Literary scenarios

Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz

Yes

Yes

Premium Edition

Aiye

aiye

Young male voice

Literary scenarios

Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

Yes

Premium Edition

Aiting

aiting

Radio-style female voice

Literary scenarios

Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

Yes

Premium Edition

Aifan

aifan

Emotional female voice

Literary scenarios

Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

Yes

Premium Edition

Lydia

lydia

Bilingual (English-Chinese) female voice

English scenarios

English and English-Chinese mixed scenarios

8 kHz/16 kHz

Yes

No

Standard Edition

Xiaoyue

chuangirl

Sichuanese female voice

Dialect scenarios

Chinese and Chinese-English mixed scenarios

8K or 16K

No

No

Standard Edition

Aishuo

aishuo

Natural male voice

Customer service

Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

No

Standard Edition

Qingqing

qingqing

Taiwanese Mandarin female voice

Dialect scenarios

Chinese-only scenarios

8 kHz/16 kHz

No

No

Standard Edition

Cuijie

cuijie

Northeastern Mandarin female voice

Dialect scenarios

Chinese-only scenarios

8K or 16K

Yes

Yes

Standard Edition

Xiaoze

xiaoze

Hunan-accented male voice

Dialect scenarios

Chinese-only scenarios

8K/16K

No

No

Standard Edition

Ainan

ainan

Advertisement-style male voice

Literary scenarios

Chinese and Chinese-English mixed scenarios

8 K or 16 K

Yes

Yes

Premium Edition

Aihao

aihao

News-style male voice

Literary scenarios

Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

Yes

Premium Edition

Aiming

aiming

Humorous male voice

Literary scenarios

Chinese and Chinese-English mixed scenarios

8K/16K

Yes

Yes

Premium Edition

Aixiao

aixiao

News-style female voice

Literary scenarios

Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

Yes

Premium Edition

Aichu

aichu

Food documentary-style male voice

Literary scenarios

Chinese and Chinese-English mixed scenarios

8K/16K

Yes

Yes

Premium Edition

Aiqian

aiqian

News-style female voice

Literary scenarios

Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

Yes

Premium Edition

Tomoka

tomoka

Japanese female voice

Multi-language scenarios

Japanese-only scenarios

8 K or 16 K

Yes

No

Standard Edition

Tomoya

tomoya

Japanese male voice

Multi-language scenarios

Japanese-only scenarios

8K or 16K

Yes

No

Standard Edition

Annie

annie

American English female voice

English scenarios

Supports only English scenarios

8K or 16K

Yes

No

Standard Edition

Aishu

aishu

News-style male voice

Literary scenarios

Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

Yes

Premium Edition

Airu

airu

Newscast female voice

Literary scenarios

Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

Yes

Premium Edition

Jiajia

jiajia

Cantonese female voice

Dialect scenarios

Standard Cantonese (Simplified) and Cantonese-English mixed scenarios

8K / 16K

Yes

No

Standard Edition

Indah

indah

Indonesian female voice

Multi-language scenarios

Indonesian-only scenarios

8 K/16 K

No

No

Standard Edition

Peach

taozi

Cantonese female voice

Dialect scenarios

Supports Standard Cantonese (Simplified) and Cantonese-English mixed scenarios

8K or 16K

Yes

No

Standard Edition

Sales Associate

guijie

Friendly female voice

General

Supports Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz

Yes

Yes

Standard Edition

Stella

stella

Intellectual female voice

General

Supports Chinese and Chinese-English mixed scenarios

8K/16K

Yes

Yes

Standard Edition

Stanley

stanley

Calm male voice

General

Supports Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

Yes

Standard Edition

Kenny

kenny

Calm male voice

General

Supports Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

Yes

Standard Edition

Rosa

rosa

Natural female voice

General

Supports Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

Yes

Standard Edition

Farah

farah

Malay female voice

Multi-language scenarios

Supports only Malay scenarios

8 K/16 K

No

No

Standard Edition

Mashu

mashu

Children's drama male voice

General

General

8K/16K

Yes

No

Standard Edition

Zhiqi

zhiqi

Gentle female voice

UHD scenarios

Supports Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz/48 kHz

Yes

No

Premium Edition

Zhichu

zhichu

Food documentary-style male voice

UHD scenarios

Supports Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz/48 kHz

Yes

Yes

Premium Edition

Xiaoxian

xiaoxian

Friendly female voice

Livestreaming scenarios

Supports Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz

Yes

Yes

Standard Edition

Yuer

yuer

Children's drama female voice

General

Supports only Chinese scenarios

8K or 16K

Yes

No

Standard Edition

Maoxiaomei

maoxiaomei

Energetic female voice

Livestreaming scenarios

Supports Chinese and Chinese-English mixed scenarios

8 K/16 K

Yes

Yes

Standard Edition

Zhixiang

zhixiang

Magnetic male voice

UHD scenarios

Supports Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz/48 kHz

Yes

No

Premium Edition

Zhijia

zhijia

Standard female voice

UHD scenarios

Supports Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz/48 kHz

Yes

No

Premium Edition

Zhinan

zhinan

Advertisement-style male voice

UHD scenarios

Supports Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz/48 kHz

Yes

No

Premium Edition

Zhiqian

zhiqian

News-style female voice

UHD scenarios

Supports Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz/48 kHz

Yes

No

Premium Edition

Zhiru

zhiru

Newscast female voice

UHD scenarios

Supports Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz/48 kHz

Yes

No

Premium Edition

Zhide

zhide

Newscast male voice

UHD scenarios

Supports Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz/24 kHz/48 kHz

Yes

No

Premium Edition

Zhifei

zhifei

Passionate narrator voice

UHD scenarios

Supports Chinese and Chinese-English mixed scenarios

8K/16K

Yes

No

Premium Edition

Aifei

aifei

Passionate narrator voice

Livestreaming scenarios

Supports Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

Yes

Standard Edition

Subgroup

yaqun

Store broadcast voice

Livestreaming scenarios

Supports Chinese and Chinese-English mixed scenarios

8K/16K

Yes

Yes

Standard Edition

Qiaowei

qiaowei

Store broadcast voice

Livestreaming scenarios

Supports Chinese and Chinese-English mixed scenarios

8K/16K

Yes

Yes

Standard Edition

Dahu

dahu

Northeastern Mandarin male voice

Dialect scenarios

Supports Chinese and Chinese-English mixed scenarios

8 K / 16 K

Yes

Yes

Standard Edition

ava

ava

American English female voice

English scenarios

Supports only English scenarios

8 K/16 K

Yes

No

Standard Edition

Zhilun

zhilun

Suspense narrator voice

UHD scenarios

Supports Chinese and Chinese-English mixed scenarios

8K or 16K

Yes

No

Premium Edition

Ailun

ailun

Suspense narrator voice

Livestreaming scenarios

Supports Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz

Yes

Yes

Standard Edition

Jielidou

jielidou

Healing a Child's Voice

Child voice scenarios

Supports only Chinese scenarios

8K / 16K

Yes

Yes

Standard Edition

Zhiwei

zhiwei

Childlike Female Voice

UHD scenarios

Supports only Chinese scenarios

8 kHz/16 kHz/24 kHz/48 kHz

Yes

No

Premium Edition

Laotie

laotie

Northeastern Buddy

Livestreaming scenarios

Supports only Chinese scenarios

8K/16K

Yes

Yes

Standard Edition

Younger sister

laomei

Hawking female voice

Livestreaming scenarios

Supports only Chinese scenarios

8 K/16 K

Yes

Yes

Standard Edition

Aikan

aikan

Tianjin-accented male voice

Dialect scenarios

Supports only Chinese scenarios

8 K/16 K

Yes

Yes

Standard Edition

Tala

tala

Filipino female voice

Multi-language scenarios

Supports only Filipino scenarios

8K or 16K

No

No

Standard Edition

Zhitian

zhitian

Sweet female voice

General

Supports Chinese and Chinese-English mixed scenarios

8 K/16 K

Yes

No

Premium Edition

Zhiqing

zhiqing

Taiwanese Mandarin female voice

Dialect scenarios

Supports only Chinese scenarios

8K / 16K

Yes

No

Premium Edition

Tien

tien

Vietnamese female voice

Multi-language scenarios

Supports only Vietnamese scenarios

8 K/16 K

No

No

Standard Edition

Becca

becca

American English customer service female voice

American English

Supports only English scenarios

8K or 16K

No

No

Standard Edition

Kyong

Kyong

Korean female voice

Korean scenarios

Korean

8 kHz/16 kHz

No

No

Standard Edition

masha

masha

Russian female voice

Russian scenarios

Russian

8K/16K

No

No

Standard Edition

camila

camila

Spanish female voice

Spanish scenarios

Spanish

8 kHz/16 kHz

No

No

Standard Edition

perla

perla

Italian female voice

Italian scenarios

Italian

8 kHz/16 kHz

No

No

Standard Edition

Zhimao

zhimao

Mandarin female voice

Livestreaming

Chinese

8 kHz/16 kHz

Yes

No

Standard Edition

Zhiyuan

zhiyuan

Mandarin female voice

General

Chinese

8 kHz/16 kHz

Yes

No

Standard Edition

Zhiya

zhiya

Mandarin female voice

Customer service

Chinese

8 kHz/16 kHz

Yes

No

Standard Edition

Zhiyue

zhiyue

Mandarin female voice

General

Chinese

8 kHz/16 kHz

Yes

No

Standard Edition

Zhida

zhida

Mandarin male voice

General

Chinese and Chinese-English mixed scenarios

8 kHz/16 kHz

Yes

No

Standard Edition

Zhisha

zhistella

Mandarin female voice

General

Chinese

8 kHz/16 kHz

Yes

No

Standard Edition

Kelly

kelly

Hong Kong Cantonese female voice

Dialect scenarios

Hong Kong Cantonese

8 kHz/16 kHz

Yes

No

Standard Edition

clara

clara

French female voice

General

French

8 kHz/16 kHz

No

No

Standard Edition

hanna

hanna

German female voice

General

German

8 kHz/16 kHz

No

No

Standard Edition

waan

waan

Thai female voice

General

Thai

8 kHz/16 kHz

No

No

Standard Edition

betty

betty

American English female voice

General

American English

8 kHz/16 kHz

Yes

No

Standard Edition

beth

beth

American English female voice

General

American English

8 kHz/16 kHz

Yes

No

Standard Edition

cindy

cindy

American English female voice

General

American English

8 kHz/16 kHz

Yes

No

Standard Edition

donna

donna

American English female voice

General

American English

8 kHz/16 kHz

Yes

No

Standard Edition

eva

eva

American English female voice

General

American English

8 kHz/16 kHz

Yes

No

Standard Edition

brian

brian

American English male voice

General

American English

8 kHz/16 kHz

Yes

No

Standard Edition

david

david

American English male voice

General

American English

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

abby_ecmix

abby_ecmix

American English female voice

General

English and English-Chinese mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

annie_ecmix

annie_ecmix

American English female voice

General

English and English-Chinese mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

andy_ecmix

andy_ecmix

American English male voice

General

English and English-Chinese mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

ava_ecmix

ava_ecmix

American English female voice

General

English and English-Chinese mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

betty_ecmix

betty_ecmix

American English female voice

General

English and English-Chinese mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

beth_ecmix

beth_ecmix

American English female voice

General

English and English-Chinese mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

brian_ecmix

brian_ecmix

American English male voice

General

English and English-Chinese mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

cindy_ecmix

cindy_ecmix

American English female voice

General

English and English-Chinese mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

cally_ecmix

cally_ecmix

American English female voice

General

English and English-Chinese mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

donna_ecmix

donna_ecmix

American English female voice

General

English and English-Chinese mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

david_ecmix

david_ecmix

American English male voice

General

English and English-Chinese mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

eva_ecmix

eva_ecmix

American English female voice

General

English and English-Chinese mixed scenarios

8 kHz/16 kHz/24 kHz

Yes

No

Standard Edition

Multi-emotion voice support

Only multi-emotion voice models support emotion selection. The supported emotions are listed in the following table. The supported emotion categories vary by voice and include neutral, happy, angry, sad, fear, hate, surprise, arousal, serious, disgust, jealousy, embarrassed, frustrated, affectionate, gentle, newscast, customer-service, story, and lively.

Voice name

voice parameter value

Emotion categorization

Zhifeng_Multi-emotion

zhifeng_emo

angry, fear, happy, neutral, sad, surprise

Zhibing_Multi-emotion

zhibing_emo

angry, fear, happy, neutral, sad, surprise

Zhimiao_Multi-emotion

zhimiao_emo

serious, sad, disgust, jealousy, embarrassed, happy, fear, surprise, neutral, frustrated, affectionate, gentle, angry, newscast, customer-service, story, living

Zhimi_Multi-emotion

zhimi_emo

angry, fear, happy, hate, neutral, sad, surprise

Zhiyan_Multi-emotion

zhiyan_emo

neutral, happy, angry, sad, fear, hate, surprise, arousal

Zhibei_Multi-emotion

zhibei_emo

neutral, happy, angry, sad, fear, hate, surprise

Zhitian_Multi-emotion

zhitian_emo

neutral, happy, angry, sad, fear, hate, surprise

Call instructions

  • The input text must be UTF-8 encoded.

  • Long text-to-speech and standard speech synthesis share many similarities.

Service endpoints

Access type

Description

URL

Host

Public network access

All servers can use the public endpoint URL. The SDK is configured with the public endpoint URL by default.

https://nls-gateway-cn-shanghai.aliyuncs.com/rest/v1/tts/async

nls-gateway-cn-shanghai.aliyuncs.com

Internal access from Alibaba Cloud ECS in Shanghai

If you use an Alibaba Cloud ECS instance in the China (Shanghai) region, you can use the internal endpoint URL. ECS instances in the classic network cannot access AnyTunnel and therefore cannot access Voice Service over the internal network. To use AnyTunnel, create a VPC and access the service from within the VPC.

Note

  • Using the internal endpoint does not incur public data transfer costs for your ECS instance.

  • For more information about ECS network types, see Network types.

http://nls-gateway-cn-shanghai-internal.aliyuncs.com/rest/v1/tts/async

nls-gateway-cn-shanghai-internal.aliyuncs.com

Interaction flow

The client sends an HTTPS POST request containing text to the server, and the server returns a response. The client can then handle the response in two ways:

  • Poll the synthesis status until the task is complete.

  • Wait for the server to complete the synthesis and send a callback to the webhook address that you configured. Your client program can then perform subsequent actions.

Note
  • Unlike the RESTful API for standard speech synthesis, the RESTful API for long text-to-speech does not directly return the synthesized audio data to the client. Instead, it returns an HTTP URL that you can use to download or play the audio file.

  • The server response header also includes the `task_id` parameter, which is the unique ID for the request.

Request parameters

The request parameters for asynchronous long text-to-speech are described in the following table.

When you send an HTTP/HTTPS request, include these parameters in the request body.

Name

Type

Required

Description

appkey

String

Yes

The AppKey of your application. For more information, see Manage projects.

token

String

No

The authentication token for the service.

text

String

Yes

The text to synthesize. The text must be UTF-8 encoded.

Note

To use the multi-emotion feature of a voice, add the SSML emotion tag to the text. For more information, see <emotion>.

Only voices that support multiple emotions can use the <emotion> tag. Otherwise, an `Illegal ssml text` error is reported.

format

String

Yes

The audio encoding format. Valid values: pcm, wav, and mp3. Default value: pcm.

sample_rate

Integer

Yes

The audio sample rate. Valid values: 16000 Hz and 8000 Hz. Default value: 16000 Hz.

voice

String

No

The voice. Default value: xiaoyun. For more voices, see Voice list.

volume

Integer

No

The volume. The value must be in the range of 0 to 100. Default value: 50.

speech_rate

Integer

No

The speech rate. The value must be in the range of -500 to 500. Default value: 0.

pitch_rate

Integer

No

The pitch. The value must be in the range of -500 to 500. Default value: 0.

enable_subtitle

Boolean

No

Specifies whether to enable the sentence-level timestamp feature. Default value: false.

enable_notify

Boolean

Yes

Specifies whether to enable the callback feature. Default value: false.

notify_url

String

No

The webhook address for callbacks. This parameter is required if enable_notify is set to true.

The URL must use the HTTP or HTTPS protocol. The host cannot be an IP address.

Service status codes

Each service response includes a `status` field, which is the service status code. The following sections describe the meanings of different status codes.

General-purpose error codes

Status code

Status message

Cause

Solution

40000000

The default client error code. This code corresponds to multiple error messages.

Invalid parameters or call logic was used.

Compare your code with the sample code in the official documentation to test and verify it.

40000001

The token 'xxx' has expired.

The token 'xxx' is invalid

Invalid parameters or call logic was used. This is a general-purpose client error code that usually indicates an incorrect token, such as an expired or invalid token.

Compare your code with the sample code in the official documentation to test and verify it.

40000002

Gateway:MESSAGE_INVALID:Can't process message in state'FAILED'!

The message is invalid or incorrect.

Compare your code with the sample code in the official documentation to test and verify it.

40000003

PARAMETER_INVALID

Failed to decode url params

The parameters passed by the user are incorrect. This error is common for RESTful API calls.

Compare your code with the sample code in the official documentation to test and verify it.

40000005

Gateway:TOO_MANY_REQUESTS:Too many requests!

Too many concurrent requests.

If you are using the Free Edition, you can upgrade to a commercial version to increase the concurrency.

If you are already using a commercial version, you can purchase a concurrency resource plan to increase your concurrency quota.

40000009

Invalid wav header!

The message header is invalid.

If you send a WAV audio file and set the format parameter to wav, check whether the WAV header of the audio file is correct. If the header is incorrect, the server may reject the request.

40000009

Too large wav header!

The WAV header of the transmitted audio is invalid.

You can send the audio stream in a format such as PCM or OPUS. If you use the WAV format, make sure that the WAV header of the audio file contains the correct data length.

40000010

Gateway:FREE_TRIAL_EXPIRED:The free trial has expired!

The trial period has ended, and the commercial version is not activated or your account has an overdue payment.

You can log on to the console to check the service activation status and your account balance.

40010001

Gateway:NAMESPACE_NOT_FOUND:RESTful url path illegal

The operation or parameter is not supported.

Check whether the parameters passed in the call are consistent with the requirements in the official documentation. You can compare them with the error message to identify and set the correct parameters.

For example, if you are using a curl command to make a RESTful API request, check whether the URL you constructed is valid.

40010003

Gateway:DIRECTIVE_INVALID:[xxx]

A general-purpose client-side error code.

This error indicates that the client passed an incorrect parameter or instruction. Detailed error messages are available for different operations. You can refer to the corresponding documentation to set the parameters correctly.

40010004

Gateway:CLIENT_DISCONNECT:Client disconnected before task finished!

The client actively terminated the connection before the request was processed.

None. Alternatively, you can close the connection after the server responds.

40010005

Gateway:TASK_STATE_ERROR:Got stop directive while task is stopping!

The client sent a message instruction that is not currently supported.

Compare your code with the sample code in the official documentation to test and verify it.

40020105

Meta:APPKEY_NOT_EXIST:Appkey not exist!

A non-existent Appkey was used.

Confirm whether a non-existent Appkey was used. You can log on to the console and view the project configuration to find the Appkey.

40020106

Meta:APPKEY_UID_MISMATCH:Appkey and user mismatch!

The Appkey and token passed in the call were not created by the same Alibaba Cloud account UID. This causes a mismatch.

Check whether you are using resources from two different accounts. Do not use an Appkey from Account A with a token generated from Account B.

403

Forbidden

The token is invalid. For example, the token does not exist or has expired.

Set a valid token. Tokens have an expiration period. You must obtain a new token before the current one expires.

41000003

MetaInfo doesn't have end point info

Failed to retrieve the routing information for this Appkey.

Check whether you are using resources from two different accounts. Do not use an Appkey from Account A with a token generated from Account B.

41010101

UNSUPPORTED_SAMPLE_RATE

The sample rate is not supported.

Real-time speech recognition currently supports only audio with a sample rate of 8000 Hz or 16000 Hz.

41040201

Realtime:GET_CLIENT_DATA_TIMEOUT:Client data does not send continuously!

Failed to retrieve data from the client due to a timeout.

When you call real-time speech recognition, the client must send data at a real-time rate and close the connection promptly after the data is sent.

50000000

GRPC_ERROR:Grpc error!

An exception caused by factors such as machine load or network issues. This error usually occurs randomly.

You can retry the call to resolve the issue.

50000001

GRPC_ERROR:Grpc error!

An exception caused by factors such as machine load or network issues. This error usually occurs randomly.

You can retry the call to resolve the issue.

52010001

GRPC_ERROR:Grpc error!

An exception caused by factors such as machine load or network issues. This error usually occurs randomly.

You can retry the call to resolve the issue.

Speech synthesis/Long text-to-speech error codes

Status code

Status message

Cause

Solution

21050000

SUCCESS

Success.

None.

21050001

RUNNING

The asynchronous long text-to-speech task is running.

You can send a GET request later to query the detection result.

21050002

QUEUEING

The asynchronous long text-to-speech task is in the queue.

Send a GET request later to query the detection result.

40000001

Gateway:ACCESS_DENIED:No privilege to this voice!

An invalid voice name was set.

Set a valid voice name based on the official documentation.

40000004

Gateway:IDLE_TIMEOUT:Websocket session is idle for too long time,the last directive is 'StartSynthesis'!

After a connection is established, no data is sent for an extended period. The server returns this error message after 10 seconds.

Close the connection promptly after the request is processed. This error may also occur when the server is under high momentary pressure and cannot return data in time. In this case, you can retry the call.

40010003

Gateway:DIRECTIVE_INVALID:No text specified!

No valid text was set for synthesis.

Set the text to be synthesized based on the sample code in the official documentation.

41020001

Speech synthesis call client-side error

Multiple error messages may be displayed. Resolve each error based on the corresponding message.

  • If the error message is Engine return error code: 424., it indicates that the background music or concatenated recording is in an invalid format. Set the background music to a supported format as described in the documentation.

  • If the error message is Engine return error code:418, it indicates that an unsupported voice name was passed.

  • If the error message is Engine return error code: 413, it indicates that the SSML format is invalid.

  • If the error message is Request json illegal,failed to parse request., it indicates that the JSON format is invalid.

  • If the error message is SSML text length should be less than 300., it indicates that the synthesis text is too long. We recommend that you use the long text-to-speech API.

51020001

TTS:TtsServerError

An exception occurred due to factors such as machine load or network issues. This error is usually intermittent.

Retry the call.