API reference

Updated at:

The speech synthesis service converts input text to binary audio data over a WebSocket connection.

← Back to Speech Synthesis

Billing and concurrency limits

Features

  • Supports PCM, WAV, and MP3 output formats.

  • Supports configuring the speaking rate, pitch, and volume.

  • Supports voices for different scenarios and styles. See Voices.

  • Supports synthesizing up to 300 characters per request. A Chinese character, English letter, punctuation mark, or interword space counts as one character. Content beyond 300 characters is truncated.

  • Input text must be UTF-8 encoded.

  • Supports multi-emotion voices. For usage information, see Introduction to SSML for the <emotion> tag. Tags are not counted as characters.If a rare character is mispronounced, use the SSML <phoneme> tag to specify its pinyin, or replace the character with a homophone. For more information, see Introduction to SSML.

Note
  • Timestamp support: The service can return the position of each Chinese character or English word in the audio. You can use the timestamps for digital-human lip synchronization and video subtitles. For more information, see Introduction to the speech synthesis timestamp feature.

  • Replace the original audio track in a video or audio file with a preset TTS voice:

    1. Call the TTS API to generate an audio file and use the voice parameter to specify the voice. To synchronize the audio by timestamp, set enable_subtitle=true and select a voice that supports timestamps, such as siqi, siyue, or sicheng. The default voice xiaoyun does not support timestamps. Timestamps are returned only over WebSocket.

    2. Use ffmpeg to merge the audio with the video:ffmpeg -i input.mp4 -i tts_output.wav -c:v copy -c:a aac -map 0:v:0 -map 1:a:0 output.mp4.

  • For voices intended for literature scenarios, see Long-text speech synthesis API reference.

  • To use the Android or iOS SDK, see Mobile SDK API reference.

Intelligent access to the nearest region

Speech synthesis supports intelligent access to the nearest region through the domain name nls-gateway.aliyuncs.com.

For end users, use intelligent access to the nearest region. The domain name resolves to the nearest regional server based on the client's location. For example, a request from Beijing resolves to the same server as the endpoint nls-gateway-cn-beijing.aliyuncs.com.

Endpoint

Access type

Description

URL

Internet access (China (Shanghai) by default)

All servers can access the service over the Internet. The SDK uses the Internet endpoint by default.

  • China (Shanghai): wss://nls-gateway-cn-shanghai.aliyuncs.com/ws/v1

  • China (Beijing): wss://nls-gateway-cn-beijing.aliyuncs.com/ws/v1

  • China (Shenzhen): wss://nls-gateway-cn-shenzhen.aliyuncs.com/ws/v1

ECS internal-network access

ECS instances in the China (Shanghai), China (Beijing), and China (Shenzhen) regions can use the internal endpoints. ECS instances in the classic network cannot access AnyTunnel. To access AnyTunnel over an internal network, use a virtual private cloud (VPC).

Note

  • Internal-network access does not incur Internet data transfer fees for ECS instances.

  • For information about ECS network types, see What is VPC.

  • China (Shanghai): ws://nls-gateway-cn-shanghai-internal.aliyuncs.com:80/ws/v1

  • China (Beijing): ws://nls-gateway-cn-beijing-internal.aliyuncs.com:80/ws/v1

  • China (Shenzhen): ws://nls-gateway-cn-shenzhen-internal.aliyuncs.com:80/ws/v1

Interaction flow

image
Note
  • The preceding figure shows the WebSocket interaction flow. For the RESTful API interaction flow, see RESTful API.

  • In addition to the audio stream, each server response contains the task_id parameter in the header. The parameter uniquely identifies the request.

  • To play the audio stream in real time, use a player that supports streaming playback, such as ffmpeg, pyaudio for Python, AudioFormat for Java, or MediaSource for JavaScript.

  1. Authenticate

    The client uses a token for authentication when it establishes a WebSocket connection to the server. For information about how to obtain a token, see Obtain a token.

  2. Start synthesis

    The client sends a speech synthesis request and configures the request parameters. The following table describes the parameters.

    Parameter

    Type

    Required

    Description

    appkey

    String

    Yes

    consoleThe appkey of the project created in the console.

    text

    String

    Yes

    The text to synthesize. The text must be UTF-8 encoded and cannot exceed 300 characters. Separate English words with spaces.

    Note

    To synthesize speech with an emotion-supported voice, add the SSML emotion tag to the text parameter. For more information, see Introduction to SSML.

    Only voices that support multiple emotions can use the <emotion> tag. Otherwise, the request returns the error Illegal ssml text.

    voice

    String

    No

    The voice. Default: xiaoyun.

    format

    String

    No

    The audio format. Valid values: .pcm, .wav, and .mp3. Default: pcm.

    sample_rate

    Integer

    No

    The audio sample rate. Default: 16000 Hz.

    volume

    Integer

    No

    The volume. Valid values: 0 to 100. Default: 50.

    speech_rate

    Integer

    No

    The speaking rate. Valid values: -500 to 500. Default: 0.

    The values [-500, 0, 500] correspond to the speed multipliers [0.5, 1.0, 2.0].

    1. -500 specifies 0.5 times the default speaking rate.

    2. 0 specifies the default speaking rate. The default rate varies by voice and is approximately four Chinese characters per second.

    3. 500 specifies twice the default speaking rate.

    Use the following formulas:

    1. 0.8× speed: (1-1/0.8)/0.002 = -125

    2. 1.2× speed: (1-1/1.2)/0.001 = 166

    Note
    1. For a speed multiplier below 1, use the coefficient 0.002.

    2. For a speed multiplier above 1, use the coefficient 0.001.

    The algorithm rounds the result to an approximate value.

    pitch_rate

    Integer

    No

    The pitch. Valid values: -500 to 500. Default: 0.

    enable_subtitle

    Boolean

    No

    Specifies whether to enable character-level timestamps. For more information, see Introduction to the speech synthesis timestamp feature.

  3. Receive synthesized audio

    The server returns synthesized audio as a binary stream. Use the SDK to receive and process it.

  4. Complete synthesis

    After synthesis is complete, the server sends a SynthesisCompleted event, as shown in the following example.

    {
        "header": {
            "message_id": "05450bf69c53413f8d88aed1ee60****",
            "task_id": "640bc797bb684bd6960185651307****",
            "namespace": "SpeechSynthesizer",
            "name": "SynthesisCompleted",
            "status": 20000000,
            "status_message": "GATEWAY|SUCCESS|Success."
        }
    }
    Note

    The sample saves the synthesized audio to a file. For lower playback latency, use streaming playback to play the audio as it arrives.

  5. Handle synthesis failures

    If synthesis fails because of invalid parameters or another error, the client receives a TaskFailed event, as shown in the following example. The underlying connection is closed after the event is received.

    {
       "header":{
          "namespace":"Default",
          "name":"TaskFailed",
          "status":41020001,
          "message_id":"62c126f7d9b340deb82b5b7eaca0****",
          "task_id":"4552df26d1f547aab9a2c4a94678****",
          "status_text":"TTS:TtsClientError:[tts]Engine return error code: 418"
       }
    }

Voices

Name

Voice parameter

Type

Scenario

Supported language

Sample rate (Hz)

Timestamp support

Erhua support

Voice quality

Abin

abin

Guangdong-accented Mandarin

Conversational digital human

Chinese or bilingual Chinese and English

8K/16K/24K/48K

No

No

Standard

Zhixiaobai

zhixiaobai

Mandarin female

Conversational digital human

Chinese or bilingual Chinese and English

8K/16K/24K/48K

No

Yes

Standard

Zhixiaoxia

zhixiaoxia

Mandarin female

Conversational digital human

Chinese or bilingual Chinese and English

8K/16K/24K/48K

No

Yes

Standard

Zhixiaomei

zhixiaomei

Mandarin female

Live-streaming digital human

Chinese or bilingual Chinese and English

8K/16K/24K

Yes

Yes

Standard

Zhigui

zhigui

Mandarin female

Live-streaming digital human

Chinese or bilingual Chinese and English

8K/16K

Yes

Yes

Standard

Zhishuo

zhishuo

Mandarin male

Customer-service digital human

Chinese or bilingual Chinese and English

8K/16K

Yes

Yes

Standard

Aixia

aixia

Amiable female

Customer service

Chinese or bilingual (Chinese and English)

8K/16K

Yes

Yes

Standard

Cally

cally

American English female

English conversational digital human

English only

8K/16K

Yes

Yes

Standard

Zhifeng Emo

zhifeng_emo

Multi-emotion male

General

Chinese or bilingual Chinese and English

8K/16K/24K

Yes

Yes

Standard

Zhibing Emo

zhibing_emo

Multi-emotion male

General

General scenarios

Chinese only

Chinese-only scenarios

8K/16K/24K

Yes

Yes

Yes

Standard

Zhimiao Emo

zhimiao_emo

Multi-emotion female

Chinese and English

Chinese and English

8K/16K

Yes

Yes

Standard

Zhimi Emo

zhimi_emo

Multi-emotion female

General

Chinese or bilingual Chinese and English

8K/16K

Yes

No

Standard

Zhiyan Emo

zhiyan_emo

Multi-emotion female

General

Chinese or bilingual Chinese and English

8K/16K

Yes

No

Standard

Zhibei Emo

zhibei_emo

Multi-emotion child

General

Chinese or bilingual Chinese and English

8K/16K

Yes

No

Standard

Zhitian Emo

zhitian_emo

Multi-emotion female

General

Chinese or bilingual Chinese and English

8K/16K

Yes

No

Standard

Xiaoyun

xiaoyun

Standard female

All scenarios

Chinese or bilingual (Chinese and English)

8K/16K

No

No

Lite

Xiaogang

xiaogang

Standard male

All scenarios

Chinese or bilingual (Chinese and English)

8K/16K

No

No

Lite

Ruoxi

ruoxi

Gentle female

All scenarios

Chinese or bilingual (Chinese and English)

8K/16K/24K

No

No

Standard

Siqi

siqi

Gentle female

All scenarios

Chinese or bilingual (Chinese and English)

8K/16K/24K

Yes

No

Standard

Sijia

sijia

Standard female

All scenarios

Chinese or bilingual (Chinese and English)

8K/16K/24K

No

No

Standard

Sicheng

sicheng

Standard male

All scenarios

Chinese or bilingual (Chinese and English)

8K/16K/24K

Yes

No

Standard

Aiqi

aiqi

Gentle female

All scenarios

Chinese or bilingual (Chinese and English)

8K/16K

Yes

No

Standard

Aijia

aijia

Standard female

All scenarios

Chinese or bilingual (Chinese and English)

8K/16K

Yes

No

Standard

Aicheng

aicheng

Standard male

All scenarios

Chinese or bilingual (Chinese and English)

8K/16K

Yes

No

Standard

Aida

aida

Standard male

All scenarios

Chinese or bilingual (Chinese and English)

8K/16K

Yes

No

Standard

Ning'er

ninger

Standard female

All scenarios

Chinese only

8K/16K/24K

No

No

Standard

Ruilin

ruilin

Standard female

All scenarios

Chinese only

8K/16K/24K

No

No

Standard

Siyue

siyue

Gentle female

Customer service

Chinese or bilingual (Chinese and English)

8K/16K/24K

Yes

No

Standard

Aiya

aiya

Harsh female

Customer service

Chinese or bilingual (Chinese and English)

8K/16K

Yes

No

Standard

Aimei

aimei

Sweet female

Customer service

Chinese or bilingual (Chinese and English)

8K/16K

Yes

No

Standard

Aiyu

aiyu

Natural female

Customer service

Chinese or bilingual (Chinese and English)

8K/16K

Yes

No

Standard

Aiyue

aiyue

Gentle female

Customer service

Chinese or bilingual (Chinese and English)

8K/16K

Yes

No

Standard

Aijing

aijing

Harsh female

Customer service

Chinese or bilingual (Chinese and English)

8K/16K

Yes

No

Standard

Xiaomei

xiaomei

Sweet female

Customer service

Chinese or bilingual (Chinese and English)

8K/16K/24K

No

No

Standard

Aina

aina

Female, Zhejiang accent

Customer service

Chinese only

8K/16K

Yes

No

Standard

Yina

yina

Female, Zhejiang accent

Customer service

Chinese only

8K/16K/24K

No

No

Standard

Sijing

sijing

Harsh female

Customer service

Chinese only

8K/16K/24K

Yes

No

Standard

Sitong

sitong

Child voice

Child voice scenarios

Chinese only

8K/16K/24K

No

No

Standard

Xiaobei

xiaobei

Little girl voice

Child voice scenarios

Chinese only

8K/16K/24K

Yes

No

Standard

Aitong

aitong

Child voice

Child voice scenarios

Chinese only

8K/16K

Yes

No

Standard

Aiwei

aiwei

Little girl voice

Child voice scenarios

Chinese only

8K/16K

Yes

No

Standard

Aibao

aibao

Little girl voice

Child voice scenarios

Chinese only

8K/16K

Yes

No

Standard

Harry

harry

Male, British accent

English only

English only

8K/16K

No

No

Standard

Abby

abby

Female, American accent

English only

English only

8K/16K

Yes

No

Standard

Andy

andy

Male, American accent

English only

English only

8K/16K

Yes

No

Standard

Eric

eric

Male, British accent

English only

English only

8K/16K

Yes

No

Standard

Emily

emily

Female, British accent

English only

English only

8K/16K

Yes

No

Standard

Luna

luna

Female, British accent

English only

English only

8K/16K

Yes

No

Standard

Luca

luca

Male, British accent

English only

English only

8K/16K

Yes

No

Standard

Wendy

wendy

Female, British accent

English only

English only

8K/16K/24K

No

No

Standard

William

william

Male, British accent

English only

English only

8K/16K/24K

No

No

Standard

Olivia

olivia

Female, British accent

English only

English only

8K/16K/24K

No

No

Standard

Shanshan

shanshan

Female, Cantonese

Dialect scenarios

Cantonese (simplified) and bilingual (Cantonese and English)

8K/16K/24K

No

No

Standard

Aiyuan

aiyuan

Caring female

Literature

Chinese or bilingual Chinese and English

8K/16K

Yes

Yes

Premium

Aiying

aiying

Cute child

Literature

Chinese or bilingual Chinese and English

8K/16K

Yes

Yes

Premium

Aixiang

aixiang

Resonant male

Literature

Chinese or bilingual Chinese and English

8K/16K

Yes

Yes

Premium

Aimo

aimo

Emotional male

Literature

Chinese or bilingual Chinese and English

8K/16K

Yes

Yes

Premium

Aiye

aiye

Young male

Literature

Chinese or bilingual Chinese and English

8K/16K

Yes

Yes

Premium

Aiting

aiting

Radio female

Literature

Chinese or bilingual Chinese and English

8K/16K

Yes

Yes

Premium

Aifan

aifan

Emotional female

Literature

Chinese or bilingual Chinese and English

8K/16K

Yes

Yes

Premium

Lydia

lydia

Female, bilingual (Chinese and English)

English only

English only

8K/16K

Yes

No

Standard

Chuangirl

chuangirl

Sichuan-accented female

Dialect

Chinese or bilingual Chinese and English

8K/16K

No

No

Standard

Aishuo

aishuo

Natural male

Customer service

Chinese or bilingual (Chinese and English)

8K/16K

Yes

No

Standard

Qingqing

qingqing

Female, Taiwanese

Dialect scenarios

Chinese only

8K/16K

No

No

Standard

Cuijie

cuijie

Female, Northeastern Mandarin

Dialect scenarios

Chinese only

8K/16K

Yes

Yes

Standard

Xiaoze

xiaoze

Male, Hunan accent

Dialect scenarios

Chinese only

8K/16K

No

No

Standard

Ainan

ainan

Advertising male

Literature

Chinese or bilingual Chinese and English

8K/16K

Yes

Yes

Premium

Aihao

aihao

Information male

Literature

Chinese or bilingual Chinese and English

8K/16K

Yes

Yes

Premium

Aiming

aiming

Humorous male

Literature

Chinese or bilingual Chinese and English

8K/16K

Yes

Yes

Premium

Aixiao

aixiao

Information female

Literature

Chinese or bilingual Chinese and English

8K/16K

Yes

Yes

Premium

Aichu

aichu

Food-documentary male

Literature

Chinese or bilingual Chinese and English

8K/16K

Yes

Yes

Premium

Aiqian

aiqian

Information female

Literature

Chinese or bilingual Chinese and English

8K/16K

Yes

Yes

Premium

Tomoka

tomoka

Japanese female

Multilingual

Japanese only

8K/16K

Yes

No

Standard

Tomoya

tomoya

Japanese male

Multilingual

Japanese only

8K/16K

Yes

No

Standard

Annie

annie

American English female

English

English only

8K/16K

Yes

No

Standard

Aishu

aishu

Information male

Literature

Chinese or bilingual Chinese and English

8K/16K

Yes

Yes

Premium

Airu

airu

News female

Literature

Chinese or bilingual Chinese and English

8K/16K

Yes

Yes

Premium

Jiajia

jiajia

Cantonese female

Dialect

Simplified Cantonese or bilingual Cantonese and English

8K/16K

Yes

No

Standard

Indah

indah

Indonesian female

Multilingual

Indonesian only

8K/16K

No

No

Standard

Taozi

taozi

Cantonese female

Dialect

Simplified Cantonese or bilingual Cantonese and English

8K/16K

Yes

No

Standard

Guijie

guijie

Friendly female

General

Chinese or bilingual Chinese and English

8K/16K

Yes

Yes

Standard

Stella

stella

Intellectual female

General

Chinese or bilingual Chinese and English

8K/16K

Yes

Yes

Standard

Stanley

stanley

Calm male

General

Chinese or bilingual Chinese and English

8K/16K

Yes

Yes

Standard

Kenny

kenny

Calm male

General

Chinese or bilingual Chinese and English

8K/16K

Yes

Yes

Standard

Rosa

rosa

Natural female

General

Chinese or bilingual Chinese and English

8K/16K

Yes

Yes

Standard

Farah

farah

Malay female

Multilingual

Malay only

8K/16K

No

No

Standard

Mashu

mashu

Children's drama male

General

General

8K/16K

Yes

No

Standard

Zhiqi

zhiqi

Gentle female

Ultra-HD

Chinese or bilingual Chinese and English

8K/16K/24K/48K

Yes

No

Premium

Zhichu

zhichu

Food-documentary male

Ultra-HD

Chinese or bilingual Chinese and English

8K/16K/24K/48K

Yes

Yes

Premium

Xiaoxian

xiaoxian

Friendly female

Live streaming

Chinese or bilingual Chinese and English

8K/16K

Yes

Yes

Standard

Yuer

yuer

Children's drama female

General

Chinese only

8K/16K

Yes

No

Standard

Maoxiaomei

maoxiaomei

Energetic female

Live streaming

Chinese or bilingual Chinese and English

8K/16K

Yes

Yes

Standard

Zhixiang

zhixiang

Resonant male

Ultra-HD

Chinese or bilingual Chinese and English

8K/16K/24K/48K

Yes

No

Premium

Zhijia

zhijia

Standard female

Ultra-HD

Chinese or bilingual Chinese and English

8K/16K/24K/48K

Yes

No

Premium

Zhinan

zhinan

Advertising male

Ultra-HD

Chinese or bilingual Chinese and English

8K/16K/24K/48K

Yes

No

Premium

Zhiqian

zhiqian

Information female

Ultra-HD

Chinese or bilingual Chinese and English

8K/16K/24K/48K

Yes

No

Premium

Zhiru

zhiru

News female

Ultra-HD

Chinese or bilingual Chinese and English

8K/16K/24K/48K

Yes

No

Premium

Zhide

zhide

News male

Ultra-HD

Chinese or bilingual Chinese and English

8K/16K/24K/48K

Yes

No

Premium

Zhifei

zhifei

Passionate narration

Ultra-HD

Chinese or bilingual Chinese and English

8K/16K

Yes

No

Premium

Aifei

aifei

Passionate narration

Live streaming

Chinese or bilingual Chinese and English

8K/16K

Yes

Yes

Standard

Yaqun

yaqun

Retail announcement

Live streaming

Chinese or bilingual Chinese and English

8K/16K

Yes

Yes

Standard

Qiaowei

qiaowei

Retail announcement

Live streaming

Chinese or bilingual Chinese and English

8K/16K

Yes

Yes

Standard

Dahu

dahu

Northeastern Mandarin male

Dialect

Chinese or bilingual Chinese and English

8K/16K

Yes

Yes

Standard

ava

ava

American English female

English

English only

8K/16K

Yes

No

Standard

Zhilun

zhilun

Suspense narration

Ultra-HD

Chinese or bilingual Chinese and English

8K/16K

Yes

No

Premium

Ailun

ailun

Suspense narration

Live streaming

Chinese or bilingual Chinese and English

8K/16K

Yes

Yes

Standard

Jielidou

jielidou

Soothing child

Child voice

Chinese only

8K/16K

Yes

Yes

Standard

Zhiwei

zhiwei

Young girl

Ultra-HD

Chinese only

8K/16K/24K/48K

Yes

No

Premium

Laotie

laotie

Northeastern Mandarin male

Live streaming

Chinese only

8K/16K

Yes

Yes

Standard

Laomei

laomei

Vendor-call female

Live streaming

Chinese only

8K/16K

Yes

Yes

Standard

Aikan

aikan

Tianjin-accented male

Dialect

Chinese only

8K/16K

Yes

Yes

Standard

Tala

tala

Filipino female

Multilingual

Filipino only

8K/16K

No

No

Standard

Zhitian

zhitian

Sweet female

General

Chinese or bilingual Chinese and English

8K/16K

Yes

No

Premium

Zhiqing

zhiqing

Taiwan-accented female

Dialect

Chinese only

8K/16K

Yes

No

Premium

Tien

tien

Vietnamese female

Multilingual

Vietnamese only

8K/16K

No

No

Standard

Becca

becca

American English customer-service female

American English

English only

8K/16K

No

No

Standard

Kyong

Kyong

Korean female

Korean

Korean

8K/16K

No

No

Standard

masha

masha

Russian female

Russian

Russian

8K/16K

No

No

Standard

camila

camila

Spanish female

Spanish

Spanish

8K/16K

No

No

Standard

perla

perla

Italian female

Italian

Italian

8K/16K

No

No

Standard

Zhimao

zhimao

Mandarin female

Live streaming

Chinese

8K/16K

Yes

No

Standard

Zhiyuan

zhiyuan

Mandarin female

General

Chinese

8K/16K

Yes

No

Standard

Zhiya

zhiya

Mandarin female

Customer service

Chinese

8K/16K

Yes

No

Standard

Zhiyue

zhiyue

Mandarin female

General

Chinese

8K/16K

Yes

No

Standard

Zhida

zhida

Mandarin male

General

Chinese or bilingual Chinese and English

8K/16K

Yes

No

Standard

Zhistella

zhistella

Mandarin female

General

Chinese

8K/16K

Yes

No

Standard

Kelly

kelly

Hong Kong Cantonese female

Dialect

Hong Kong Cantonese

8K/16K

Yes

No

Standard

clara

clara

French female

General

French

8K/16K

No

No

Standard

hanna

hanna

German female

General

German

8K/16K

No

No

Standard

waan

waan

Thai female

General

Thai

8K/16K

No

No

Standard

betty

betty

American English female

General

American English

8K/16K

Yes

No

Standard

beth

beth

American English female

General

American English

8K/16K

Yes

No

Standard

cindy

cindy

American English female

General

American English

8K/16K

Yes

No

Standard

donna

donna

American English female

General

American English

8K/16K

Yes

No

Standard

eva

eva

American English female

General

American English

8K/16K

Yes

No

Standard

brian

brian

American English male

General

American English

8K/16K

Yes

No

Standard

david

david

American English male

General

American English

8K/16K/24K

Yes

No

Standard

abby_ecmix

abby_ecmix

American English female

General

English or bilingual English and Chinese

8K/16K/24K

Yes

No

Standard

annie_ecmix

annie_ecmix

American English female

General

English or bilingual English and Chinese

8K/16K/24K

Yes

No

Standard

andy_ecmix

andy_ecmix

American English male

General

English or bilingual English and Chinese

8K/16K/24K

Yes

No

Standard

ava_ecmix

ava_ecmix

American English female

General

English or bilingual English and Chinese

8K/16K/24K

Yes

No

Standard

betty_ecmix

betty_ecmix

American English female

General

English or bilingual English and Chinese

8K/16K/24K

Yes

No

Standard

beth_ecmix

beth_ecmix

American English female

General

English or bilingual English and Chinese

8K/16K/24K

Yes

No

Standard

brian_ecmix

brian_ecmix

American English male

General

English or bilingual English and Chinese

8K/16K/24K

Yes

No

Standard

cindy_ecmix

cindy_ecmix

American English female

General

English or bilingual English and Chinese

8K/16K/24K

Yes

No

Standard

cally_ecmix

cally_ecmix

American English female

General

English or bilingual English and Chinese

8K/16K/24K

Yes

No

Standard

donna_ecmix

donna_ecmix

American English female

General

English or bilingual English and Chinese

8K/16K/24K

Yes

No

Standard

david_ecmix

david_ecmix

American English male

General

English or bilingual English and Chinese

8K/16K/24K

Yes

No

Standard

eva_ecmix

eva_ecmix

American English female

General

English or bilingual English and Chinese

8K/16K/24K

Yes

No

Standard

Multi-emotion voice support

Only multi-emotion voice models support emotion selection. Supported emotions vary by voice and can include neutral, happy, angry, sad, fear, hate, surprise, arousal, serious, disgust, jealousy, embarrassed, frustrated, affectionate, gentle, newscast, customer-service, story, and living.

Voice name

Voice parameter

Emotion category

Zhifeng (multi-emotion)

zhifeng_emo

angry, fear, happy, neutral, sad, surprise

Zhibing (multi-emotion)

zhibing_emo

angry, fear, happy, neutral, sad, surprise

Zhimiao (multi-emotion)

zhimiao_emo

serious, sad, disgust, jealousy, embarrassed, happy, fear, surprise, neutral, frustrated, affectionate, gentle, angry, newscast, customer-service, story, living

Zhimi (multi-emotion)

zhimi_emo

angry, fear, happy, hate, neutral, sad, surprise

Zhiyan (multi-emotion)

zhiyan_emo

neutral, happy, angry, sad, fear, hate, surprise, arousal

Zhibei (multi-emotion)

zhibei_emo

neutral, happy, angry, sad, fear, hate, surprise

Zhitian (multi-emotion)

zhitian_emo

neutral, happy, angry, sad, fear, hate, surprise

Status codes

Each response contains a status field that indicates the service status code. The following tables describe the status codes.

General-purpose error codes

Status code

Status message

Cause

Solution

40000000

The default client error code. This code corresponds to multiple error messages.

Invalid parameters or call logic was used.

Compare your code with the sample code in the official documentation to test and verify it.

40000001

The token 'xxx' has expired.

The token 'xxx' is invalid

Invalid parameters or call logic was used. This is a general-purpose client error code that usually indicates an incorrect token, such as an expired or invalid token.

Compare your code with the sample code in the official documentation to test and verify it.

40000002

Gateway:MESSAGE_INVALID:Can't process message in state'FAILED'!

The message is invalid or incorrect.

Compare your code with the sample code in the official documentation to test and verify it.

40000003

PARAMETER_INVALID

Failed to decode url params

The parameters passed by the user are incorrect. This error is common for RESTful API calls.

Compare your code with the sample code in the official documentation to test and verify it.

40000005

Gateway:TOO_MANY_REQUESTS:Too many requests!

Too many concurrent requests.

If you are using the Free Edition, you can upgrade to a commercial version to increase the concurrency.

If you are already using a commercial version, you can purchase a concurrency resource plan to increase your concurrency quota.

40000009

Invalid wav header!

The message header is invalid.

If you send a WAV audio file and set format to wav, make sure that the WAV header is valid. Otherwise, the server may reject the file.

40000009

Too large wav header!

The WAV header of the transmitted audio is invalid.

You can send the audio stream in a format such as PCM or OPUS. If you use the WAV format, make sure that the WAV header of the audio file contains the correct data length.

40000010

Gateway:FREE_TRIAL_EXPIRED:The free trial has expired!

The trial period has ended, and the commercial version is not activated or your account has an overdue payment.

You can log on to the console to check the service activation status and your account balance.

  1. On the Service Management and Activation page of the Intelligent Speech Interaction console, verify that the service used by the API or SDK is upgraded to the Commercial Edition.

  2. Speech synthesis and streaming text-to-speech are separate services and must be upgraded to the Commercial Edition separately. Upgrading one service does not upgrade the other.

  3. Make sure that the SDK matches the activated service. For example, before you use the streaming text-to-speech SDK, upgrade streaming text-to-speech to the Commercial Edition.

40010001

Gateway:NAMESPACE_NOT_FOUND:RESTful url path illegal

The operation or parameter is not supported.

Check whether the parameters passed in the call are consistent with the requirements in the official documentation. You can compare them with the error message to identify and set the correct parameters.

For example, if you are using a curl command to make a RESTful API request, check whether the URL you constructed is valid.

40010003

Gateway:DIRECTIVE_INVALID:[xxx]

A general-purpose client-side error code.

This error indicates that the client passed an incorrect parameter or instruction. Detailed error messages are available for different operations. You can refer to the corresponding documentation to set the parameters correctly.

40010004

Gateway:CLIENT_DISCONNECT:Client disconnected before task finished!

The client actively terminated the connection before the request was processed.

None. Alternatively, you can close the connection after the server responds.

40010005

Gateway:TASK_STATE_ERROR:Got stop directive while task is stopping!

The client sent a message instruction that is not currently supported.

Compare your implementation with the sample code in the documentation. If the message empty body data is returned, troubleshoot the issue as follows:

  1. Check how the audio is passed. If you use a URL, make sure that the URL is valid and the audio file has been generated. If you use a binary stream, compare the original audio data with the data sent by the client and forwarded by the server. The data must be non-empty and identical.

  2. Run the official sample after replacing AccessKeyId, AccessKeySecret, and Appkey with your credentials to rule out an implementation issue.

  3. Use fixed test parameters and a fixed audio source, such as a fixed OSS URL, to reproduce the issue and determine whether the cause is the network, data preparation, or the API call.

40020105

Meta:APPKEY_NOT_EXIST:Appkey not exist!

A non-existent Appkey was used.

Confirm whether a non-existent Appkey was used. You can log on to the console and view the project configuration to find the Appkey.

40020106

Meta:APPKEY_UID_MISMATCH:Appkey and user mismatch!

The Appkey and token passed in the call were not created by the same Alibaba Cloud account UID. This causes a mismatch.

Check whether you are using resources from two different accounts. Do not use an Appkey from Account A with a token generated from Account B.

403

Forbidden

The token is invalid. For example, the token does not exist or has expired.

Set a valid token. Tokens have an expiration period. You must obtain a new token before the current one expires.

41000003

MetaInfo doesn't have end point info

Failed to retrieve the routing information for this Appkey.

Check whether you are using resources from two different accounts. Do not use an Appkey from Account A with a token generated from Account B.

41010101

UNSUPPORTED_SAMPLE_RATE

The sample rate is not supported.

Real-time speech recognition currently supports only audio with a sample rate of 8000 Hz or 16000 Hz.

41040201

Realtime:GET_CLIENT_DATA_TIMEOUT:Client data does not send continuously!

Failed to retrieve data from the client due to a timeout.

When you call real-time speech recognition, the client must send data at a real-time rate and close the connection promptly after the data is sent.

50000000

GRPC_ERROR:Grpc error!

An exception caused by factors such as machine load or network issues. This error usually occurs randomly.

You can retry the call to resolve the issue.

50000001

GRPC_ERROR:Grpc error!

An exception caused by factors such as machine load or network issues. This error usually occurs randomly.

You can retry the call to resolve the issue.

52010001

GRPC_ERROR:Grpc error!

An exception caused by factors such as machine load or network issues. This error usually occurs randomly.

You can retry the call to resolve the issue.

Speech synthesis/Long-text speech synthesis error codes

Status code

Status message

Cause

Solution

40000001

Gateway:ACCESS_DENIED:No privilege to this voice!

An incorrect speaker name was set.

You can refer to the official documentation to set the correct speaker.

40000004

Gateway:IDLE_TIMEOUT:Websocket session is idle for too long time,the last directive is 'StartSynthesis'!

After a connection is established, the server returns this error message if no data is sent for more than 10 seconds.

Close the connection promptly after the request is processed. This error may also occur if the server is under high instantaneous pressure and cannot return data in time. In this case, you can retry the request to resolve the issue.

40010003

Gateway:DIRECTIVE_INVALID:No text specified!

No valid text for synthesis was set.

You can refer to the sample code in the official documentation to set the text for synthesis.

41020001

Speech synthesis client error

Multiple error messages may be returned. Adjust your code based on the specific error message.

  • The message Engine return error code: 424. indicates that the background music or concatenated recording has an invalid format. Configure the background audio as described in the documentation.

  • The message Engine return error code:418 indicates that the specified voice is unsupported. If this error is returned for a CosyVoice-V2 voice, use the WebSocket endpoint of the China (Beijing) region (wss://nls-gateway-cn-beijing.aliyuncs.com/ws/v1). CosyVoice large-model voices are available only in the China (Beijing) region.

  • A cloned VoiceName does not support the enable_subtitle timestamp feature. Only some official long-text voices, such as voices whose names start with long, support this feature.

  • The message Engine return error code: 413 indicates that SSML tags of the same type are nested. Tags of the same type, such as s, p, and break, cannot be nested. Tags of different types can be nested. Place tags of the same type at the same level. For example, do not nest a p tag inside an s tag within speak. Instead, place separate s tags for the text segments at the same level within speak.

  • The message Request json illegal,failed to parse request. indicates that the JSON payload is invalid.

  • The message SSML text length should be less than 300. indicates that the synthesis text is too long. Use long-text speech synthesis.

51020001

TTS:TtsServerError

An exception caused by factors such as machine load or network issues. This error usually occurs randomly.

You can retry the call to resolve the issue.