Connection guide

Updated at:

This document describes the WebSocket API for real-time multimodal interaction. The WebSocket protocol is the recommended connection method due to its low latency and minimal resource consumption.

WebSocket is a network protocol that supports full-duplex communication. A client and server establish a persistent connection through a single handshake. This allows both parties to actively push data to each other, offering significant advantages in real-time performance and efficiency.

Many WebSocket libraries and examples are available for common programming languages, such as:

  • Go: gorilla/websocket
  • PHP: Ratchet
  • Node.js: ws

Before you start development, you should understand the basic principles and technical details of WebSocket.

NoteFor RTOS or some older Linux systems, you must configure the TLS tunnel for secure WebSocket communication as follows:

  • The TLS version must be TLS 1.2 or later.
  • Enable SNI (Server Name Indication).
  • Configure the CA certificate (GlobalSign Root CA - R3). You can also download it from the official GlobalSign website.

Prerequisites

You must activate the service and obtain an API key. To prevent security risks from code leaks, we recommend that you configure the API key as an environment variable instead of hard coding it in your application.

NoteHandling an API key on the client poses a security risk. Instead, your server should use the API key to obtain a temporary authentication token and send it to the client. For more information, see Obtain a temporary authentication token.

Call sequence diagram

Key workflow: A Started message from the server confirms that the session is created, but the client must not send audio immediately. The client must wait for the DialogStateChanged event and can only start sending the audio stream when the state is Listening.

image

Endpoint

wss://dashscope.aliyuncs.com/api-ws/v1/inference

Authentication

To initiate a WebSocket handshake (HTTP Upgrade) request, include the API key in the HTTP header. Replace your_api_key with your actual API key:

"Authorization": "Bearer your_api_key"

Voice interaction

Enabling voice interaction adds speech recognition and speech synthesis capabilities to your multimodal interaction application.

Speech recognition supports the following models: Paraformer, FunASR, Qwen-Audio-3.0-ASR-Flash-Streaming, qwen3-asr-flash-realtime, and AppSpecificASR-Realtime.

Speech synthesis supports the following models: cosyvoice-v2, cosyvoice-v3-flash, cosyvoice-v3-plus, cosyvoice-v3.5-flash, cosyvoice-v3.5-plus, qwen3-tts, qwen3-tts-instruct, qwen3-tts-vd, qwen3-tts-vc, sambert, and AppSpecificTTS.

After you select a speech synthesis model in the console, click the voice list in the upper-right corner of the voice interaction area to view the supported voices.

You can also find lists of official voices in the documentation. For the voices of cosyvoice-v2, cosyvoice-v3-plus, and cosyvoice-v3-flash, see the CosyVoice voice list. For qwen3-tts voices, see Supported voices. For sambert voices, see the Java SDK. For the voice parameter, use the name from the SDK, but remove the sambert- prefix and v1 suffix.

Before using a custom voice, ensure its status is "OK". To check the status, see Query a specific voice.

Message types

Binary messages

Currently, binary messages contain only audio data.

Upload audio

To upload audio, you can convert the raw audio directly into a binary stream without extra processing.

The audio uploaded for speech recognition must be 16-bit, single-channel, signed, little-endian PCM. For the sample rate, refer to the description of the parameters.upstream.sample_rate parameter in the Start message.

To reduce network traffic and bandwidth consumption, you can encode the PCM audio into Opus format and set the upload audio format to raw-opus.

When you upload audio, the action to take depends on the upstream.mode setting in the Start message:

  • If mode is tap2talk or duplex, the client must continuously upload audio, and the server automatically detects voice activity. We recommend that you upload data every 100 ms. An interval that is too long or too short can negatively affect latency and processing efficiency.

    • Calculate the bytes per upload with the following formula: Bytes per upload = sample rate * bit depth / 8 * time interval (ms) / 1000
    • For example, with a 16 kHz sample rate, 16-bit bit depth, and a 100 ms interval, upload 16000 * (16/8) * 100 / 1000 = 3200 bytes of PCM data each time.
  • If the mode is push2talk, the client does not need to continuously upload audio. Instead, you must use SendSpeech and StopSpeech to notify the server of the start and end of audio recognition. You must upload the audio immediately after you send SendSpeech. Otherwise, the processing time increases.

Receive audio

The server sends the large model's response to the Text-to-Speech (TTS) service to generate audio, which is then sent to the client.

  • The received audio is 16-bit single-channel. The sample rate and encoding format are defined by the parameters in the Start message.
  • The delivery speed depends on the TTS service performance and is usually faster than the playback speed.
  • The server sends a RespondingStarted event before it sends the audio and a RespondingEnded event after.
  • After playback is complete, the client must send LocalRespondingEnded to signal that playback has finished.

Text messages

Text messages are JSON-formatted strings and are classified into two types based on their direction:

  • Input Message: A directive sent from the client to the server that specifies an action for the server to execute.
  • Output Message: An event sent from the server to the client that indicates the result or progress of a server action.

Text messages mark key points in the interaction flow. They control the process and transmit critical information. For more information about the interaction sequence of different messages, see the time series chart.

A text message consists of two parts: header and payload.

  • payload: The content varies based on the message type.

  • header: The content is fixed and includes the following parameters:

    Parameter

    Type

    Required

    Description

    task_id

    string

    Yes

    The client generates this unique identifier for the connection to track task execution in the engineering pipeline. We recommend a 36-character UUID string, for example: "f894c16f-f20e-4c1d-837e-89e0fbc63a43"

    streaming

    string

    Yes

    The input and output type. For multimodal interaction, this must be set to "duplex", which indicates streaming input and streaming output.

    action

    string

    Yes

    The type of input message for the model:

    • run-task: The first input message of a task.

    • finish-task: The last input message of a task.

    • continue-task: Other input messages in the task.

Connection keepalive policy

If 60 seconds pass without the server sending a message to the client, Model Studio terminates the connection and returns a ResponseTimeout error.

To maintain the connection during periods of inactivity, you must periodically send a HeartBeat message. The server responds to the heartbeat to keep the connection active and prevent a timeout.

The mobile and C++ software development kits (SDKs) have built-in keepalive logic. You do not need to send heartbeats manually.

Text message types

Start a session

Start - Input Message

The client sends this message to request the start of a session. After the server receives a Start message, it sends a Started message to the client.

Level 1 Parameter

Level 2 Parameter

Type

Required

Description

task_group

string

Yes

The name of the task group. This parameter is fixed to aigc.

task

string

Yes

The name of the task. This parameter is fixed to multimodal-generation.

function

string

Yes

The function to call. This parameter is fixed to generation.

model

string

Yes

The name of the Alibaba Cloud Model Studio model. This parameter is fixed to multimodal-dialog.

input

directive

string

Yes

The name of the instruction: Start

workspace_id

string

Yes

Your workspace ID from Alibaba Cloud Model Studio (Workspace ID). To find this ID, go to the multimodal interaction console, click the workspace name in the lower-left corner, and look in Workspace Details. Currently, only the default workspace of an Alibaba Cloud account is supported.

app_id

string

Yes

Your App ID (App ID). You can find this on the My Applications page of the multimodal interaction console.

dialog_id

string

No

The dialog ID. Omit this parameter to start a new session. The server generates a dialog ID and returns it in an event. The format is a 36-character string, for example: "12345678-1234-1234-1234-1234567890ab". To continue a previous conversation, pass the dialog_id from that session.

parameters

upstream

object

Yes

For details, see the parameters.upstream table below.

downstream

object

No

For details, see the parameters.downstream table below.

client_info

object

Yes

For details, see the parameters.client_info table below.

biz_params

object

No

For details, see the parameters.biz_params table below.

The following table describes the parameters for parameters.upstream.

Parameter

Type

Required

Description

type

string

Yes

The upstream type.

AudioOnly: A voice-only call.

mode

string

No

The mode the client uses. Default is tap2talk.

Options:

  • push2talk: client-controlled mode.

  • tap2talk: tap-to-talk mode.

  • duplex: duplex mode.

For a comparison of the three client modes, see the Comparison of the three client modes table below.

audio_format

string

No

The audio format. Supported formats are pcm and raw-opus. The default is pcm.

sample_rate

int

No

The sample rate for speech recognition. Supported values:

  • 8000

  • 16000

  • 24000

  • 48000

The default is 16000.

vocabulary_id

string

No

The custom vocabulary ID. If you set this parameter, it overrides the custom vocabulary configuration in the console. If the custom vocabularies in the console do not meet your needs, you can manage them programmatically by using the OpenAPI. For more information, see the custom vocabulary API documentation.

language

string

No

The speech recognition language. By default, it matches the language selected in the console. To specify multiple languages, separate them with commas, for example: zh,en.

The following table describes the parameters for parameters.downstream:

Parameter

Type

Required

Description

voice

string

No

The voice for speech synthesis. The supported voices depend on the speech synthesis model that you select in the console.

sample_rate

int

No

The sample rate for speech synthesis. Supported values:

  • 8000

  • 16000

  • 24000

  • 48000

The default is 24000.

The Qwen-TTS and Qwen3-TTS models support only 24000.

audio_format

string

No

The audio format. Supported formats are pcm, opus, mp3, and raw-opus. The default is pcm.

The Qwen-TTS model supports only pcm.

Note: The difference between opus and raw-opus is that each packet in the opus format has an extra Ogg encapsulation (RFC 7845).

frame_size

int

No

The frame size of the synthesized audio, in milliseconds. Valid values:

  • 10

  • 20

  • 40

  • 60

  • 100

  • 120

The default is 60.

This parameter applies only when the audio format for speech synthesis is opus or raw-opus.

volume

int

No

The volume of the synthesized audio. Valid range: 0–100. Default: 50.

speech_rate

int

No

The speech rate of the synthesized audio, as a percentage of the normal speed. Valid range: 50–200. Default: 100.

pitch_rate

int

No

The pitch of the synthesized audio. Valid range: 50–200. Default: 100.

bit_rate

int

No

The bit rate of the synthesized audio, in kbps. Valid range: 6–510. Default: 32. This parameter applies only when the audio format for speech synthesis is opus or raw-opus.

intermediate_text

string

No

Controls which intermediate text is returned to the user:

  • transcript: Returns the user's speech recognition results.

  • dialog: Returns intermediate results from the dialog system's response.

You can specify multiple types, separated by commas. The default is transcript.

word_timestamp_enabled

boolean

No

Specifies whether to return timestamps for the synthesized audio. If set to true, timestamp information is returned in RespondingContent.extra_info, which can be used for client-side features such as displaying subtitles. The default is false.

This parameter requires intermediate_text to be set to dialog.

Timestamps are returned only for cloned voices and for voices that are specified to support timestamps in the CosyVoice voice list.

transmit_rate_limit

int

No

The rate limit for sending downstream audio, in bytes per second.

incremental_response

boolean

No

Specifies whether to stream results from the large model. If true, results are sent incrementally. If false, the complete result is sent at once. The default is false.

instruction

string

No

Sets instructions to control synthesis effects such as dialect and emotion. This feature applies to qwen3-tts-instruct-flash-realtime, cosyvoice-v3.5-plus, cosyvoice-v3.5-flash, cosyvoice-v3-plus, and cosyvoice-v3-flash.

The instruction parameter has a fixed format. For more information, see the description of the "instruction" parameter in the Java SDK.

The following table describes the parameters for parameters.client_info:

Parameter

Sub-parameter

Type

Required

Description

user_id

string

Yes

The end-user ID. Generate this ID based on your business rules to implement customized features for different end users. The maximum length is 36 characters.

device

uuid

string

No

A globally unique ID for the client device. You must generate this and pass it to the SDK. The maximum length is 40 characters. An end user can have multiple devices, each with a different uuid but the same user_id.

network

ip

string

No

The public IP address of the caller.

location

latitude

string

No

The latitude of the caller. Submit this for business scenarios that require the client's precise location.

longitude

string

No

The longitude of the caller. Submit this for business scenarios that require the client's precise location.

city_name

string

No

The city where the caller is located. This indicates the client's approximate location.

The following table describes the parameters for parameters.biz_params:

Parameter

Type

Required

Description

user_defined_params

json object

No

Pass-through parameters for the agent. For parameters required by different agents, see the Call an Official Agent documentation. You can set dialog extension parameters in the extra_config sub-node. Currently, this supports enable_web_search to enable or disable web search. These settings override the configurations in the console.

user_prompt_params

json object

No

User-defined variables for the prompt. You can define the keys and values in the JSON object. For details on how to configure custom prompt variables in the console, see Application Configuration - Prompts.

user_query_params

json object

No

User-defined dialogue variables. You can define the keys and values in the JSON object. For details on how to configure dialogue variables in the console, see Application Configuration - Dialogue Variables.

Client mode comparison

Comparison item

push2talk

tap2talk

duplex

Type

Client-controlled mode

Tap-to-talk mode

Duplex mode

Audio upload method

On-demand

Continuous

An error occurs if the client does not upload audio for more than 20 seconds in the Listening state.

Continuous

An error occurs if the client does not upload audio for more than 20 seconds in any state.

VAD detection

Client

Server

Server

Interruption method

Interruption via a RequestToSpeak message

Interruption via a RequestToSpeak message

Voice interruption

Use case

The client application controls when to send audio for recognition. This is suitable for push-to-talk scenarios where the user presses a button to speak and releases it to stop.

The client continuously uploads audio, and the server detects voice activity. This mode does not support voice interruption; you must send a RequestToSpeak message to interrupt.

The client continuously uploads audio, and the server detects voice activity. The user can interrupt the model's output by speaking at any time.

Example:

{
    "header": {
        "action":"run-task",
        "task_id": "f894c16f-f20e-4c1d-837e-89e0fbc63a43",
        "streaming":"duplex"
    },
    "payload": {
        "task_group":"aigc",
        "task":"multimodal-generation", // Task: multimodal generation
        "function":"generation",
        "model":"multimodal-dialog", // Model: multimodal dialog (note this is different from the task)
        "input":{
          "directive": "Start",
          "workspace_id": "llm-***********",
          "app_id": "****************"
        },
        "parameters":{
          "upstream":{
            "type": "AudioOnly",
            "mode": "duplex"
          },
          "downstream":{
            "voice": "longxiaochun_v2",
            "sample_rate": 24000
          },
          "client_info":{
            "user_id": "bin********207",
            "device":{
              "uuid": "432k*********k449"
            },
            "network":{
              "ip": "203.0.113.10"
            },
            "location":{
              "city_name": "Beijing"
            }
          },
          "biz_params":{
            "user_defined_params": {
                "extra_config": {
                    "enable_web_search": false
                },
                "agent_id_xxxxx": {
                    "name": "value"
                }
            },
            "user_prompt_params": {
                "name": "value"
            },
            "user_query_params": {
                "name": "value"
            }
          }
        }
    }
}

Started - Output Message

NoteAfter receiving the Started message, do not send audio immediately. You must wait for a DialogStateChanged message that confirms the session has entered the Listening state before sending audio.

Level 1 Parameter

Level 2 Parameter

Type

Description

output

event

string

The name of the event: Started

dialog_id

string

The dialog ID.

Example:

Sample Started response:

{
    "header": {
        "event":"result-generated",
        "task_id": "f894c16f-f20e-4c1d-837e-89e0fbc63a43"
    },
    "payload": {
        "output":{
          "event": "Started",
          "dialog_id": "dd84xxxx-xxxx-xxxx-xxxx-xxxxb7bb"
        }
    }
}

End a session

Stop - Input Message

Level 1 Parameter

Level 2 Parameter

Type

Required

Description

input

directive

string

Yes

The name of the instruction: Stop

dialog_id

string

No

The dialog ID.

Example:

{
    "header": {
        "action":"finish-task",
        "task_id": "f894c16f-f20e-4c1d-837e-89e0fbc63a43",
        "streaming":"duplex"
    },
    "payload": {
        "input":{
          "directive": "Stop",
          "dialog_id": "dd84xxxx-xxxx-xxxx-xxxx-xxxxb7bb"
        }
    }
}

Stopped - Output Message

Level 1 Parameter

Level 2 Parameter

Type

Required

Description

output

event

string

Yes

The name of the event: Stopped

dialog_id

string

Yes

The dialog ID.

Example:

{
    "header": {
        "event":"result-generated",
        "task_id": "f894c16f-f20e-4c1d-837e-89e0fbc63a43"
    },
    "payload": {
        "output":{
          "event": "Stopped",
          "dialog_id": "dd84xxxx-xxxx-xxxx-xxxx-xxxxb7bb"
        }
    }
}

Server state change event

DialogStateChanged - Output Message

Level 1 Parameter

Level 2 Parameter

Type

Required

Description

output

event

string

Yes

The name of the event: DialogStateChanged

state

string

Yes

The AI interaction state. Valid values: Listening, Thinking, and Responding.

Note: The Listening state indicates that the SDK can send audio to the server, but it does not indicate whether the client's microphone is active.

dialog_id

string

Yes

The dialog ID.

Example:

{
    "header": {
        "event":"result-generated",
        "task_id": "f894c16f-f20e-4c1d-837e-89e0fbc63a43"
    },
    "payload": {
        "output":{
          "event": "DialogStateChanged",
          "state": "Listening",
          "dialog_id": "dd84xxxx-xxxx-xxxx-xxxx-xxxxb7bb"
        }
    }
}

Request to upload voice

RequestToSpeak - Input Message

If a user wants to speak when the state is not Listening, they can send this directive to interrupt the large language model's (LLM) response.

The specific user action that triggers this event varies by interaction type, such as pressing a button or interrupting with a voice command (requires duplex mode).

Level 1 Parameter

Level 2 Parameter

Type

Required

Description

input

directive

string

Yes

The name of the instruction: RequestToSpeak

dialog_id

string

No

The dialog ID.

Example:

{
    "header": {
        "action":"continue-task",
        "task_id": "f894c16f-f20e-4c1d-837e-89e0fbc63a43",
        "streaming":"duplex"
    },
    "payload": {
        "input":{
          "directive": "RequestToSpeak",
          "dialog_id": "dd84xxxx-xxxx-xxxx-xxxx-xxxxb7bb"
        }
    }
}

RequestAccepted - Output Message

Level 1 Parameter

Level 2 Parameter

Type

Required

Description

output

event

string

Yes

The name of the event: RequestAccepted

dialog_id

string

Yes

The dialog ID.

Example:

{
    "header": {
        "event":"result-generated",
        "task_id": "f894c16f-f20e-4c1d-837e-89e0fbc63a43"
    },
    "payload": {
        "output":{
          "event": "RequestAccepted",
          "dialog_id": "dd84xxxx-xxxx-xxxx-xxxx-xxxxb7bb"
        }
    }
}

Upload voice command

SendSpeech - Input Message

When the Start message specifies the push2talk mode, the client must notify the server when the user starts speaking. In the Listening state, the client sends a SendSpeech message when the user presses a button. This message signals that an audio upload is about to begin. The client must send the audio data immediately after this message.

Level 1 Parameter

Level 2 Parameter

Type

Required

Description

input

directive

string

Yes

The name of the instruction: SendSpeech

dialog_id

string

No

The dialog ID.

Example:

{
    "header": {
        "action":"continue-task",
        "task_id": "f894c16f-f20e-4c1d-837e-89e0fbc63a43",
        "streaming":"duplex"
    },
    "payload": {
        "input":{
          "directive": "SendSpeech",
          "dialog_id": "dd84xxxx-xxxx-xxxx-xxxx-xxxxb7bb"
        }
    }
}

StopSpeech - Input Message

When the Start message specifies the push2talk mode, the client must send a StopSpeech message when the user finishes speaking and releases the button to signal the end of the voice input.

Level 1 Parameter

Level 2 Parameter

Type

Required

Description

input

directive

string

Yes

The name of the instruction: StopSpeech

dialog_id

string

No

The dialog ID.

Example:

{
    "header": {
        "action":"continue-task",
        "task_id": "f894c16f-f20e-4c1d-837e-89e0fbc63a43",
        "streaming":"duplex"
    },
    "payload": {
        "input":{
          "directive": "StopSpeech",
          "dialog_id": "dd84xxxx-xxxx-xxxx-xxxx-xxxxb7bb"
        }
    }
}

CancelSpeech - Input message

In push2talk or tap2talk mode, the client can send a CancelSpeech message to terminate the voice input. The server then stops the recognition process and returns to an idle state.

Level 1 parameter

Level 2 parameter

Type

Required

Description

input

directive

string

Yes

Directive name: CancelSpeech

dialog_id

string

No

dialog ID

Example:

{
    "header": {
        "action":"continue-task",
        "task_id": "f894c16f-f20e-4c1d-837e-89e0fbc63a43",
        "streaming":"duplex"
    },
    "payload": {
        "input":{
          "directive": "CancelSpeech",
          "dialog_id": "dd84xxxx-xxxx-xxxx-xxxx-xxxxb7bb"
        }
    }
}

Speech recognition start/end

SpeechStarted - Output Message

The server sends this event when it detects the start of speech for automatic speech recognition (ASR).

Level 1 Parameter

Level 2 Parameter

Type

Required

Description

output

event

string

Yes

The name of the event: SpeechStarted

dialog_id

string

Yes

The dialog ID.

Example:

{
    "header": {
        "event":"result-generated",
        "task_id": "f894c16f-f20e-4c1d-837e-89e0fbc63a43"
    },
    "payload": {
        "output":{
          "event": "SpeechStarted",
          "dialog_id": "dd84xxxx-xxxx-xxxx-xxxx-xxxxb7bb"
        }
    }
}

SpeechEnded - Output Message

The server sends this event when it detects the end of speech for ASR. If the client is still uploading audio, it must stop after it receives this event.

Level 1 Parameter

Level 2 Parameter

Type

Required

Description

output

event

string

Yes

The name of the event: SpeechEnded

dialog_id

string

Yes

The dialog ID.

Example:

{
    "header": {
        "event":"result-generated",
        "task_id": "f894c16f-f20e-4c1d-837e-89e0fbc63a43"
    },
    "payload": {
        "output":{
          "event": "SpeechEnded",
          "dialog_id": "dd84xxxx-xxxx-xxxx-xxxx-xxxxb7bb"
        }
    }
}

Specify response content

RequestToRespond - Input Message

In the Listening state, this message notifies the server to proactively interact with the user. Based on the type field, the server either directly converts the text in the text field to speech and sends it, or calls the large model and then converts the result to speech before sending it.

Level 1 Parameter

Level 2 Parameter

Type

Required

Description

input

directive

string

Yes

The name of the instruction: RequestToRespond

dialog_id

string

Yes

The dialog ID.

type

string

Yes

Specifies the interaction type. Two values are supported:

  • transcript: Directly convert text to speech.

  • prompt: Send the text to the large model for a response.

text

string

Yes

The text to process. Must not be null.

  • When calling some agents, text can be an empty string ("") if the server only needs the images or biz_params parameters. For details, see Calling Official Agents.

parameters

images

list[]

No

Information about the images to be analyzed.

biz_params

object

No

For details, see the parameters.biz_params table below.

The following table describes parameters.biz_params:

Parameter

Type

Required

Description

videos

list[]

No

Manages video call connections.

Example:

Setting payload.biz_params.videos.action to connect enters video mode.

Setting payload.biz_params.videos.action to exit exits video mode.

Other parameters

No

This parameter is the same as parameters.biz_params in the Start message, and is used to pass custom parameters to the dialog system. The biz_params parameter in RequestToRespond is effective only for the current request.

NoteIn addition to the videos parameter, parameters.biz_params is the same as the parameters.biz_params in the Start message and is used to pass custom parameters to the dialog system. The biz_params parameter of RequestToRespond is valid only for the current request.

{
    "header": {
        "action":"continue-task",
        "task_id": "f894c16f-f20e-4c1d-837e-89e0fbc63a43",
        "streaming":"duplex"
    },
    "payload": {
        "input":{
          "directive": "RequestToRespond",
          "dialog_id": "dd84xxxx-xxxx-xxxx-xxxx-xxxxb7bb",
          "type": "prompt",
          "text": "Hello, what would you like to talk about?"
        },
        "parameters":{
          "images":[{
              "type": "base64",
              "value": "aGVsbG8gd29ybGQ="
          }],
          "biz_params":{
            "user_defined_params":{},
            "videos": [
                  {
                    "action": "connect/exit",
                    "type": "voicechat_video_channel"
                  }
                ]
          }
        }
    }
}

AI voice response status

RespondingStarted - Output Message

The AI voice response has started. The SDK must prepare to receive audio data from the server.

Level 1 Parameter

Level 2 Parameter

Type

Required

Description

output

event

string

Yes

The name of the event: RespondingStarted

dialog_id

string

Yes

The dialog ID.

Example:

{
    "header": {
        "event":"result-generated",
        "task_id": "f894c16f-f20e-4c1d-837e-89e0fbc63a43"
    },
    "payload": {
        "output":{
          "event": "RespondingStarted",
          "dialog_id": "dd84xxxx-xxxx-xxxx-xxxx-xxxxb7bb"
        }
    }
}

RespondingEnded - Output Message

The AI voice response has ended.

Level 1 Parameter

Level 2 Parameter

Type

Required

Description

output

event

string

Yes

The name of the event: RespondingEnded

dialog_id

string

Yes

The dialog ID.

Example:

{
    "header": {
        "event":"result-generated",
        "task_id": "f894c16f-f20e-4c1d-837e-89e0fbc63a43"
    },
    "payload": {
        "output":{
          "event": "RespondingEnded",
          "dialog_id": "dd84xxxx-xxxx-xxxx-xxxx-xxxxb7bb"
        }
    }
}

Client-side playback event

LocalRespondingStarted - Input Message

The client starts playing the audio sent from the server.

Level 1 Parameter

Level 2 Parameter

Type

Required

Description

input

directive

string

Yes

The name of the instruction: LocalRespondingStarted

dialog_id

string

No

The dialog ID.

Example:

{
    "header": {
        "action":"continue-task",
        "task_id": "f894c16f-f20e-4c1d-837e-89e0fbc63a43",
        "streaming":"duplex"
    },
    "payload": {
        "input":{
          "directive": "LocalRespondingStarted",
          "dialog_id": "dd84xxxx-xxxx-xxxx-xxxx-xxxxb7bb"
        }
    }
}

LocalRespondingEnded - Input Message

The client has finished playing audio from the server. This message confirms that playback is complete, prompting the server to end the current question and answer session and return to the Listening state.

Level 1 Parameter

Level 2 Parameter

Type

Required

Description

input

directive

string

Yes

The name of the instruction: LocalRespondingEnded

dialog_id

string

No

The dialog ID.

Example:

{
    "header": {
        "action":"continue-task",
        "task_id": "f894c16f-f20e-4c1d-837e-89e0fbc63a43",
        "streaming":"duplex"
    },
    "payload": {
        "input":{
          "directive": "LocalRespondingEnded",
          "dialog_id": "dd84xxxx-xxxx-xxxx-xxxx-xxxxb7bb"
        }
    }
}

Text delivery event

SpeechContent - Output Message

Level 1 Parameter

Level 2 Parameter

Type

Required

Description

output

event

string

Yes

The name of the event: SpeechContent

dialog_id

string

Yes

The dialog ID.

text

string

Yes

The text from speech recognition, returned as a streaming output. Each message contains the full, cumulative text recognized so far. For example, you might first receive "text": "Hello", followed by "text": "Hello world".

finished

bool

Yes

Indicates whether the output is finished.

Example:

{
    "header": {
        "event":"result-generated",
        "task_id": "f894c16f-f20e-4c1d-837e-89e0fbc63a43"
    },
    "payload": {
        "output":{
          "event": "SpeechContent",
          "dialog_id": "dd84xxxx-xxxx-xxxx-xxxx-xxxxb7bb",
          "text": "One two three",
          "finished": false
        }
    }
}

RespondingContent - Output Message

Level 1 Parameter

Level 2 Parameter

Type

Required

Description

output

event

string

Yes

The event name: RespondingContent.

dialog_id

string

Yes

The dialog ID.

round_id

string

Yes

The ID of the current interaction round.

llm_request_id

string

Yes

The request_id for the large language model (LLM) call.

text

string

Yes

The text output by the system. This is a full streaming output.

spoken

string

Yes

The text used for speech synthesis. This is a full streaming output.

finished

bool

Yes

Indicates whether the output is finished.

extra_info

object

No

Other extension information. Currently supports:

  • commands: A command string. This field is a JSON-formatted string that you must parse separately. For command strings used by various agents, see Calling Official Agents.

  • agent_info: Agent information.

  • tool_calls: Information returned from a plugin.

  • word_timestamps: Word-level timestamps for the corresponding text-to-speech (TTS) output.

Omit this field if no extra information is available.

Example:

{
    "header": {
        "event": "result-generated",
        "task_id": "f894c16f-f20e-4c1d-837e-89e0fbc63a43"
    },
    "payload": {
        "output": {
            "event": "RespondingContent",
            "dialog_id": "dd84xxxx-xxxx-xxxx-xxxx-xxxxb7bb",
            "text": "You entered the number sequence \"12345\". If you have any questions about these numbers or need me to perform a task with them, please provide more details, and I'll do my best to help you.",
            "spoken": "You entered the number sequence \"12345\". If you have any questions about these numbers or need me to perform a task with them, please provide more details, and I'll do my best to help you.",
            "finished": true,
            "extra_info": {
                "commands": "[{\"name\":\"VOLUME_SET\",\"params\":[{\"name\":\"series\",\"normValue\":\"70\",\"value\":\"70\"}]}]",
                "tool_calls": [
                    {
                        "id": "",
                        "type": "function",
                        "function": {
                            "name": "function_name",
                            "arguments": "{\"id\": \"123\", \"name\": \"test\"}",
                            "outputs": "{\"result\": \"success\"}",
                            "status": {
                                "code": 200,
                                "message": "Success."
                            }
                        }
                    }
                ]
            }
        }
    }
}

Client-side update event

UpdateInfo - Input Message

Level 1 Parameter

Level 2 Parameter

Level 3 Parameter

Type

Required

Description

input

directive

string

Yes

The name of the instruction: UpdateInfo

dialog_id

string

No

The dialog ID.

parameters

images

list[]

No

Image data.

client_info

status

object

No

The current status of the client.

biz_params

object

No

Specifies custom parameters for the dialog system, similar to biz_params in the Start message. Each parameter provided here overwrites the one with the same name set in the Start message. The updated parameters remain in effect for all subsequent dialogs in the current connection.

upstream

language

string

No

Updates the speech recognition language. If omitted, the language remains unchanged. To specify multiple languages, separate them with commas, for example: zh,en.

downstream

language

string

No

Updates the speech synthesis language. If omitted, the language remains unchanged.

Example:

{
    "header": {
        "action":"continue-task",
        "task_id": "f894c16f-f20e-4c1d-837e-89e0fbc63a43",
        "streaming":"duplex"
    },
    "payload": {
        "input":{
          "directive": "UpdateInfo",
          "dialog_id": "dd84xxxx-xxxx-xxxx-xxxx-xxxxb7bb"
        },
        "parameters":{
          "images":[{
              "type": "base64",
              "value": "base64String"
          }],
          "client_info": {
              "status": {
                  "bluetooth_announcement": {
                      "status": "stopped"
                  },
                  "stream_media_playback": {
                      "status": "stopped"
                  },
                  "phone_ringing": {
                      "status": "stopped"
                  },
                  "stream_media_playback_qq": {
                      "status": "stopped"
                  }
              }
          },
          "biz_params":{
          }
        }
    }
}

Heartbeat event

Periodically send this message to the server to prevent connection timeouts. To account for network latency and processing time, we recommend sending it every 50 seconds. The server will respond with the same message, which the client can safely ignore.

HeartBeat - Input Message

Level 1 Parameter

Level 2 Parameter

Type

Required

Description

input

directive

string

Yes

The name of the instruction: HeartBeat

dialog_id

string

No

The dialog ID.

Example:

{
  "header": {
    "action": "continue-task",
    "task_id": "f894c16f-f20e-4c1d-837e-89e0fbc63a43",
    "streaming": "duplex"
  },
  "payload": {
    "input": {
      "directive": "HeartBeat",
      "dialog_id": "dd84xxxx-xxxx-xxxx-xxxx-xxxxb7bb"
    }
  }
}

HeartBeat - Output Message

Level 1 Parameter

Level 2 Parameter

Type

Required

Description

output

event

string

Yes

The name of the event: HeartBeat

dialog_id

string

Yes

The dialog ID.

Example:

{
    "header": {
        "event":"result-generated",
        "task_id": "f894c16f-f20e-4c1d-837e-89e0fbc63a43"
    },
    "payload": {
        "output":{
          "event": "HeartBeat",
          "dialog_id": "dd84xxxx-xxxx-xxxx-xxxx-xxxxb7bb"
        }
    }
}

Error event

Error - Output Message

Level 1 Parameter

Level 2 Parameter

Type

Required

Description

output

error_code

int

Yes

The error code.

error_name

string

Yes

The name of the error.

error_message

string

Yes

The error message.

{
    "header": {
        "event":"result-generated",
        "task_id": "f894c16f-f20e-4c1d-837e-89e0fbc63a43"
    },
    "payload": {
        "output":{
          "event": "Error",
          "dialog_id": "dd84xxxx-xxxx-xxxx-xxxx-xxxxb7bb",
          "error_code": 500,
          "error_name": "InternalLLMError",
          "error_message": "Internal LLM error"
        }
    }
}

Error codes

For troubleshooting, see multimodal interaction service - error codes.

If the issue persists, contact technical support with the full request_id and dialog_id.

Glossary

VAD: Voice Activity Detection

ASR: Automatic Speech Recognition

TTS: Text-to-Speech

LLM: Large Language Model