Call Official Agents
The Multimodal Interaction Development Kit provides a set of agents for daily life, learning, and work scenarios.
News Radio
The News Radio agent delivers trending news daily across technology, entertainment, and other fields. Two AI streamers present and discuss the news.
You can select a program to listen to. You can also interrupt the broadcast at any time to join the discussion with the two AI streamers.
This agent is supported only in multimodal interaction applications.
Configuration
Input parameter | Configuration method | Required | Description |
|---|---|---|---|
podcast_channel | sdk | No | Supported values: technology, entertainment, social |
podcast_name | Console | No | The custom radio program name. The name is announced in the first sentence of the program. The default name is News Radio. You can configure this in the console. |
speaker_1_voice | Console | No | The voice for the first speaker. Valid values depend on the list of voices supported by the text-to-speech (TTS) model configured in your application. |
speaker_2_voice | Console | No | The voice for the second speaker. Valid values depend on the list of voices supported by the TTS model configured in your application. |
To specify the news category to play, set the parameter in the payload.biz_params node of the Start message, as shown in the following example:
{
"user_defined_params": {
"news_radio": {
"podcast_channel": "technology"
}
}
}
How to use
Say “Enter News Radio” during a voice call.
Voice Translation
Note: This agent supports only duplex mode. You must continuously upload audio data throughout the session, whether or not the user is speaking.This agent provides real-time translation similar to simultaneous interpretation. It outputs translations while continuously receiving audio input. It works well for long speech segments, such as presentations and meetings.
Real-time voice playback of translations is currently supported for the following languages:
- Chinese or English: The system uses the voice you specify.
- Japanese or Korean: The system uses the default voice.
Configuration
To enable voice playback of translations, select Translate Voice in the Voice Translation agent settings in the console.
To set the default translation language, pass parameters in the Start instruction when starting a session:
Input parameter | Configuration method | Required | Description |
|---|---|---|---|
modelName | sdk | No | The translation model. Only gummy is supported. |
sourceLanguage | sdk | No | The source language—the language spoken by the user. Supported language codes include the following:
|
translationLanguages | sdk | No | The target translation language. This is a list, but only one language is supported. Language codes match those used for the Chinese → English, Chinese → Japanese, Chinese → Korean, English → Chinese, English → Japanese, English → Korean, (Japanese, Korean, Cantonese, German, French, Russian, Italian, Spanish, Thai, Malay, Indonesian) → (Chinese, English) |
Translate Voice | Console | No | Set whether to convert translation results into speech. Default is disabled. |
To set the source and target languages, add parameters to the payload.biz_params node of the Start message, as shown in the following example:
{
"user_defined_params": {
"voice_translate": {
"modelName": "gummy",
"translationLanguages": "[\"en\"]",
"sourceLanguage": "zh"
}
}
}
How to use
- Say “Turn on Voice Translation” during a voice call. If the SDK request includes both source and target languages, the system enters translation mode immediately. Otherwise, it asks which language to translate from and to.
- Say “Turn on Voice Translation and translate Chinese into English” during a voice call. The system enters translation mode immediately using the specified languages.
- Say “Exit Voice Translation” during translation mode to exit.
Photo Q&A
When the system detects that you want to understand what is in the current view—for example, if you ask “What is in front of me?”—it automatically sends a photo-capture instruction. It then analyzes the image and replies. You do not need to take a photo manually.
This agent is supported only in multimodal interaction applications.
Configuration
In the Photo Q&A agent configuration page in the console, you can set the agent persona and add custom trigger instructions.
How to use
Say “Take a photo to see what I am holding” during a voice call. The service sends a Photo Q&A instruction. After your client receives the instruction, wait for the StateChanged message. When the state changes to Listening, upload the image in the required format. The service analyzes the image and returns the result in speech.
An example of a Photo Q&A instruction from the server appears in the RespondingContent event. The value of output.extra_info.commands is a JSON string where name="visual_qa":
{
"extra_info": {
"commands": "[{\"name\":\"visual_qa\",\"params\":[{\"name\":\"shot\",\"value\":\"Take a photo\",\"normValue\":\"True\"}]}]"
}
}
An example of an image upload message from the client:
{
"header": {
"action":"continue-task",
"task_id": "9B32878******************3D053",
"streaming":"duplex"
},
"payload": {
"input":{
"directive": "RequestToRespond",
"dialog_id": "b3939********************81f7",
"type": "prompt",
"text": "What am I holding?"
},
"parameters":{
"images":[{
"type": "url",
"value": "https://your.server.name/path/abc.jpg"
}]
}
}
}
Video Call
This agent enables real-time video conversations with users. It analyzes visual information such as user actions and surroundings. Your client must upload camera screenshots to the service for analysis using the UpdateInfo message over a WebSocket connection. Upload frequency is recommended at two frames per second.
This agent is supported only in multimodal interaction applications.
Configuration
In the Video Call agent configuration page in the console, you can set the agent persona and add custom start and exit instructions.
How to use
To start a video call: While in Listening state, send a RequestToRespond instruction with video call information in biz_params. The service enters video call mode and sends a prompt. After the prompt finishes playing and the StateChanged message confirms the state has returned to Listening, the video call mode is active.
{
"biz_params": {
"videos": [
{
"action": "connect",
"type": "voicechat_video_channel"
}
]
}
}
After entering video call mode, the conversation flow is the same as a regular voice call. In addition, your client must upload one camera screenshot every 500 ms using the UpdateInfo instruction. We recommend JPEG images at 720p or 480p resolution, no larger than 180 KB. Encode the image binary data in Base64 and include it directly in the parameter. Do not add extra metadata. Note that each instruction supports only one image.
{
"parameters":{
"images":[{
"type": "base64",
"value": "base64String"
}]
}
}
To end a video call: While in Listening state, send a RequestToRespond instruction with video call exit information in biz_params. The service exits video call mode and sends a prompt.
{
"biz_params": {
"videos": [
{
"action": "exit",
"type": "voicechat_video_channel"
}
]
}
}
Helper instructions
To enter or exit video call mode naturally during regular conversations, the service provides intent recognition instructions. These notify your client when the user expresses the relevant intent. After receiving the instruction, wait for the StateChanged message. When the state changes to Listening, send the RequestToRespond request.
For example, if the user says “Enter Video Call” during a voice call, the service detects the intent and sends the following instruction in the RespondingContent message:
{
"extra_info": {
"commands": "[{\"name\":\"open_videochat\",\"params\":[]}]"
}
}
If the user says “Exit Video Call” during a video call, the service detects the intent and sends the following instruction in the RespondingContent message:
{
"extra_info": {
"commands": "[{\"name\":\"quit_videochat\",\"params\":[]}]"
}
}
Ultra-Fast Video Call
Note: This agent supports only duplex mode in multimodal interaction applications. It is unavailable in voice-only applications or other modes. You must continuously upload audio data throughout the session, whether or not the user is speaking. During video calls, you must upload image data as required.Built on the Qwen-Omni model, this agent supports real-time video conversations and analyzes visual information such as user actions and surroundings. Your client must upload camera screenshots to the service for analysis using the UpdateInfo message over a WebSocket connection. Upload frequency is recommended at two frames per second.
Compared to the Video Call agent, Ultra-Fast Video Call responds faster but supports only casual chat. It does not support instructions or extensions.
Configuration
To enable Ultra-Fast Video Call, add Ultra-Fast Video Call in the Console under Agents > Model Studio.
In the Ultra-Fast Video Call agent settings, you can also add custom start and exit instructions.
How to use
Say “Turn on Ultra-Fast Video Call” during a voice call. After the service correctly detects the intent, it sends a prompt and an instruction to start sending video data. Then it enters video call mode. After your client receives the instruction, wait for the dialog state to change to Listening before uploading camera screenshot data.
{
"extra_info": {
"commands": "[{\"name\":\"send_video_stream\",\"params\":[]}]"
}
}
Say “Exit Ultra-Fast Video Call” during a video call. After the service correctly detects the intent, it sends a prompt and an instruction to stop sending video data. Then it exits video call mode. After your client receives the instruction, stop uploading camera screenshot data.
{
"extra_info": {
"commands": "[{\"name\":\"stop_video_stream\",\"params\":[]}]"
}
}
Children's Stories
The Children's Stories agent includes a rich story library. Web search and Large Language Model (LLM) capabilities expand the available stories.
You can request specific stories on demand. You can interrupt the story at any time to ask questions about it or request a rewrite. After the Q&A ends, the story resumes seamlessly from where it left off.
This agent is supported only in multimodal interaction applications.
Configuration
Input parameter | Configuration method | Required | Description |
|---|---|---|---|
speaker_1_voice | Console | No | The voice for story narration. Valid values depend on the list of voices supported by the TTS model configured in your application. |
To set the narration voice, add the parameter to the payload.biz_params node of the Start message, as shown in the following example:
{
"user_defined_params": {
"children_story": {
"speaker_1_voice": "Replace with the voice ID for the TTS model configured for Children's Stories"
}
}
}
How to use
Say “Turn on Story Mode” during a voice call.
Smart Meeting Notes
The Smart Meeting Notes agent analyzes and summarizes real-time or offline audio recordings.
You can use this agent in the Multimodal Interaction Development Kit and the Tongyi Tingwu Agent’s Smart Meeting Notes.
Configuration
Full integration requires coordination between your client-side app and backend services. For complete integration guidance, see: