Qwen-Audio client events
Client event reference for the Qwen-Audio Realtime API.
User guide: Realtime Audio Chat (Qwen-Audio-Realtime). For event interaction sequences, see WebSocket API.
session.update
Description: After a connection is established, send this event to update the default session configuration. Include only the fields you want to change; omitted fields retain their current values. If any parameter is invalid, the server returns an error. If all parameters are valid, the server applies the changes and returns the full configuration.
Noteturn_detection can only be modified before the first audio is sent (IDLE state).
type Event type. Fixed value: session Session configuration. | Language configuration (qwen-audio-3.1-realtime-plus)After connecting to Function Calling: Voiceprint registration (voiceprint_audio_urls): |
input_audio_buffer.append
Description: Appends audio data to the input buffer. Send this event continuously at a high frequency — for example, one chunk every 20–40 ms. The server sends no acknowledgment for this event.
type Event type. Fixed value: audio Base64-encoded audio data. | |
input_audio_buffer.commit
Description: Push-to-talk mode only. Commits the buffered audio as a user message. This doesn't automatically trigger inference. Send response.create to trigger inference manually.
This event is ignored in server_vad and smart_turn modes.
type Event type. Fixed value: | |
input_audio_buffer.clear
Description: Push-to-talk mode only. Clears uncommitted audio from the buffer. This event is ignored in server_vad and smart_turn modes. The server responds with an input_audio_buffer.cleared event.
type Event type. Fixed value: | |
conversation.item.create
Description: Inserts a conversation item into the conversation context. Use this event to inject historical context or add text content, or to return Function Calling results.
NoteIf item.id already exists in the conversation, an error is returned and the item isn't created.
type Event type. Fixed value: previous_item_id Specifies the conversation item after which the new item is inserted. If not provided, the item is appended to the end of the conversation. item The conversation item to create. | Inject a user text message: Return a Function Calling result: |
conversation.item.retrieve
Description: Retrieves a conversation item stored on the server. Audio-type content in the response contains only the transcript (transcript), not the original audio data.
type Event type. Fixed value: item_id ID of the conversation item to retrieve. The server returns the result in a | |
conversation.item.delete
Description: Deletes a conversation item from the conversation context. The server confirms the deletion with a conversation.item.deleted event.
type Event type. Fixed value: item_id ID of the conversation item to delete. | |
response.create
Description: Triggers model inference. Behavior varies by mode:
- Push-to-talk mode: Must be called manually. Commit the buffered audio with
input_audio_buffer.commitfirst, or return afunction_call_outputresult before triggering. Can't be called while a response is being generated. - server_vad mode: Typically triggered automatically by the server. Clients can also call it manually when no response is being generated. Can't be called while a response is being generated.
- smart_turn mode: Can be called while waiting for the next user turn. Can't be called during an active turn (between
input_audio_buffer.speech_startedandresponse.done).
The optional response field overrides the session defaults for the current inference round. In Function Calling scenarios, after the client returns a function_call_output, this event triggers the second inference round.
NoteIn server_vad and smart_turn modes, manually triggered inferences can still be interrupted by new speech.
type Event type. Fixed value: response Overrides the session defaults for the current inference round. If not provided, the current session configuration is used. Properties modalities Overrides the output modalities for the current round. Valid values are the same as voice Overrides the TTS voice for the current round. | |
response.cancel
Description: Cancels the current inference. Any text generated so far is saved to the item list. The server then returns a response.done event with status=cancelled.
An error is returned if no inference is in progress.
type Event type. Fixed value: | |