The Qwen-Audio-3.0-ASR-Flash-Streaming/Fun-ASR-Realtime real-time speech recognition service pushes server-side events to the client over WebSocket. This topic describes the data structures and field descriptions of the four event types: task-started, result-generated, task-finished, and task-failed.
User guide: For model descriptions and selection guidance, see Speech-to-text.
Event interaction flow: To understand the event interaction sequence, see WebSocket API.
task-started
Description: The task has started successfully. The client can begin sending audio data.
header | |
payload This value is always |
result-generated
Description: The recognition result. It includes intermediate results (sentence_end=false) and final results (sentence_end=true).
header | |
payload |
task-finished
Description: The task has ended normally. You can close the connection or reuse it.
header | |
payload You can ignore this field. This value is usually |
task-failed
Description: The task has failed. The connection is closed and cannot be reused.
header | |
payload This value is always |