Speech-to-speech overview
Choose a model for voice conversation, speech translation, or simultaneous interpretation.
Migrate from closed-source models
Map your current OpenAI Realtime or Gemini Live setup to an equivalent Bailian model.
Closed-source examples | Bailian recommendation | |
|---|---|---|
Real-time conversation | OpenAI GPT Realtime, Gemini 3.1 Live |
|
Cost-sensitive conversation | OpenAI gpt-4o-mini Realtime |
|
Real-time translation | Gemini 3.1 Live |
|
This page covers speech-to-speech. For visual understanding, audio/video analysis, and content moderation, see Omni-modal. For tasks that require reasoning and text output, such as video analysis and content labeling, see Qwen3.8-Omni-Flash.
S2S (speech-to-speech) vs. pipeline
There are two approaches to building voice applications:
S2S | Pipeline (ASR + LLM + TTS) | |
|---|---|---|
Latency | Low -- single-model stream processing | Higher -- three-stage serial processing |
Audio understanding | End-to-end -- perceives tone and emotion and responds accordingly | Converts to text before processing, losing subtle audio cues |
Voice customization | Selection of preset voices via a system prompt | Voice cloning and voice design (CosyVoice) |
- Use S2S for low latency, audio-aware responses, and interactive conversation.
- Use a pipeline when you need voice customization or want to select ASR, LLM, and TTS models independently.
This page covers the S2S single-model approach (Omni and Livetranslate series). For the pipeline approach, select each component separately:
- ASR (speech recognition):Speech-to-text
- LLM (large language model):Text generation
- TTS (text-to-speech):Speech synthesis
Real-time or file mode?
- Real-time (WebSocket): Voice assistants, call centers, and simultaneous interpretation. Streams audio input and speech output.
- File mode (HTTP): Higher latency but better quality. Ideal for video dubbing, podcast translation, and offline processing. Also supports function calling, web search, thinking mode, and video context (see Companion capabilities below).
Choose a model by use case (S2S single-model approach)
All use cases below use the S2S single-model approach. For the pipeline approach, use the ASR, LLM, and TTS guides linked above.
Use case | Recommended model | API |
|---|---|---|
Voice assistants and customer-service conversations |
| WebSocket |
Cost-sensitive conversations |
| WebSocket |
Simultaneous interpretation and live translation |
| WebSocket |
Video dubbing and podcast translation |
| Chat Completions |
Semantic VAD voice assistants and smart customer service (with Function Calling support) |
| WebSocket |
Companion capabilities of the S2S single-model approach
The following sections cover tool calling, web search, and text reasoning in voice applications.
Function calling
To query knowledge bases, check schedules, or trigger workflows based on audio/video content, use Qwen-Audio Realtime (WebSocket).
Web search
To retrieve current information and generate spoken responses, use Qwen3.5-Omni (WebSocket or HTTP, both Plus and Flash). The model decides autonomously whether to search. Qwen-Audio Realtime does not support this capability.
Not supported by Qwen3-Omni-Flash or the Livetranslate model.
Thinking mode
Qwen3-Omni-Flash does not support voice generation in thinking mode.
Speech translation
The following model series support speech translation:
- Qwen3.8-Livetranslate: Supports 60 source languages, speech output in 29 languages, and audio and image input. See Model information.
- Qwen3.5-Livetranslate: 60 languages (29 with audio+text output, 31 text-only). Covers Chinese, English, French, German, Russian, Japanese, Korean, Spanish, Portuguese, Arabic, and more.
- Qwen3-Livetranslate: 18 languages and 5 Chinese dialects (~3 s latency). File mode accepts video input for context-aware translations. 7 languages produce text-only output.
- Qwen3.5-Omni: 29 output languages and 8 Chinese dialects. Strong audio/video understanding and web search. Inject terminology and domain context via system prompt. Real-time and file modes.
- Qwen3-Omni-Flash: 11 output languages and 8 Chinese dialects. Inject terminology and domain context via system prompt. Real-time and file modes.
NoteUse Livetranslate for ready-to-use translation. Choose Qwen3.5-Omni for offline speech output with web search and terminology injection.
Supported languages
Language | Qwen3.5-Livetranslate | Qwen3-Livetranslate | Qwen3.5-Omni | Qwen3-Omni-Flash |
|---|---|---|---|---|
English | Supported | Supported | Supported | Supported |
Chinese (Mandarin) | Supported | Supported | Supported | Supported |
Cantonese | Text-only | Supported | Supported | Supported |
Sichuan dialect | Supported | Supported | Supported | Supported |
Shanghainese | Supported | Supported | Supported | Supported |
Beijing dialect | Supported | Supported | Supported | Supported |
Tianjin dialect | Supported | Supported | Supported | Supported |
Nanjing dialect | -- | -- | Supported | Supported |
Shaanxi dialect | -- | -- | Supported | Supported |
Minnan dialect | -- | -- | Supported | Supported |
French | Supported | Supported | Supported | Supported |
German | Supported | Supported | Supported | Supported |
Russian | Supported | Supported | Supported | Supported |
Italian | Supported | Supported | Supported | Supported |
Spanish | Supported | Supported | Supported | Supported |
Portuguese | Supported | Supported | Supported | Supported |
Japanese | Supported | Supported | Supported | Supported |
Korean | Supported | Supported | Supported | Supported |
Thai | Supported | Text-only | Supported | Supported |
Indonesian | Supported | Text-only | Supported | -- |
Vietnamese | Supported | Text-only | Supported | -- |
Arabic | Supported | Text-only | Supported | -- |
Hindi | Supported | Text-only | Supported | -- |
Turkish | Supported | Text-only | Supported | -- |
Finnish | Supported | -- | Supported | -- |
Polish | Supported | -- | Supported | -- |
Dutch | Supported | -- | Supported | -- |
Czech | Supported | -- | Supported | -- |
Urdu | Supported | -- | Supported | -- |
Tagalog | Supported | -- | Supported | -- |
Swedish | Supported | -- | Supported | -- |
Danish | Supported | -- | Supported | -- |
Hebrew | Supported | -- | Supported | -- |
Icelandic | Supported | -- | Supported | -- |
Malay | Supported | -- | Supported | -- |
Norwegian | Supported | -- | Supported | -- |
Persian | Supported | -- | Supported | -- |
Greek | Text-only | Text-only | -- | -- |
Afrikaans | Text-only | -- | -- | -- |
Asturian | Text-only | -- | -- | -- |
Belarusian | Text-only | -- | -- | -- |
Bulgarian | Text-only | -- | -- | -- |
Bengali | Text-only | -- | -- | -- |
Bosnian | Text-only | -- | -- | -- |
Catalan | Text-only | -- | -- | -- |
Cebuano | Text-only | -- | -- | -- |
Estonian | Text-only | -- | -- | -- |
Galician | Text-only | -- | -- | -- |
Gujarati | Text-only | -- | -- | -- |
Croatian | Text-only | -- | -- | -- |
Hungarian | Text-only | -- | -- | -- |
Javanese | Text-only | -- | -- | -- |
Kazakh | Text-only | -- | -- | -- |
Kannada | Text-only | -- | -- | -- |
Kyrgyz | Text-only | -- | -- | -- |
Latvian | Text-only | -- | -- | -- |
Macedonian | Text-only | -- | -- | -- |
Malayalam | Text-only | -- | -- | -- |
Marathi | Text-only | -- | -- | -- |
Punjabi | Text-only | -- | -- | -- |
Romanian | Text-only | -- | -- | -- |
Slovak | Text-only | -- | -- | -- |
Slovenian | Text-only | -- | -- | -- |
Swahili | Text-only | -- | -- | -- |
Tajik | Text-only | -- | -- | -- |
Azerbaijani | Text-only | -- | -- | -- |
Ukrainian | Text-only | -- | -- | -- |
"Supported" = speech + text output. "Text-only" = text output only, no speech.
Qwen3.5-Omni supports 113 input languages and dialects.
Qwen3.5-Livetranslate supports 60 languages (29 with audio and text, 31 text only).
The legacy qwen-omni-turbo model supports only Chinese and English.
Recommended models
The table lists the entry-point model in each series. To pin a dated version for regression testing or stability, see All models below.
Model | API | Use cases |
|---|---|---|
qwen3.5-omni-plus-realtime / qwen3.5-omni-flash-realtime | WebSocket | Realtime audio/video conversation |
qwen-audio-3.0-realtime-plus / qwen-audio-3.0-realtime-flash | WebSocket | Realtime voice conversation |
qwen3.5-omni-plus / qwen3.5-omni-flash | Chat Completions | Offline speech output |
qwen3.8-livetranslate-flash-realtime | WebSocket | Realtime translation |
qwen3-livetranslate-flash | Chat Completions | Audio/video file translation |
All models
Qwen-Audio
Model | API | Input | Function calling | Web search | Thinking mode | Translation |
|---|---|---|---|---|---|---|
| WebSocket | audio, text | Supported | -- | -- | -- |
| WebSocket | audio, text | Supported | -- | -- | -- |
Qwen3.5-Omni
Model | API | Input | Function calling | Web search | Thinking mode |
|---|---|---|---|---|---|
| WebSocket | Text, audio, image, video | Supported | Supported | -- |
| WebSocket | Text, audio, image, video | Supported | Supported | -- |
| Chat Completions | Text, audio, image, video | Supported (Beijing, text output) | Supported | -- |
| Chat Completions | Text, audio, image, video | Supported (Beijing, text output) | Supported | -- |
| WebSocket | Text, audio, image, video | Supported | Supported | -- |
| WebSocket | Text, audio, image, video | Supported | Supported | -- |
| Chat Completions | Text, audio, image, video | Supported (Beijing, text output) | Supported | -- |
| Chat Completions | Text, audio, image, video | Supported (Beijing, text output) | Supported | -- |
Qwen3-Omni
Model | API | Input | Function calling | Web search | Thinking mode |
|---|---|---|---|---|---|
| WebSocket | Text, audio, image, video | -- | -- | -- |
| WebSocket | Text, audio, image, video | -- | -- | -- |
| WebSocket | Text, audio, image, video | -- | -- | -- |
| Chat Completions | Text, audio, image, video | Supported | -- | Supported |
| Chat Completions | Text, audio, image, video | Supported | -- | Supported |
| Chat Completions | Text, audio, image, video | Supported | -- | Supported |
Qwen3.8-Livetranslate
Model | API | Input | Languages |
|---|---|---|---|
WebSocket | Audio, image | 60 |
Qwen3.5-Livetranslate
Model | API | Input | Languages |
|---|---|---|---|
| WebSocket | Audio, images | 60 |
| WebSocket | Audio | 60 |
Qwen3-Livetranslate
Model | API | Input | Languages |
|---|---|---|---|
| WebSocket | Audio | 18 |
| WebSocket | Audio | 18 |
| Chat Completions | Audio, video | 18 |
| Chat Completions | Audio, video | 18 |
Legacy models
These models are no longer updated. For new projects, use Qwen3.5-Omni.
Model | Input | API |
|---|---|---|
| Text, audio, image, video | HTTP |
| Text, audio, image, video | HTTP |
| Text, audio, image, video | HTTP |
| Text, audio, image, video | HTTP |
| Text, audio | WebSocket |
| Text, audio | WebSocket |
| Text, audio | WebSocket |
What's next
API documentation by model series:
- Qwen-Audio Realtime (WebSocket, real-time voice conversation): Realtime Audio Chat (Qwen-Audio-Realtime)
- Qwen3.5-Omni and Qwen3-Omni (WebSocket, real-time): Qwen-Omni-Realtime
- Qwen3.5-Omni (HTTP, offline speech output): Non-real-time (Qwen-Omni)
- Qwen3.8-Livetranslate / Qwen3.5-Livetranslate (real-time): Real-time audio and video translation - Qwen
- Qwen3-Livetranslate (HTTP, file): Audio and video file translation (Qwen)