Omni-modal overview
Choose a model for voice conversation, audio and video analysis, content moderation, voice translation, and other omni-modal tasks.
Migrate from closed-source models
Map your current GPT or Gemini omni-modal/real-time capabilities to an equivalent Bailian model.
Closed-source examples | Bailian recommendation | |
|---|---|---|
Audio/video understanding and text generation | ||
Real-time translation | Gemini 3.5 Live Translate |
|
NoteFor speech output, use qwen3.5-omni-plus.
Use cases
Omni-modal models understand text, audio, images, and video. Qwen3.8-Omni-Flash supports audio/video analysis, meeting summaries, and subtitle generation, with thinking mode, tool calling, and web search. Choose a model for your use case:
Use case | Recommended model | User guide |
|---|---|---|
Audio/video understanding and text generation: Analyze audio/video content and generate meeting summaries, subtitles, and text answers | Qwen3.8-Omni-Flash (Chat Completions / Responses) | |
Real-time voice/video conversation: Interact with AI through a microphone and camera (voice assistants, customer service bots, visual Q&A, live-stream analysis) | Qwen3.5-Omni Realtime (WebSocket) | |
Offline audio output: Upload audio or video files and generate speech responses | Qwen3.5-Omni (Chat Completions) | |
Real-time voice translation: Simultaneous interpretation with approximately 3-second latency, supporting 60 languages (live interpretation, multilingual meetings) | Qwen3.5-Livetranslate (WebSocket) | |
Audio and video file translation: Upload audio or video files and translate them into a target language (video dubbing, podcast translation) | Qwen3-Livetranslate (Chat Completions) | |
Voice cloning: Provide a reference audio clip and the AI generates speech responses in that voice | Qwen3.5-Omni Plus / Flash (Chat Completions / Realtime API) | |
Real-time voice conversation with semantic VAD: End-to-end voice interaction, semantic turn detection (smart_turn), non-meaningful utterances won't interrupt the conversation, supports Function Calling (voice assistants, smart customer service) | Qwen-Audio (WebSocket) |
- When using Qwen3.5-Omni for content analysis, it supports audio up to 3 hours and video up to 1 hour per request.
- Qwen3.8-Omni-Flash (Chat Completions and Responses), Qwen3.5-Omni Plus / Flash (HTTP, text output), Qwen3-Omni-Flash (HTTP only), and Qwen-Audio Realtime (WebSocket) support function calling.
- Qwen3.8-Omni-Flash (Chat Completions / Responses) and Qwen3.5-Omni (Chat Completions / Realtime API) support web search. Qwen3.5-Omni cannot enable web search and function calling at the same time.
Translation
For audio/video translation into text or subtitles, use Qwen3.8-Omni-Flash. For spoken translations, choose a model below based on latency and speech output requirements.
NoteQuick setup: Qwen3.5-Livetranslate (60 languages, ~3 s latency, out-of-the-box). Speech output with web search and term injection: Qwen3.5-Omni (29 output languages, web search and term injection).
Supported languages
Language | Qwen3.5-Livetranslate | Qwen3-Livetranslate | Qwen3.5-Omni | Qwen3-Omni-Flash |
|---|---|---|---|---|
English | Supported | Supported | Supported | Supported |
Chinese (Mandarin) | Supported | Supported | Supported | Supported |
Cantonese | Supported | Supported | Supported Text only | Supported |
Sichuanese | Supported | Supported | Supported | Supported |
Shanghainese | Supported | Supported | Supported | Supported |
Beijingese | Supported | Supported | Supported | Supported |
Tianjinese | Supported | Supported | Supported | Supported |
Nanjingese | Unsupported Text only | Unsupported | Supported | Supported |
Shaanxi dialect | Unsupported Text only | Unsupported | Supported | Supported |
Hokkien | Unsupported Text only | Unsupported | Supported | Supported |
French | Supported | Supported | Supported | Supported |
German | Supported | Supported | Supported | Supported |
Russian | Supported | Supported | Supported | Supported |
Italian | Supported | Supported | Supported | Supported |
Spanish | Supported | Supported | Supported | Supported |
Portuguese | Supported | Supported | Supported | Supported |
Japanese | Supported | Supported | Supported | Supported |
Korean | Supported | Supported | Supported | Supported |
Thai | Supported | Supported Text only | Supported | Supported |
Indonesian | Supported | Supported Text only | Supported | Unsupported |
Vietnamese | Supported | Supported Text only | Supported | Unsupported |
Arabic | Supported | Supported Text only | Supported | Unsupported |
Hindi | Supported | Supported Text only | Supported | Unsupported |
Turkish | Supported | Supported Text only | Supported | Unsupported |
Finnish | Unsupported Text only | Unsupported | Supported | Unsupported |
Polish | Unsupported Text only | Unsupported | Supported | Unsupported |
Dutch | Unsupported Text only | Unsupported | Supported | Unsupported |
Czech | Unsupported Text only | Unsupported | Supported | Unsupported |
Urdu | Unsupported Text only | Unsupported | Supported | Unsupported |
Tagalog | Unsupported Text only | Unsupported | Supported | Unsupported |
Swedish | Unsupported Text only | Unsupported | Supported | Unsupported |
Danish | Unsupported Text only | Unsupported | Supported | Unsupported |
Hebrew | Unsupported Text only | Unsupported | Supported | Unsupported |
Icelandic | Unsupported Text only | Unsupported | Supported | Unsupported |
Malay | Unsupported Text only | Unsupported | Supported | Unsupported |
Norwegian | Unsupported Text only | Unsupported | Supported | Unsupported |
Persian | Unsupported Text only | Unsupported | Supported | Unsupported |
Greek | Supported Text only | Supported Text only | Unsupported | Unsupported |
"Supported" means both speech and text output. "Text only" means text output without speech.
Qwen3.5-Livetranslate supports 60 languages in total: 29 with both audio and text output, and 31 with text-only output.
Qwen3.8-Omni-Flash and Qwen3.5-Omni support 113 input languages and dialects. For the full input-language list, see model selection.
The legacy qwen-omni-turbo supports only Chinese and English.
Recommended models
Model | API | Use cases |
|---|---|---|
qwen3.8-omni-flash | Chat Completions / Responses | Audio/video understanding, text generation, thinking, function calling, web search |
qwen3.5-omni-plus-realtime / qwen3.5-omni-flash-realtime | Realtime API (WebSocket) | Realtime audio/video conversation |
qwen3.5-omni-plus / qwen3.5-omni-flash | Chat Completions | Offline speech output and voice cloning |
qwen3.5-livetranslate-flash-realtime | Realtime API (WebSocket) | Realtime translation |
qwen3-livetranslate-flash | Chat Completions | Audio/video file translation |
qwen-audio-3.0-realtime-plus / qwen-audio-3.0-realtime-flash | Realtime API (WebSocket) | Realtime voice conversation |
All models
Qwen3.8-Omni
qwen3.8-omni-flash accepts text, image, audio, and video input and produces text only through Chat Completions or Responses. For audio/video understanding and content analysis, see Qwen3.8-Omni-Flash.
| Model ID | API | Input | Output | Function Calling | Web search | Thinking mode |
|---|---|---|---|---|---|---|
qwen3.8-omni-flash | Chat Completions / Responses | Text, audio, images, video | Text | Supported | Supported | Enabled by default |
Supports a 1M-token context window, multichannel spatial audio, implicit caching, and Responses Session caching. See model details for supported regions, limits, and capability guides.
Qwen3.5-Omni
Model ID | API | Input | Function calling | Web search | Thinking mode |
|---|---|---|---|---|---|
| Realtime API (WebSocket) | Text, audio, images, video | Supported | Supported | Unsupported |
| Realtime API (WebSocket) | Text, audio, images, video | Supported | Supported | Unsupported |
| Chat Completions | Text, audio, images, video | Supported (Beijing, text output) | Supported | Unsupported |
| Chat Completions | Text, audio, images, video | Supported (Beijing, text output) | Supported | Unsupported |
| Realtime API (WebSocket) | Text, audio, images, video | Supported | Supported | Unsupported |
| Realtime API (WebSocket) | Text, audio, images, video | Supported | Supported | Unsupported |
| Chat Completions | Text, audio, images, video | Supported (Beijing, text output) | Supported | Unsupported |
| Chat Completions | Text, audio, images, video | Supported (Beijing, text output) | Supported | Unsupported |
Qwen3-Omni
Model ID | API | Input | Function calling | Web search | Thinking mode |
|---|---|---|---|---|---|
| Realtime API (WebSocket) | Text, audio, images, video | Unsupported | Unsupported | Unsupported |
| Realtime API (WebSocket) | Text, audio, images, video | Unsupported | Unsupported | Unsupported |
| Realtime API (WebSocket) | Text, audio, images, video | Unsupported | Unsupported | Unsupported |
| Chat Completions | Text, audio, images, video | Supported | Unsupported | Supported |
| Chat Completions | Text, audio, images, video | Supported | Unsupported | Supported |
| Chat Completions | Text, audio, images, video | Supported | Unsupported | Supported |
Qwen3.5-Livetranslate
Model ID | API | Input | Languages |
|---|---|---|---|
| Realtime API (WebSocket) | Audio, images | 60 |
| Realtime API (WebSocket) | Audio | 60 |
Qwen3-Livetranslate
Model ID | API | Input | Languages |
|---|---|---|---|
| Realtime API (WebSocket) | Audio | 18 |
| Realtime API (WebSocket) | Audio | 18 |
| Chat Completions | Audio, video | 18 |
| Chat Completions | Audio, video | 18 |
Qwen-Audio
Model ID | API | Input | Function calling | Web search | Thinking mode |
|---|---|---|---|---|---|
| Realtime API (WebSocket) | Audio, text | Supported | Unsupported | Unsupported |
| Realtime API (WebSocket) | Audio, text | Supported | Unsupported | Unsupported |
Legacy models
The following models are no longer updated. For new projects, use Qwen3.8-Omni-Flash for audio/video understanding and text generation, or Qwen3.5-Omni for offline speech output.
Model ID | Input | API |
|---|---|---|
| Text, audio, images, video | Chat Completions |
| Text, audio, images, video | Chat Completions |
| Text, audio, images, video | Chat Completions |
| Text, audio, images, video | Chat Completions |
| Text, audio | Realtime API (WebSocket) |
| Text, audio | Realtime API (WebSocket) |
| Text, audio | Realtime API (WebSocket) |