Omni-modal overview

Updated at:

Choose a model for voice conversation, audio and video analysis, content moderation, voice translation, and other omni-modal tasks.

Migrate from closed-source models

Map your current GPT or Gemini omni-modal/real-time capabilities to an equivalent Bailian model.

Closed-source examples

Bailian recommendation

Audio/video understanding and text generation

Gemini 3.8 Flash

qwen3.8-omni-flash

Real-time translation

Gemini 3.5 Live Translate

qwen3.5-livetranslate-flash-realtime

NoteFor speech output, use qwen3.5-omni-plus.

Use cases

Omni-modal models understand text, audio, images, and video. Qwen3.8-Omni-Flash supports audio/video analysis, meeting summaries, and subtitle generation, with thinking mode, tool calling, and web search. Choose a model for your use case:

Use case

Recommended model

User guide

Audio/video understanding and text generation: Analyze audio/video content and generate meeting summaries, subtitles, and text answers

Qwen3.8-Omni-Flash (Chat Completions / Responses)

Qwen3.8-Omni-Flash guide

Real-time voice/video conversation: Interact with AI through a microphone and camera (voice assistants, customer service bots, visual Q&A, live-stream analysis)

Qwen3.5-Omni Realtime (WebSocket)

Qwen-Omni-Realtime

Offline audio output: Upload audio or video files and generate speech responses

Qwen3.5-Omni (Chat Completions)

Non-real-time (Qwen-Omni)

Real-time voice translation: Simultaneous interpretation with approximately 3-second latency, supporting 60 languages (live interpretation, multilingual meetings)

Qwen3.5-Livetranslate (WebSocket)

Real-time audio and video translation - Qwen

Audio and video file translation: Upload audio or video files and translate them into a target language (video dubbing, podcast translation)

Qwen3-Livetranslate (Chat Completions)

Audio and video file translation (Qwen)

Voice cloning: Provide a reference audio clip and the AI generates speech responses in that voice

Qwen3.5-Omni Plus / Flash (Chat Completions / Realtime API)

Voice cloning

Real-time voice conversation with semantic VAD: End-to-end voice interaction, semantic turn detection (smart_turn), non-meaningful utterances won't interrupt the conversation, supports Function Calling (voice assistants, smart customer service)

Qwen-Audio (WebSocket)

Realtime Audio Chat (Qwen-Audio-Realtime)

  • When using Qwen3.5-Omni for content analysis, it supports audio up to 3 hours and video up to 1 hour per request.
  • Qwen3.8-Omni-Flash (Chat Completions and Responses), Qwen3.5-Omni Plus / Flash (HTTP, text output), Qwen3-Omni-Flash (HTTP only), and Qwen-Audio Realtime (WebSocket) support function calling.
  • Qwen3.8-Omni-Flash (Chat Completions / Responses) and Qwen3.5-Omni (Chat Completions / Realtime API) support web search. Qwen3.5-Omni cannot enable web search and function calling at the same time.

Translation

For audio/video translation into text or subtitles, use Qwen3.8-Omni-Flash. For spoken translations, choose a model below based on latency and speech output requirements.

NoteQuick setup: Qwen3.5-Livetranslate (60 languages, ~3 s latency, out-of-the-box). Speech output with web search and term injection: Qwen3.5-Omni (29 output languages, web search and term injection).

Supported languages

Language

Qwen3.5-Livetranslate

Qwen3-Livetranslate

Qwen3.5-Omni

Qwen3-Omni-Flash

English

Supported

Supported

Supported

Supported

Chinese (Mandarin)

Supported

Supported

Supported

Supported

Cantonese

Supported

Supported

Supported Text only

Supported

Sichuanese

Supported

Supported

Supported

Supported

Shanghainese

Supported

Supported

Supported

Supported

Beijingese

Supported

Supported

Supported

Supported

Tianjinese

Supported

Supported

Supported

Supported

Nanjingese

Unsupported Text only

Unsupported

Supported

Supported

Shaanxi dialect

Unsupported Text only

Unsupported

Supported

Supported

Hokkien

Unsupported Text only

Unsupported

Supported

Supported

French

Supported

Supported

Supported

Supported

German

Supported

Supported

Supported

Supported

Russian

Supported

Supported

Supported

Supported

Italian

Supported

Supported

Supported

Supported

Spanish

Supported

Supported

Supported

Supported

Portuguese

Supported

Supported

Supported

Supported

Japanese

Supported

Supported

Supported

Supported

Korean

Supported

Supported

Supported

Supported

Thai

Supported

Supported Text only

Supported

Supported

Indonesian

Supported

Supported Text only

Supported

Unsupported

Vietnamese

Supported

Supported Text only

Supported

Unsupported

Arabic

Supported

Supported Text only

Supported

Unsupported

Hindi

Supported

Supported Text only

Supported

Unsupported

Turkish

Supported

Supported Text only

Supported

Unsupported

Finnish

Unsupported Text only

Unsupported

Supported

Unsupported

Polish

Unsupported Text only

Unsupported

Supported

Unsupported

Dutch

Unsupported Text only

Unsupported

Supported

Unsupported

Czech

Unsupported Text only

Unsupported

Supported

Unsupported

Urdu

Unsupported Text only

Unsupported

Supported

Unsupported

Tagalog

Unsupported Text only

Unsupported

Supported

Unsupported

Swedish

Unsupported Text only

Unsupported

Supported

Unsupported

Danish

Unsupported Text only

Unsupported

Supported

Unsupported

Hebrew

Unsupported Text only

Unsupported

Supported

Unsupported

Icelandic

Unsupported Text only

Unsupported

Supported

Unsupported

Malay

Unsupported Text only

Unsupported

Supported

Unsupported

Norwegian

Unsupported Text only

Unsupported

Supported

Unsupported

Persian

Unsupported Text only

Unsupported

Supported

Unsupported

Greek

Supported Text only

Supported Text only

Unsupported

Unsupported

"Supported" means both speech and text output. "Text only" means text output without speech.

Qwen3.5-Livetranslate supports 60 languages in total: 29 with both audio and text output, and 31 with text-only output.

Qwen3.8-Omni-Flash and Qwen3.5-Omni support 113 input languages and dialects. For the full input-language list, see model selection.

The legacy qwen-omni-turbo supports only Chinese and English.

Model

API

Use cases

qwen3.8-omni-flash

Chat Completions / Responses

Audio/video understanding, text generation, thinking, function calling, web search

qwen3.5-omni-plus-realtime / qwen3.5-omni-flash-realtime

Realtime API (WebSocket)

Realtime audio/video conversation

qwen3.5-omni-plus / qwen3.5-omni-flash

Chat Completions

Offline speech output and voice cloning

qwen3.5-livetranslate-flash-realtime

Realtime API (WebSocket)

Realtime translation

qwen3-livetranslate-flash

Chat Completions

Audio/video file translation

qwen-audio-3.0-realtime-plus / qwen-audio-3.0-realtime-flash

Realtime API (WebSocket)

Realtime voice conversation

All models

Qwen3.8-Omni

qwen3.8-omni-flash accepts text, image, audio, and video input and produces text only through Chat Completions or Responses. For audio/video understanding and content analysis, see Qwen3.8-Omni-Flash.

Model IDAPIInputOutputFunction CallingWeb searchThinking mode
qwen3.8-omni-flashChat Completions / ResponsesText, audio, images, videoTextSupportedSupportedEnabled by default

Supports a 1M-token context window, multichannel spatial audio, implicit caching, and Responses Session caching. See model details for supported regions, limits, and capability guides.

Qwen3.5-Omni

Model ID

API

Input

Function calling

Web search

Thinking mode

qwen3.5-omni-plus-realtime

Realtime API (WebSocket)

Text, audio, images, video

Supported

Supported

Unsupported

qwen3.5-omni-plus-realtime-2026-03-15

Realtime API (WebSocket)

Text, audio, images, video

Supported

Supported

Unsupported

qwen3.5-omni-plus

Chat Completions

Text, audio, images, video

Supported (Beijing, text output)

Supported

Unsupported

qwen3.5-omni-plus-2026-03-15

Chat Completions

Text, audio, images, video

Supported (Beijing, text output)

Supported

Unsupported

qwen3.5-omni-flash-realtime

Realtime API (WebSocket)

Text, audio, images, video

Supported

Supported

Unsupported

qwen3.5-omni-flash-realtime-2026-03-15

Realtime API (WebSocket)

Text, audio, images, video

Supported

Supported

Unsupported

qwen3.5-omni-flash

Chat Completions

Text, audio, images, video

Supported (Beijing, text output)

Supported

Unsupported

qwen3.5-omni-flash-2026-03-15

Chat Completions

Text, audio, images, video

Supported (Beijing, text output)

Supported

Unsupported

Qwen3-Omni

Model ID

API

Input

Function calling

Web search

Thinking mode

qwen3-omni-flash-realtime

Realtime API (WebSocket)

Text, audio, images, video

Unsupported

Unsupported

Unsupported

qwen3-omni-flash-realtime-2025-12-01

Realtime API (WebSocket)

Text, audio, images, video

Unsupported

Unsupported

Unsupported

qwen3-omni-flash-realtime-2025-09-15

Realtime API (WebSocket)

Text, audio, images, video

Unsupported

Unsupported

Unsupported

qwen3-omni-flash

Chat Completions

Text, audio, images, video

Supported

Unsupported

Supported

qwen3-omni-flash-2025-12-01

Chat Completions

Text, audio, images, video

Supported

Unsupported

Supported

qwen3-omni-flash-2025-09-15

Chat Completions

Text, audio, images, video

Supported

Unsupported

Supported

Qwen3.5-Livetranslate

Model ID

API

Input

Languages

qwen3.5-livetranslate-flash-realtime

Realtime API (WebSocket)

Audio, images

60

qwen3.5-livetranslate-flash-realtime-2026-05-19

Realtime API (WebSocket)

Audio

60

Qwen3-Livetranslate

Model ID

API

Input

Languages

qwen3-livetranslate-flash-realtime

Realtime API (WebSocket)

Audio

18

qwen3-livetranslate-flash-realtime-2025-09-22

Realtime API (WebSocket)

Audio

18

qwen3-livetranslate-flash

Chat Completions

Audio, video

18

qwen3-livetranslate-flash-2025-12-01

Chat Completions

Audio, video

18

Qwen-Audio

Model ID

API

Input

Function calling

Web search

Thinking mode

qwen-audio-3.0-realtime-plus

Realtime API (WebSocket)

Audio, text

Supported

Unsupported

Unsupported

qwen-audio-3.0-realtime-flash

Realtime API (WebSocket)

Audio, text

Supported

Unsupported

Unsupported

Legacy models

The following models are no longer updated. For new projects, use Qwen3.8-Omni-Flash for audio/video understanding and text generation, or Qwen3.5-Omni for offline speech output.

Model ID

Input

API

qwen2.5-omni-7b

Text, audio, images, video

Chat Completions

qwen-omni-turbo

Text, audio, images, video

Chat Completions

qwen-omni-turbo-latest

Text, audio, images, video

Chat Completions

qwen-omni-turbo-2025-03-26

Text, audio, images, video

Chat Completions

qwen-omni-turbo-realtime

Text, audio

Realtime API (WebSocket)

qwen-omni-turbo-realtime-latest

Text, audio

Realtime API (WebSocket)

qwen-omni-turbo-realtime-2025-05-08

Text, audio

Realtime API (WebSocket)