qwen3-omni-flash

Updated at:

Qwen3-Omni-Flash multimodal large-scale model, based on the Thinker–Talker Mixed Expert (MoE) architecture, supports efficient understanding and speech generation of text, images, audio, and video. It can interact with text in 119 languages ​​and speech in 20 languages, generating human-like speech for precise cross-lingual communication. The model boasts powerful command-following and system prompt customization capabilities, flexibly adapting to conversational styles and character settings. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interaction experience.This model version is functionally equivalent to the snapshot model qwen3-omni-flash-2025-12-01.

Inference Service Provider

The inference service provider for qwen3-omni-flash is Alibaba Cloud Model Studio.

Model Capabilities

China (Beijing)

CapabilitySupportCapabilitySupport

Input Modality

Text Image Video Audio

Output Modality

Text Audio

Model Experience

Supported

Function Calling

Supported

Structured Outputs

Unsupported

Web Search

Unsupported

Prefix Completion

Unsupported

Context Caching

Unsupported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Singapore

Scope: International

CapabilitySupportCapabilitySupport

Input Modality

Text Image Audio Video

Output Modality

Text Audio

Model Experience

Supported

Function Calling

Unsupported

Structured Outputs

Unsupported

Web Search

Unsupported

Prefix Completion

Unsupported

Context Caching

Unsupported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Context Limits

ParameterValueParameterValue

Max Input Length

49152

Max Output Length

16384

Context Window

65536

Max Input Length (Thinking Mode)

16384

Max Output Length (Thinking Mode)

16384

Max Chain-of-Thought Length

32768

Pricing

This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.

China (Beijing)

Billing ItemPrice (CNY)Unit

Input: Text

1.8

Per 1M tokens

Input: Audio

15.8

Per 1M tokens

Input: Vision

3.3

Per 1M tokens

Output: Text (When input contains only text)

6.9

Per 1M tokens

Output: Text (When input contains images/audio/video)

12.7

Per 1M tokens

Output: Text&Audio (Output text is not charged)

62.6

Per 1M tokens

Input: Text(Thinking)

1.8

Per 1M tokens

Input: Audio(Thinking)

15.8

Per 1M tokens

Input: Vision(Thinking)

3.3

Per 1M tokens

Output: Text (in thinking mode, when input contains only text)

6.9

Per 1M tokens

Output: Text (in thinking mode, when the input contains images/audio/video)

12.7

Per 1M tokens

Singapore

Scope: International

Billing ItemPrice (CNY)Unit

Input: Text

3.156

Per 1M tokens

Input: Audio

27.962

Per 1M tokens

Input: Vision

5.725

Per 1M tokens

Output: Text (When input contains only text)

12.183

Per 1M tokens

Output: Text (When input contains images/audio/video)

22.458

Per 1M tokens

Output: Text&Audio (Output text is not charged)

110.896

Per 1M tokens

Input: Text(Thinking)

3.156

Per 1M tokens

Input: Audio(Thinking)

27.962

Per 1M tokens

Input: Vision(Thinking)

5.725

Per 1M tokens

Output: Text (in thinking mode, when input contains only text)

12.183

Per 1M tokens

Output: Text (in thinking mode, when the input contains images/audio/video)

22.458

Per 1M tokens

Rate Limits

China (Beijing)

ParameterValue

RPM (Requests Per Minute)

60

TPM (Tokens Per Minute)

100,000

Singapore

Scope: International

ParameterValue

RPM (Requests Per Minute)

60

TPM (Tokens Per Minute)

100,000

Snapshot Versions

qwen3-omni-flash-2025-12-01

Qwen3-Omni-Flash multimodal large-scale model, based on the Thinker-Talker Hybrid Expert (MoE) architecture. It supports efficient understanding and speech generation of text, images, audio, and video, enabling text interaction in 119 languages ​​and voice interaction in 20 languages. It supports 49 voice timbres and generates human-like speech for accurate cross-language communication. The model features powerful command following and system prompt customization capabilities, flexibly adapting to dialogue styles and role settings. It is widely used in text creation, voice assistants, multimedia analysis, and other scenarios, providing a natural and smooth multimodal interactive experience. This version is a snapshot from December 1, 2025.

Inference Service Provider

The inference service provider for qwen3-omni-flash-2025-12-01 is Alibaba Cloud Model Studio.

Model Capabilities

CapabilitySupportCapabilitySupport

Input Modality

Text Image Audio Video

Output Modality

Text Audio

Model Experience

Supported

Function Calling

Supported

Structured Outputs

Unsupported

Web Search

Unsupported

Prefix Completion

Unsupported

Context Caching

Unsupported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Context Limits

ParameterValueParameterValue

Max Input Length

49152

Max Output Length

16384

Context Window

65536

Max Input Length (Thinking Mode)

16384

Max Output Length (Thinking Mode)

16384

Max Chain-of-Thought Length

32768

Pricing

This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.

China (Beijing)

Billing ItemPrice (CNY)Unit

Input: Text

1.8

Per 1M tokens

Input: Audio

15.8

Per 1M tokens

Input: Vision

3.3

Per 1M tokens

Output: Text (When input contains only text)

6.9

Per 1M tokens

Output: Text (When input contains images/audio/video)

12.7

Per 1M tokens

Output: Text&Audio (Output text is not charged)

62.6

Per 1M tokens

Input: Text(Thinking)

1.8

Per 1M tokens

Input: Audio(Thinking)

15.8

Per 1M tokens

Input: Vision(Thinking)

3.3

Per 1M tokens

Output: Text (in thinking mode, when input contains only text)

6.9

Per 1M tokens

Output: Text (in thinking mode, when the input contains images/audio/video)

12.7

Per 1M tokens

Rate Limits

China (Beijing)

ParameterValue

RPM (Requests Per Minute)

60

TPM (Tokens Per Minute)

100,000

Singapore

Scope: International

ParameterValue

RPM (Requests Per Minute)

60

TPM (Tokens Per Minute)

100,000

qwen3-omni-flash-2025-09-15

Qwen3-Omni-Flash multimodal large-scale model, based on the Thinker–Talker Mixed Expert (MoE) architecture, supports efficient understanding and speech generation of text, images, audio, and video. It can interact with text in 119 languages ​​and speech in 20 languages, generating human-like speech for precise cross-lingual communication. The model boasts powerful command-following and system prompt customization capabilities, flexibly adapting to conversational styles and character settings. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interaction experience.This version is a snapshot version from September 15, 2025.

Inference Service Provider

The inference service provider for qwen3-omni-flash-2025-09-15 is Alibaba Cloud Model Studio.

Model Capabilities

CapabilitySupportCapabilitySupport

Input Modality

Text Image Video Audio

Output Modality

Text Audio

Model Experience

Supported

Function Calling

Supported

Structured Outputs

Unsupported

Web Search

Unsupported

Prefix Completion

Unsupported

Context Caching

Unsupported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Context Limits

ParameterValueParameterValue

Max Input Length

49152

Max Output Length

16384

Context Window

65536

Max Input Length (Thinking Mode)

16384

Max Output Length (Thinking Mode)

16384

Max Chain-of-Thought Length

32768

Pricing

This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.

China (Beijing)

Billing ItemPrice (CNY)Unit

Input: Text

1.8

Per 1M tokens

Input: Audio

15.8

Per 1M tokens

Input: Vision

3.3

Per 1M tokens

Output: Text (When input contains only text)

6.9

Per 1M tokens

Output: Text (When input contains images/audio/video)

12.7

Per 1M tokens

Output: Text&Audio (Output text is not charged)

62.6

Per 1M tokens

Input: Text(Thinking)

1.8

Per 1M tokens

Input: Audio(Thinking)

15.8

Per 1M tokens

Input: Vision(Thinking)

3.3

Per 1M tokens

Output: Text (in thinking mode, when input contains only text)

6.9

Per 1M tokens

Output: Text (in thinking mode, when the input contains images/audio/video)

12.7

Per 1M tokens

Singapore

Scope: International

Billing ItemPrice (CNY)Unit

Input: Text

3.156

Per 1M tokens

Input: Audio

27.962

Per 1M tokens

Input: Vision

5.725

Per 1M tokens

Output: Text (When input contains only text)

12.183

Per 1M tokens

Output: Text (When input contains images/audio/video)

22.458

Per 1M tokens

Output: Text&Audio (Output text is not charged)

110.896

Per 1M tokens

Input: Text(Thinking)

3.156

Per 1M tokens

Input: Audio(Thinking)

27.962

Per 1M tokens

Input: Vision(Thinking)

5.725

Per 1M tokens

Output: Text (in thinking mode, when input contains only text)

12.183

Per 1M tokens

Output: Text (in thinking mode, when the input contains images/audio/video)

22.458

Per 1M tokens

Rate Limits

China (Beijing)

ParameterValue

RPM (Requests Per Minute)

60

TPM (Tokens Per Minute)

100,000

Singapore

Scope: International

ParameterValue

RPM (Requests Per Minute)

60

TPM (Tokens Per Minute)

100,000