qwen3-omni-flash
Qwen3-Omni-Flash multimodal large-scale model, based on the Thinker–Talker Mixed Expert (MoE) architecture, supports efficient understanding and speech generation of text, images, audio, and video. It can interact with text in 119 languages and speech in 20 languages, generating human-like speech for precise cross-lingual communication. The model boasts powerful command-following and system prompt customization capabilities, flexibly adapting to conversational styles and character settings. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interaction experience.This model version is functionally equivalent to the snapshot model qwen3-omni-flash-2025-12-01.
Inference Service Provider
The inference service provider for qwen3-omni-flash is Alibaba Cloud Model Studio.
Model Capabilities
China (Beijing)
| Capability | Support | Capability | Support |
|---|---|---|---|
Input Modality | Text Image Video Audio | Output Modality | Text Audio |
Model Experience | Supported | Function Calling | Supported |
Structured Outputs | Unsupported | Web Search | Unsupported |
Prefix Completion | Unsupported | Context Caching | Unsupported |
Batch Inference | Unsupported | Fine-tuning | Unsupported |
Singapore
Scope: International
| Capability | Support | Capability | Support |
|---|---|---|---|
Input Modality | Text Image Audio Video | Output Modality | Text Audio |
Model Experience | Supported | Function Calling | Unsupported |
Structured Outputs | Unsupported | Web Search | Unsupported |
Prefix Completion | Unsupported | Context Caching | Unsupported |
Batch Inference | Unsupported | Fine-tuning | Unsupported |
Context Limits
| Parameter | Value | Parameter | Value |
|---|---|---|---|
Max Input Length | 49152 | Max Output Length | 16384 |
Context Window | 65536 | Max Input Length (Thinking Mode) | 16384 |
Max Output Length (Thinking Mode) | 16384 | Max Chain-of-Thought Length | 32768 |
Pricing
This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.
China (Beijing)
| Billing Item | Price (CNY) | Unit |
|---|---|---|
Input: Text | 1.8 | Per 1M tokens |
Input: Audio | 15.8 | Per 1M tokens |
Input: Vision | 3.3 | Per 1M tokens |
Output: Text (When input contains only text) | 6.9 | Per 1M tokens |
Output: Text (When input contains images/audio/video) | 12.7 | Per 1M tokens |
Output: Text&Audio (Output text is not charged) | 62.6 | Per 1M tokens |
Input: Text(Thinking) | 1.8 | Per 1M tokens |
Input: Audio(Thinking) | 15.8 | Per 1M tokens |
Input: Vision(Thinking) | 3.3 | Per 1M tokens |
Output: Text (in thinking mode, when input contains only text) | 6.9 | Per 1M tokens |
Output: Text (in thinking mode, when the input contains images/audio/video) | 12.7 | Per 1M tokens |
Singapore
Scope: International
| Billing Item | Price (CNY) | Unit |
|---|---|---|
Input: Text | 3.156 | Per 1M tokens |
Input: Audio | 27.962 | Per 1M tokens |
Input: Vision | 5.725 | Per 1M tokens |
Output: Text (When input contains only text) | 12.183 | Per 1M tokens |
Output: Text (When input contains images/audio/video) | 22.458 | Per 1M tokens |
Output: Text&Audio (Output text is not charged) | 110.896 | Per 1M tokens |
Input: Text(Thinking) | 3.156 | Per 1M tokens |
Input: Audio(Thinking) | 27.962 | Per 1M tokens |
Input: Vision(Thinking) | 5.725 | Per 1M tokens |
Output: Text (in thinking mode, when input contains only text) | 12.183 | Per 1M tokens |
Output: Text (in thinking mode, when the input contains images/audio/video) | 22.458 | Per 1M tokens |
Rate Limits
China (Beijing)
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) | 60 |
TPM (Tokens Per Minute) | 100,000 |
Singapore
Scope: International
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) | 60 |
TPM (Tokens Per Minute) | 100,000 |
Snapshot Versions
qwen3-omni-flash-2025-12-01
Qwen3-Omni-Flash multimodal large-scale model, based on the Thinker-Talker Hybrid Expert (MoE) architecture. It supports efficient understanding and speech generation of text, images, audio, and video, enabling text interaction in 119 languages and voice interaction in 20 languages. It supports 49 voice timbres and generates human-like speech for accurate cross-language communication. The model features powerful command following and system prompt customization capabilities, flexibly adapting to dialogue styles and role settings. It is widely used in text creation, voice assistants, multimedia analysis, and other scenarios, providing a natural and smooth multimodal interactive experience. This version is a snapshot from December 1, 2025.
Inference Service Provider
The inference service provider for qwen3-omni-flash-2025-12-01 is Alibaba Cloud Model Studio.
Model Capabilities
| Capability | Support | Capability | Support |
|---|---|---|---|
Input Modality | Text Image Audio Video | Output Modality | Text Audio |
Model Experience | Supported | Function Calling | Supported |
Structured Outputs | Unsupported | Web Search | Unsupported |
Prefix Completion | Unsupported | Context Caching | Unsupported |
Batch Inference | Unsupported | Fine-tuning | Unsupported |
Context Limits
| Parameter | Value | Parameter | Value |
|---|---|---|---|
Max Input Length | 49152 | Max Output Length | 16384 |
Context Window | 65536 | Max Input Length (Thinking Mode) | 16384 |
Max Output Length (Thinking Mode) | 16384 | Max Chain-of-Thought Length | 32768 |
Pricing
This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.
China (Beijing)
| Billing Item | Price (CNY) | Unit |
|---|---|---|
Input: Text | 1.8 | Per 1M tokens |
Input: Audio | 15.8 | Per 1M tokens |
Input: Vision | 3.3 | Per 1M tokens |
Output: Text (When input contains only text) | 6.9 | Per 1M tokens |
Output: Text (When input contains images/audio/video) | 12.7 | Per 1M tokens |
Output: Text&Audio (Output text is not charged) | 62.6 | Per 1M tokens |
Input: Text(Thinking) | 1.8 | Per 1M tokens |
Input: Audio(Thinking) | 15.8 | Per 1M tokens |
Input: Vision(Thinking) | 3.3 | Per 1M tokens |
Output: Text (in thinking mode, when input contains only text) | 6.9 | Per 1M tokens |
Output: Text (in thinking mode, when the input contains images/audio/video) | 12.7 | Per 1M tokens |
Rate Limits
China (Beijing)
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) | 60 |
TPM (Tokens Per Minute) | 100,000 |
Singapore
Scope: International
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) | 60 |
TPM (Tokens Per Minute) | 100,000 |
qwen3-omni-flash-2025-09-15
Qwen3-Omni-Flash multimodal large-scale model, based on the Thinker–Talker Mixed Expert (MoE) architecture, supports efficient understanding and speech generation of text, images, audio, and video. It can interact with text in 119 languages and speech in 20 languages, generating human-like speech for precise cross-lingual communication. The model boasts powerful command-following and system prompt customization capabilities, flexibly adapting to conversational styles and character settings. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interaction experience.This version is a snapshot version from September 15, 2025.
Inference Service Provider
The inference service provider for qwen3-omni-flash-2025-09-15 is Alibaba Cloud Model Studio.
Model Capabilities
| Capability | Support | Capability | Support |
|---|---|---|---|
Input Modality | Text Image Video Audio | Output Modality | Text Audio |
Model Experience | Supported | Function Calling | Supported |
Structured Outputs | Unsupported | Web Search | Unsupported |
Prefix Completion | Unsupported | Context Caching | Unsupported |
Batch Inference | Unsupported | Fine-tuning | Unsupported |
Context Limits
| Parameter | Value | Parameter | Value |
|---|---|---|---|
Max Input Length | 49152 | Max Output Length | 16384 |
Context Window | 65536 | Max Input Length (Thinking Mode) | 16384 |
Max Output Length (Thinking Mode) | 16384 | Max Chain-of-Thought Length | 32768 |
Pricing
This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.
China (Beijing)
| Billing Item | Price (CNY) | Unit |
|---|---|---|
Input: Text | 1.8 | Per 1M tokens |
Input: Audio | 15.8 | Per 1M tokens |
Input: Vision | 3.3 | Per 1M tokens |
Output: Text (When input contains only text) | 6.9 | Per 1M tokens |
Output: Text (When input contains images/audio/video) | 12.7 | Per 1M tokens |
Output: Text&Audio (Output text is not charged) | 62.6 | Per 1M tokens |
Input: Text(Thinking) | 1.8 | Per 1M tokens |
Input: Audio(Thinking) | 15.8 | Per 1M tokens |
Input: Vision(Thinking) | 3.3 | Per 1M tokens |
Output: Text (in thinking mode, when input contains only text) | 6.9 | Per 1M tokens |
Output: Text (in thinking mode, when the input contains images/audio/video) | 12.7 | Per 1M tokens |
Singapore
Scope: International
| Billing Item | Price (CNY) | Unit |
|---|---|---|
Input: Text | 3.156 | Per 1M tokens |
Input: Audio | 27.962 | Per 1M tokens |
Input: Vision | 5.725 | Per 1M tokens |
Output: Text (When input contains only text) | 12.183 | Per 1M tokens |
Output: Text (When input contains images/audio/video) | 22.458 | Per 1M tokens |
Output: Text&Audio (Output text is not charged) | 110.896 | Per 1M tokens |
Input: Text(Thinking) | 3.156 | Per 1M tokens |
Input: Audio(Thinking) | 27.962 | Per 1M tokens |
Input: Vision(Thinking) | 5.725 | Per 1M tokens |
Output: Text (in thinking mode, when input contains only text) | 12.183 | Per 1M tokens |
Output: Text (in thinking mode, when the input contains images/audio/video) | 22.458 | Per 1M tokens |
Rate Limits
China (Beijing)
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) | 60 |
TPM (Tokens Per Minute) | 100,000 |
Singapore
Scope: International
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) | 60 |
TPM (Tokens Per Minute) | 100,000 |