qwen3.5-omni-flash
Qwen 3.5-Omni is a Qwen multimodal large model, supporting text, image, audio, and audio-visual understanding and interaction. As a comprehensive evolution of Qwen 3-Omni, it supports over 10 hours of audio understanding and over 400 seconds of 720P (1 FPS) audio-visual understanding and dialogue. It further expands the language range, supporting audio input in 60+ languages and speech output in 30+ languages. It also possesses powerful structured audio-visual understanding capabilities and is widely used in text creation, voice assistants, multimedia analysis, and other scenarios, providing a natural and fluent multimodal understanding and interactive experience.This model version is functionally equivalent to the snapshot model qwen3.5-omni-flash-2026-03-15.
Inference Service Provider
The inference service provider for qwen3.5-omni-flash is Alibaba Cloud Model Studio.
Model Capabilities
China (Beijing)
| Capability | Support | Capability | Support |
|---|---|---|---|
Input Modality | Text Image Video Audio | Output Modality | Text Audio |
Model Experience | Unsupported | Function Calling | Supported (Chat Completions, text output) |
Structured Outputs | Unsupported | Web Search | Supported |
Prefix Completion | Unsupported | Context Caching | Unsupported |
Batch Inference | Supported | Fine-tuning | Unsupported |
Singapore
Scope: International
| Capability | Support | Capability | Support |
|---|---|---|---|
Input Modality | Text Image Video Audio | Output Modality | Text Audio |
Model Experience | Unsupported | Function Calling | Unsupported |
Structured Outputs | Unsupported | Web Search | Supported |
Prefix Completion | Unsupported | Context Caching | Unsupported |
Batch Inference | Unsupported | Fine-tuning | Unsupported |
Context Limits
| Parameter | Value | Parameter | Value |
|---|---|---|---|
Max Input Length | 196608 | Max Output Length | 65536 |
Context Window | 262144 |
Pricing
This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.
China (Beijing)
| Billing Item | Price (CNY) | Unit |
|---|---|---|
Input: Audio | 18 | Per 1M tokens |
Output: Text&Audio (Output text is not charged) | 72 | Per 1M tokens |
input:Text/Image/Video | 2.2 | Per 1M tokens |
Output: Text | 13.3 | Per 1M tokens |
Input:Audio(Batch File) | 9 | Per 1M tokens |
Input:Text/Image/Video(Batch File) | 1.1 | Per 1M tokens |
Output:Text(Batch File) | 6.65 | Per 1M tokens |
Input:Audio(Batch Chat) | 18 | Per 1M tokens |
Input:Text/Image/Video(Batch Chat) | 2.2 | Per 1M tokens |
Output:Text(Batch Chat) | 13.3 | Per 1M tokens |
Singapore
Scope: International
| Billing Item | Price (CNY) | Unit |
|---|---|---|
Input: Audio | 22.48 | Per 1M tokens |
Output: Text&Audio (Output text is not charged) | 89.18 | Per 1M tokens |
input:Text/Image/Video | 3 | Per 1M tokens |
Output: Text | 16.49 | Per 1M tokens |
Rate Limits
China (Beijing)
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) | 60 |
TPM (Tokens Per Minute) | 100,000 |
Singapore
Scope: International
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) | 60 |
TPM (Tokens Per Minute) | 100,000 |
Snapshot Versions
qwen3.5-omni-flash-2026-03-15
Qwen 3.5-Omni is a Qwen multimodal large model, supporting text, image, audio, and audio-visual understanding and interaction. As a comprehensive evolution of Qwen 3-Omni, it supports over 10 hours of audio understanding and over 400 seconds of 720P (1 FPS) audio-visual understanding and dialogue. It further expands the language range, supporting audio input in 60+ languages and speech output in 30+ languages. It also possesses powerful structured audio-visual understanding capabilities and is widely used in text creation, voice assistants, multimedia analysis, and other scenarios, providing a natural and fluent multimodal understanding and interactive experience.This version is a snapshot from March 15, 2026.
Inference Service Provider
The inference service provider for qwen3.5-omni-flash-2026-03-15 is Alibaba Cloud Model Studio.
Model Capabilities
| Capability | Support | Capability | Support |
|---|---|---|---|
Input Modality | Text Image Video Audio | Output Modality | Text Audio |
Model Experience | Unsupported | Function Calling | Supported (Beijing, Chat Completions, text output) |
Structured Outputs | Unsupported | Web Search | Supported |
Prefix Completion | Unsupported | Context Caching | Unsupported |
Batch Inference | Unsupported | Fine-tuning | Unsupported |
Context Limits
| Parameter | Value | Parameter | Value |
|---|---|---|---|
Max Input Length | 196608 | Max Output Length | 65536 |
Context Window | 262144 |
Pricing
This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.
China (Beijing)
| Billing Item | Price (CNY) | Unit |
|---|---|---|
Input: Audio | 18 | Per 1M tokens |
Output: Text&Audio (Output text is not charged) | 72 | Per 1M tokens |
input:Text/Image/Video | 2.2 | Per 1M tokens |
Output: Text | 13.3 | Per 1M tokens |
Singapore
Scope: International
| Billing Item | Price (CNY) | Unit |
|---|---|---|
Input: Audio | 22.48 | Per 1M tokens |
Output: Text&Audio (Output text is not charged) | 89.18 | Per 1M tokens |
input:Text/Image/Video | 3 | Per 1M tokens |
Output: Text | 16.49 | Per 1M tokens |
Rate Limits
China (Beijing)
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) | 60 |
TPM (Tokens Per Minute) | 100,000 |
Singapore
Scope: International
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) | 60 |
TPM (Tokens Per Minute) | 100,000 |