qwen3-vl-plus

Updated at:

The Qwen3 series VL models effectively integrates thinking and non-thinking modes, achieving world-leading performance in visual agent capabilities on public benchmark datasets such as OS World. This version features comprehensive upgrades in areas like visual coding, spatial perception, and multimodal reasoning, significantly enhancing visual perception and recognition abilities, and supporting the understanding of ultra-long videos.This model version is functionally equivalent to the snapshot model qwen3-vl-plus-2025-12-19.

Inference Service Provider

The inference service provider for qwen3-vl-plus is Alibaba Cloud Model Studio.

Model Capabilities

China (Beijing)

Capability Support Capability Support

Input Modality

Text Image Video

Output Modality

Text

Model Experience

Supported

Function Calling

Supported

Structured Outputs

Supported

Web Search

Unsupported

Prefix Completion

Supported

Context Caching

Supported

Batch Inference

Supported

Fine-tuning

Unsupported

Singapore

Scope: International

Capability Support Capability Support

Input Modality

Text Image Video

Output Modality

Text

Model Experience

Supported

Function Calling

Unsupported

Structured Outputs

Supported

Web Search

Unsupported

Prefix Completion

Supported

Context Caching

Supported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Germany (Frankfurt)

Scope: EU, Global

Capability Support Capability Support

Input Modality

Text Image Video

Output Modality

Text

Model Experience

Supported

Function Calling

Unsupported

Structured Outputs

Supported

Web Search

Unsupported

Prefix Completion

Supported

Context Caching

Supported

Batch Inference

Unsupported

Fine-tuning

Unsupported

US (Virginia)

Scope: Global

Capability Support Capability Support

Input Modality

Text Image Video

Output Modality

Text

Model Experience

Supported

Function Calling

Unsupported

Structured Outputs

Supported

Web Search

Unsupported

Prefix Completion

Supported

Context Caching

Supported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Context Limits

Parameter Value Parameter Value

Max Input Length

260096

Max Output Length

32768

Context Window

262144

Max Input Length (Thinking Mode)

258048

Max Output Length (Thinking Mode)

32768

Max Chain-of-Thought Length

81920

Pricing

This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.

China (Beijing)

Input<=32k

Billing Item Price (CNY) Unit

Input

1

Per 1M tokens

Output

10

Per 1M tokens

Input(Implicit Cache)

0.2

Per 1M tokens

Input(Batch File)

0.5

Per 1M tokens

Output(Batch File)

5

Per 1M tokens

Explicit Cache Creation

1.25

Per 1M tokens

Explicit Cache Read

0.1

Per 1M tokens

Input(Batch Chat)

1

Per 1M tokens

Output(Batch Chat)

10

Per 1M tokens

32k<Input<=128k

Billing Item Price (CNY) Unit

Input

1.5

Per 1M tokens

Output

15

Per 1M tokens

Input(Implicit Cache)

0.3

Per 1M tokens

Input(Batch File)

0.75

Per 1M tokens

Output(Batch File)

7.5

Per 1M tokens

Explicit Cache Creation

1.875

Per 1M tokens

Explicit Cache Read

0.15

Per 1M tokens

Input(Batch Chat)

1.5

Per 1M tokens

Output(Batch Chat)

15

Per 1M tokens

128k<Input<=256k

Billing Item Price (CNY) Unit

Input

3

Per 1M tokens

Output

30

Per 1M tokens

Input(Implicit Cache)

0.6

Per 1M tokens

Input(Batch File)

1.5

Per 1M tokens

Output(Batch File)

15

Per 1M tokens

Explicit Cache Creation

3.75

Per 1M tokens

Explicit Cache Read

0.3

Per 1M tokens

Input(Batch Chat)

3

Per 1M tokens

Output(Batch Chat)

30

Per 1M tokens

Singapore

Scope: International

Input<=32k

Billing Item Price (CNY) Unit

Input

1.468

Per 1M tokens

Output

11.743

Per 1M tokens

Input(Implicit Cache)

0.294

Per 1M tokens

Explicit Cache Creation

1.835

Per 1M tokens

Explicit Cache Read

0.147

Per 1M tokens

32k<Input<=128k

Billing Item Price (CNY) Unit

Input

2.202

Per 1M tokens

Output

17.614

Per 1M tokens

Input(Implicit Cache)

0.44

Per 1M tokens

Explicit Cache Creation

2.753

Per 1M tokens

Explicit Cache Read

0.22

Per 1M tokens

128k<Input<=256k

Billing Item Price (CNY) Unit

Input

4.404

Per 1M tokens

Output

35.228

Per 1M tokens

Input(Implicit Cache)

0.881

Per 1M tokens

Explicit Cache Creation

5.505

Per 1M tokens

Explicit Cache Read

0.44

Per 1M tokens

Germany (Frankfurt)

Scope: EU

Input<=32k

Billing Item Price (CNY) Unit

Input

1.499

Per 1M tokens

Output

11.991

Per 1M tokens

Input(Implicit Cache)

0.3

Per 1M tokens

Explicit Cache Creation

1.874

Per 1M tokens

Explicit Cache Read

0.15

Per 1M tokens

32k<Input<=128k

Billing Item Price (CNY) Unit

Input

2.248

Per 1M tokens

Output

17.986

Per 1M tokens

Input(Implicit Cache)

0.45

Per 1M tokens

Explicit Cache Creation

2.81

Per 1M tokens

Explicit Cache Read

0.225

Per 1M tokens

128k<Input<=256k

Billing Item Price (CNY) Unit

Input

4.497

Per 1M tokens

Output

35.972

Per 1M tokens

Input(Implicit Cache)

0.899

Per 1M tokens

Explicit Cache Creation

5.621

Per 1M tokens

Explicit Cache Read

0.45

Per 1M tokens

Scope: Global

Input<=32k

Billing Item Price (CNY) Unit

Input

1

Per 1M tokens

Output

10

Per 1M tokens

Input(Implicit Cache)

0.2

Per 1M tokens

Explicit Cache Creation

1.25

Per 1M tokens

Explicit Cache Read

0.1

Per 1M tokens

32k<Input<=128k

Billing Item Price (CNY) Unit

Input

1.5

Per 1M tokens

Output

15

Per 1M tokens

Input(Implicit Cache)

0.3

Per 1M tokens

Explicit Cache Creation

1.875

Per 1M tokens

Explicit Cache Read

0.15

Per 1M tokens

128k<Input<=256k

Billing Item Price (CNY) Unit

Input

3

Per 1M tokens

Output

30

Per 1M tokens

Input(Implicit Cache)

0.6

Per 1M tokens

Explicit Cache Creation

3.75

Per 1M tokens

Explicit Cache Read

0.3

Per 1M tokens

US (Virginia)

Scope: Global

Input<=32k

Billing Item Price (CNY) Unit

Input

1

Per 1M tokens

Output

10

Per 1M tokens

Input(Implicit Cache)

0.2

Per 1M tokens

Explicit Cache Creation

1.25

Per 1M tokens

Explicit Cache Read

0.1

Per 1M tokens

32k<Input<=128k

Billing Item Price (CNY) Unit

Input

1.5

Per 1M tokens

Output

15

Per 1M tokens

Input(Implicit Cache)

0.3

Per 1M tokens

Explicit Cache Creation

1.875

Per 1M tokens

Explicit Cache Read

0.15

Per 1M tokens

128k<Input<=256k

Billing Item Price (CNY) Unit

Input

3

Per 1M tokens

Output

30

Per 1M tokens

Input(Implicit Cache)

0.6

Per 1M tokens

Explicit Cache Creation

3.75

Per 1M tokens

Explicit Cache Read

0.3

Per 1M tokens

Rate Limits

China (Beijing)

Parameter Value

RPM (Requests Per Minute)

3000

TPM (Tokens Per Minute)

5,000,000

Singapore

Scope: International

Parameter Value

RPM (Requests Per Minute)

1200

TPM (Tokens Per Minute)

1,000,000

Germany (Frankfurt)

Scope: EU

Parameter Value

RPM (Requests Per Minute)

2000

TPM (Tokens Per Minute)

1,000,000

Scope: Global

Parameter Value

RPM (Requests Per Minute)

2000

TPM (Tokens Per Minute)

1,000,000

US (Virginia)

Scope: Global

Parameter Value

RPM (Requests Per Minute)

6000

TPM (Tokens Per Minute)

10,000,000

Snapshot Versions

qwen3-vl-plus-2025-12-19

The Qwen3 series of visual understanding models effectively integrates thinking and non-thinking modes. Compared to the snapshot released on September 23, this version delivers superior performance in reasoning and analysis tasks as well as style control, while also offering lower latency and faster response speeds. This version is based on a snapshot taken on December 19, 2025.

Inference Service Provider

The inference service provider for qwen3-vl-plus-2025-12-19 is Alibaba Cloud Model Studio.

Model Capabilities

China (Beijing)

Capability Support Capability Support

Input Modality

Text Image Video

Output Modality

Text

Model Experience

Supported

Function Calling

Supported

Structured Outputs

Supported

Web Search

Unsupported

Prefix Completion

Supported

Context Caching

Unsupported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Singapore

Scope: International

Capability Support Capability Support

Input Modality

Text Image Video

Output Modality

Text

Model Experience

Supported

Function Calling

Unsupported

Structured Outputs

Supported

Web Search

Unsupported

Prefix Completion

Supported

Context Caching

Unsupported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Context Limits

Parameter Value Parameter Value

Max Input Length

260096

Max Output Length

32768

Context Window

262144

Max Input Length (Thinking Mode)

258048

Max Output Length (Thinking Mode)

32768

Max Chain-of-Thought Length

81920

Pricing

This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.

China (Beijing)

Input<=32k

Billing Item Price (CNY) Unit

Input

1

Per 1M tokens

Output

10

Per 1M tokens

32k<Input<=128k

Billing Item Price (CNY) Unit

Input

1.5

Per 1M tokens

Output

15

Per 1M tokens

128k<Input<=256k

Billing Item Price (CNY) Unit

Input

3

Per 1M tokens

Output

30

Per 1M tokens

Singapore

Scope: International

Input<=32k

Billing Item Price (CNY) Unit

Input

1.468

Per 1M tokens

Output

11.743

Per 1M tokens

32k<Input<=128k

Billing Item Price (CNY) Unit

Input

2.202

Per 1M tokens

Output

17.614

Per 1M tokens

128k<Input<=256k

Billing Item Price (CNY) Unit

Input

4.404

Per 1M tokens

Output

35.228

Per 1M tokens

Rate Limits

China (Beijing)

Parameter Value

RPM (Requests Per Minute)

60

TPM (Tokens Per Minute)

100,000

Singapore

Scope: International

Parameter Value

RPM (Requests Per Minute)

60

TPM (Tokens Per Minute)

100,000

qwen3-vl-plus-2025-09-23

The Qwen3 series VL models effectively integrates thinking and non-thinking modes, achieving world-leading performance in visual agent capabilities on public benchmark datasets such as OS World. This version features comprehensive upgrades in areas like visual coding, spatial perception, and multimodal reasoning, significantly enhancing visual perception and recognition abilities, and supporting the understanding of ultra-long videos.This version is a snapshot as of September 23, 2025

Inference Service Provider

The inference service provider for qwen3-vl-plus-2025-09-23 is Alibaba Cloud Model Studio.

Model Capabilities

China (Beijing)

Capability Support Capability Support

Input Modality

Text Image Video

Output Modality

Text

Model Experience

Supported

Function Calling

Supported

Structured Outputs

Supported

Web Search

Unsupported

Prefix Completion

Supported

Context Caching

Unsupported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Singapore

Scope: International

Capability Support Capability Support

Input Modality

Text Image Video

Output Modality

Text

Model Experience

Supported

Function Calling

Unsupported

Structured Outputs

Supported

Web Search

Unsupported

Prefix Completion

Supported

Context Caching

Unsupported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Germany (Frankfurt)

Scope: Global

Capability Support Capability Support

Input Modality

Text Image Video

Output Modality

Text

Model Experience

Supported

Function Calling

Unsupported

Structured Outputs

Supported

Web Search

Unsupported

Prefix Completion

Supported

Context Caching

Unsupported

Batch Inference

Unsupported

Fine-tuning

Unsupported

US (Virginia)

Scope: Global

Capability Support Capability Support

Input Modality

Text Image Video

Output Modality

Text

Model Experience

Supported

Function Calling

Unsupported

Structured Outputs

Supported

Web Search

Unsupported

Prefix Completion

Supported

Context Caching

Unsupported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Context Limits

Parameter Value Parameter Value

Max Input Length

260096

Max Output Length

32768

Context Window

262144

Max Input Length (Thinking Mode)

258048

Max Output Length (Thinking Mode)

32768

Max Chain-of-Thought Length

81920

Pricing

This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.

China (Beijing)

Input<=32k

Billing Item Price (CNY) Unit

Input

1

Per 1M tokens

Output

10

Per 1M tokens

32k<Input<=128k

Billing Item Price (CNY) Unit

Input

1.5

Per 1M tokens

Output

15

Per 1M tokens

128k<Input<=256k

Billing Item Price (CNY) Unit

Input

3

Per 1M tokens

Output

30

Per 1M tokens

Singapore

Scope: International

Input<=32k

Billing Item Price (CNY) Unit

Input

1.468

Per 1M tokens

Output

11.743

Per 1M tokens

32k<Input<=128k

Billing Item Price (CNY) Unit

Input

2.202

Per 1M tokens

Output

17.614

Per 1M tokens

128k<Input<=256k

Billing Item Price (CNY) Unit

Input

4.404

Per 1M tokens

Output

35.228

Per 1M tokens

Germany (Frankfurt)

Scope: Global

Input<=32k

Billing Item Price (CNY) Unit

Input

1

Per 1M tokens

Output

10

Per 1M tokens

32k<Input<=128k

Billing Item Price (CNY) Unit

Input

1.5

Per 1M tokens

Output

15

Per 1M tokens

128k<Input<=256k

Billing Item Price (CNY) Unit

Input

3

Per 1M tokens

Output

30

Per 1M tokens

US (Virginia)

Scope: Global

Input<=32k

Billing Item Price (CNY) Unit

Input

1

Per 1M tokens

Output

10

Per 1M tokens

32k<Input<=128k

Billing Item Price (CNY) Unit

Input

1.5

Per 1M tokens

Output

15

Per 1M tokens

128k<Input<=256k

Billing Item Price (CNY) Unit

Input

3

Per 1M tokens

Output

30

Per 1M tokens

Rate Limits

China (Beijing)

Parameter Value

RPM (Requests Per Minute)

60

TPM (Tokens Per Minute)

100,000

Singapore

Scope: International

Parameter Value

RPM (Requests Per Minute)

120

TPM (Tokens Per Minute)

1,000,000

Germany (Frankfurt)

Scope: Global

Parameter Value

RPM (Requests Per Minute)

60

TPM (Tokens Per Minute)

100,000

US (Virginia)

Scope: Global

Parameter Value

RPM (Requests Per Minute)

60

TPM (Tokens Per Minute)

100,000