qwen3-max

更新时间:
复制 MD 格式

The Qwen 3 series Max model has undergone specialized upgrades in agent programming and tool invocation compared to the preview version. The officially released model this time has achieved state-of-the-art (SOTA) performance in its field and is better suited to meet the demands of agents operating in more complex scenarios.This model version is functionally equivalent to the snapshot model qwen3-max-2026-01-23.

Inference Service Provider

The inference service provider for qwen3-max is Alibaba Cloud Model Studio.

Model Capabilities

China (Beijing)

Capability Support Capability Support

Input Modality

Text

Output Modality

Text

Model Experience

Supported

Function Calling

Supported

Structured Outputs

Supported

Web Search

Supported

Prefix Completion

Supported

Context Caching

Supported

Batch Inference

Supported

Fine-tuning

Unsupported

Singapore

Scope: International

Capability Support Capability Support

Input Modality

Text

Output Modality

Text

Model Experience

Supported

Function Calling

Unsupported

Structured Outputs

Supported

Web Search

Supported

Prefix Completion

Supported

Context Caching

Supported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Germany (Frankfurt)

Scope: EU, Global

Capability Support Capability Support

Input Modality

Text

Output Modality

Text

Model Experience

Supported

Function Calling

Unsupported

Structured Outputs

Supported

Web Search

Unsupported

Prefix Completion

Supported

Context Caching

Supported

Batch Inference

Unsupported

Fine-tuning

Unsupported

US (Virginia)

Scope: Global

Capability Support Capability Support

Input Modality

Text

Output Modality

Text

Model Experience

Supported

Function Calling

Unsupported

Structured Outputs

Supported

Web Search

Supported

Prefix Completion

Supported

Context Caching

Supported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Context Limits

Parameter Value Parameter Value

Max Input Length

258048

Max Output Length

65536

Context Window

262144

Max Input Length (Thinking Mode)

258048

Max Output Length (Thinking Mode)

32768

Max Chain-of-Thought Length

81920

Pricing

This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.

China (Beijing)

Input<=32k

Billing Item Price (CNY) Unit

Input

2.5

Per 1M tokens

Output

10

Per 1M tokens

Input(Implicit Cache)

0.5

Per 1M tokens

Input(Batch File)

1.25

Per 1M tokens

Output(Batch File)

5

Per 1M tokens

Explicit Cache Creation

3.125

Per 1M tokens

Explicit Cache Read

0.25

Per 1M tokens

Input(Batch Chat)

2.5

Per 1M tokens

Output(Batch Chat)

10

Per 1M tokens

32k<Input<=128k

Billing Item Price (CNY) Unit

Input

4

Per 1M tokens

Output

16

Per 1M tokens

Input(Implicit Cache)

0.8

Per 1M tokens

Input(Batch File)

2

Per 1M tokens

Output(Batch File)

8

Per 1M tokens

Explicit Cache Creation

5

Per 1M tokens

Explicit Cache Read

0.4

Per 1M tokens

Input(Batch Chat)

4

Per 1M tokens

Output(Batch Chat)

16

Per 1M tokens

128k<Input<=256k

Billing Item Price (CNY) Unit

Input

7

Per 1M tokens

Output

28

Per 1M tokens

Input(Implicit Cache)

1.4

Per 1M tokens

Input(Batch File)

3.5

Per 1M tokens

Output(Batch File)

14

Per 1M tokens

Explicit Cache Creation

8.75

Per 1M tokens

Explicit Cache Read

0.7

Per 1M tokens

Input(Batch Chat)

7

Per 1M tokens

Output(Batch Chat)

28

Per 1M tokens

Singapore

Scope: International

Input<=32k

Billing Item Price (CNY) Unit

Input

8.807

Per 1M tokens

Output

44.035

Per 1M tokens

Input(Implicit Cache)

1.761

Per 1M tokens

Explicit Cache Creation

11.009

Per 1M tokens

Explicit Cache Read

0.881

Per 1M tokens

32k<Input<=128k

Billing Item Price (CNY) Unit

Input

17.614

Per 1M tokens

Output

88.071

Per 1M tokens

Input(Implicit Cache)

3.523

Per 1M tokens

Explicit Cache Creation

22.018

Per 1M tokens

Explicit Cache Read

1.761

Per 1M tokens

128k<Input<=256k

Billing Item Price (CNY) Unit

Input

22.018

Per 1M tokens

Output

110.089

Per 1M tokens

Input(Implicit Cache)

4.404

Per 1M tokens

Explicit Cache Creation

27.522

Per 1M tokens

Explicit Cache Read

2.202

Per 1M tokens

Germany (Frankfurt)

Scope: EU

Input<=32k

Billing Item Price (CNY) Unit

Input

8.993

Per 1M tokens

Output

44.965

Per 1M tokens

Input(Implicit Cache)

1.799

Per 1M tokens

Explicit Cache Creation

11.241

Per 1M tokens

Explicit Cache Read

0.899

Per 1M tokens

32k<Input<=128k

Billing Item Price (CNY) Unit

Input

17.986

Per 1M tokens

Output

89.93

Per 1M tokens

Input(Implicit Cache)

3.597

Per 1M tokens

Explicit Cache Creation

22.483

Per 1M tokens

Explicit Cache Read

1.799

Per 1M tokens

128k<Input<=256k

Billing Item Price (CNY) Unit

Input

22.483

Per 1M tokens

Output

112.413

Per 1M tokens

Input(Implicit Cache)

4.497

Per 1M tokens

Explicit Cache Creation

28.103

Per 1M tokens

Explicit Cache Read

2.248

Per 1M tokens

Scope: Global

Input<=32k

Billing Item Price (CNY) Unit

Input

2.5

Per 1M tokens

Output

10

Per 1M tokens

Input(Implicit Cache)

0.5

Per 1M tokens

Explicit Cache Creation

3.125

Per 1M tokens

Explicit Cache Read

0.25

Per 1M tokens

32k<Input<=128k

Billing Item Price (CNY) Unit

Input

4

Per 1M tokens

Output

16

Per 1M tokens

Input(Implicit Cache)

0.8

Per 1M tokens

Explicit Cache Creation

5

Per 1M tokens

Explicit Cache Read

0.4

Per 1M tokens

128k<Input<=256k

Billing Item Price (CNY) Unit

Input

7

Per 1M tokens

Output

28

Per 1M tokens

Input(Implicit Cache)

1.4

Per 1M tokens

Explicit Cache Creation

8.75

Per 1M tokens

Explicit Cache Read

0.7

Per 1M tokens

US (Virginia)

Scope: Global

Input<=32k

Billing Item Price (CNY) Unit

Input

2.5

Per 1M tokens

Output

10

Per 1M tokens

Input(Implicit Cache)

0.5

Per 1M tokens

Explicit Cache Creation

3.125

Per 1M tokens

Explicit Cache Read

0.25

Per 1M tokens

32k<Input<=128k

Billing Item Price (CNY) Unit

Input

4

Per 1M tokens

Output

16

Per 1M tokens

Input(Implicit Cache)

0.8

Per 1M tokens

Explicit Cache Creation

5

Per 1M tokens

Explicit Cache Read

0.4

Per 1M tokens

128k<Input<=256k

Billing Item Price (CNY) Unit

Input

7

Per 1M tokens

Output

28

Per 1M tokens

Input(Implicit Cache)

1.4

Per 1M tokens

Explicit Cache Creation

8.75

Per 1M tokens

Explicit Cache Read

0.7

Per 1M tokens

Rate Limits

China (Beijing)

Parameter Value

RPM (Requests Per Minute)

30000

TPM (Tokens Per Minute)

5,000,000

Singapore

Scope: International

Parameter Value

RPM (Requests Per Minute)

600

TPM (Tokens Per Minute)

1,000,000

Germany (Frankfurt)

Scope: EU

Parameter Value

RPM (Requests Per Minute)

600

TPM (Tokens Per Minute)

1,000,000

Scope: Global

Parameter Value

RPM (Requests Per Minute)

600

TPM (Tokens Per Minute)

1,000,000

US (Virginia)

Scope: Global

Parameter Value

RPM (Requests Per Minute)

600

TPM (Tokens Per Minute)

1,000,000

Snapshot Versions

qwen3-max-preview

A preview version of the Max model in the Qwen 3 series, achieving an effective integration of thinking and non-thinking modes. In thinking mode, there is a significant enhancement in capabilities such as intelligent agent programming, common-sense reasoning, and reasoning across mathematics, science, and general domains.

Inference Service Provider

The inference service provider for qwen3-max-preview is Alibaba Cloud Model Studio.

Model Capabilities

China (Beijing)

Capability Support Capability Support

Input Modality

Text

Output Modality

Text

Model Experience

Supported

Function Calling

Unsupported

Structured Outputs

Unsupported

Web Search

Unsupported

Prefix Completion

Unsupported

Context Caching

Supported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Singapore

Scope: International

Capability Support Capability Support

Input Modality

Text

Output Modality

Text

Model Experience

Supported

Function Calling

Unsupported

Structured Outputs

Unsupported

Web Search

Unsupported

Prefix Completion

Supported

Context Caching

Supported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Germany (Frankfurt)

Scope: Global

Capability Support Capability Support

Input Modality

Text

Output Modality

Text

Model Experience

Supported

Function Calling

Unsupported

Structured Outputs

Unsupported

Web Search

Unsupported

Prefix Completion

Supported

Context Caching

Supported

Batch Inference

Unsupported

Fine-tuning

Unsupported

US (Virginia)

Scope: Global

Capability Support Capability Support

Input Modality

Text

Output Modality

Text

Model Experience

Supported

Function Calling

Unsupported

Structured Outputs

Unsupported

Web Search

Unsupported

Prefix Completion

Supported

Context Caching

Supported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Context Limits

Parameter Value Parameter Value

Max Input Length

258048

Max Output Length

65536

Context Window

262144

Max Input Length (Thinking Mode)

258048

Max Output Length (Thinking Mode)

32768

Max Chain-of-Thought Length

81920

Pricing

This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.

China (Beijing)

Input<=32k

Billing Item Price (CNY) Unit

Input

6

Per 1M tokens

Output

24

Per 1M tokens

Input(Implicit Cache)

1.2

Per 1M tokens

32k<Input<=128k

Billing Item Price (CNY) Unit

Input

10

Per 1M tokens

Output

40

Per 1M tokens

Input(Implicit Cache)

2

Per 1M tokens

128k<Input<=256k

Billing Item Price (CNY) Unit

Input

15

Per 1M tokens

Output

60

Per 1M tokens

Input(Implicit Cache)

3

Per 1M tokens

Singapore

Scope: International

Input<=32k

Billing Item Price (CNY) Unit

Input

8.807

Per 1M tokens

Output

44.035

Per 1M tokens

Input(Implicit Cache)

1.761

Per 1M tokens

32k<Input<=128k

Billing Item Price (CNY) Unit

Input

17.614

Per 1M tokens

Output

88.071

Per 1M tokens

Input(Implicit Cache)

3.523

Per 1M tokens

128k<Input<=256k

Billing Item Price (CNY) Unit

Input

22.018

Per 1M tokens

Output

110.089

Per 1M tokens

Input(Implicit Cache)

4.404

Per 1M tokens

Germany (Frankfurt)

Scope: Global

Input<=32k

Billing Item Price (CNY) Unit

Input

6

Per 1M tokens

Output

24

Per 1M tokens

Input(Implicit Cache)

1.2

Per 1M tokens

32k<Input<=128k

Billing Item Price (CNY) Unit

Input

10

Per 1M tokens

Output

40

Per 1M tokens

Input(Implicit Cache)

2

Per 1M tokens

128k<Input<=256k

Billing Item Price (CNY) Unit

Input

15

Per 1M tokens

Output

60

Per 1M tokens

Input(Implicit Cache)

3

Per 1M tokens

US (Virginia)

Scope: Global

Input<=32k

Billing Item Price (CNY) Unit

Input

6

Per 1M tokens

Output

24

Per 1M tokens

Input(Implicit Cache)

1.2

Per 1M tokens

32k<Input<=128k

Billing Item Price (CNY) Unit

Input

10

Per 1M tokens

Output

40

Per 1M tokens

Input(Implicit Cache)

2

Per 1M tokens

128k<Input<=256k

Billing Item Price (CNY) Unit

Input

15

Per 1M tokens

Output

60

Per 1M tokens

Input(Implicit Cache)

3

Per 1M tokens

Rate Limits

China (Beijing)

Parameter Value

RPM (Requests Per Minute)

600

TPM (Tokens Per Minute)

1,000,000

Singapore

Scope: International

Parameter Value

RPM (Requests Per Minute)

600

TPM (Tokens Per Minute)

1,000,000

Germany (Frankfurt)

Scope: Global

Parameter Value

RPM (Requests Per Minute)

600

TPM (Tokens Per Minute)

1,000,000

US (Virginia)

Scope: Global

Parameter Value

RPM (Requests Per Minute)

600

TPM (Tokens Per Minute)

1,000,000

qwen3-max-2026-01-23

Compared with the snapshot as of September 23, 2025, the Qwen-3 series Max model in this release achieves an effective integration of thinking and non-thinking modes, resulting in a comprehensive and substantial improvement in the model’s overall performance. In thinking mode, the model simultaneously supports web search, web information extraction, and a code interpreter tool, enabling it to tackle more complex and challenging problems with greater accuracy by leveraging external tools while engaging in slow, deliberative reasoning. This version is based on a snapshot taken on January 23, 2026.

Inference Service Provider

The inference service provider for qwen3-max-2026-01-23 is Alibaba Cloud Model Studio.

Model Capabilities

China (Beijing)

Capability Support Capability Support

Input Modality

Text

Output Modality

Text

Model Experience

Supported

Function Calling

Supported

Structured Outputs

Supported

Web Search

Supported

Prefix Completion

Supported

Context Caching

Unsupported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Singapore

Scope: International

Capability Support Capability Support

Input Modality

Text

Output Modality

Text

Model Experience

Supported

Function Calling

Supported

Structured Outputs

Supported

Web Search

Supported

Prefix Completion

Supported

Context Caching

Unsupported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Germany (Frankfurt)

Scope: EU

Capability Support Capability Support

Input Modality

Text

Output Modality

Text

Model Experience

Supported

Function Calling

Supported

Structured Outputs

Supported

Web Search

Unsupported

Prefix Completion

Supported

Context Caching

Unsupported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Context Limits

Parameter Value Parameter Value

Max Input Length

258048

Max Output Length

65536

Context Window

262144

Max Input Length (Thinking Mode)

258048

Max Output Length (Thinking Mode)

32768

Max Chain-of-Thought Length

81920

Pricing

This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.

China (Beijing)

Input<=32k

Billing Item Price (CNY) Unit

Input

2.5

Per 1M tokens

Output

10

Per 1M tokens

32k<Input<=128k

Billing Item Price (CNY) Unit

Input

4

Per 1M tokens

Output

16

Per 1M tokens

128k<Input<=256k

Billing Item Price (CNY) Unit

Input

7

Per 1M tokens

Output

28

Per 1M tokens

Singapore

Scope: International

Input<=32k

Billing Item Price (CNY) Unit

Input

8.807

Per 1M tokens

Output

44.035

Per 1M tokens

32k<Input<=128k

Billing Item Price (CNY) Unit

Input

17.614

Per 1M tokens

Output

88.071

Per 1M tokens

128k<Input<=256k

Billing Item Price (CNY) Unit

Input

22.018

Per 1M tokens

Output

110.089

Per 1M tokens

Germany (Frankfurt)

Scope: EU

Input<=32k

Billing Item Price (CNY) Unit

Input

8.993

Per 1M tokens

Output

44.965

Per 1M tokens

32k<Input<=128k

Billing Item Price (CNY) Unit

Input

17.986

Per 1M tokens

Output

89.93

Per 1M tokens

128k<Input<=256k

Billing Item Price (CNY) Unit

Input

22.483

Per 1M tokens

Output

112.413

Per 1M tokens

Rate Limits

China (Beijing)

Parameter Value

RPM (Requests Per Minute)

600

TPM (Tokens Per Minute)

1,000,000

Singapore

Scope: International

Parameter Value

RPM (Requests Per Minute)

600

TPM (Tokens Per Minute)

1,000,000

Germany (Frankfurt)

Scope: EU

Parameter Value

RPM (Requests Per Minute)

60

TPM (Tokens Per Minute)

100,000

qwen3-max-2025-09-23

The Qwen 3 series Max model has undergone specialized upgrades in agent programming and tool invocation compared to the preview version. The officially released model this time has achieved state-of-the-art (SOTA) performance in its field and is better suited to meet the demands of agents operating in more complex scenarios.This version is a snapshot as of September 23, 2025.

Inference Service Provider

The inference service provider for qwen3-max-2025-09-23 is Alibaba Cloud Model Studio.

Model Capabilities

China (Beijing)

Capability Support Capability Support

Input Modality

Text

Output Modality

Text

Model Experience

Supported

Function Calling

Supported

Structured Outputs

Supported

Web Search

Supported

Prefix Completion

Supported

Context Caching

Unsupported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Singapore

Scope: International

Capability Support Capability Support

Input Modality

Text

Output Modality

Text

Model Experience

Supported

Function Calling

Unsupported

Structured Outputs

Supported

Web Search

Supported

Prefix Completion

Supported

Context Caching

Unsupported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Germany (Frankfurt)

Scope: Global

Capability Support Capability Support

Input Modality

Text

Output Modality

Text

Model Experience

Supported

Function Calling

Unsupported

Structured Outputs

Supported

Web Search

Unsupported

Prefix Completion

Supported

Context Caching

Unsupported

Batch Inference

Unsupported

Fine-tuning

Unsupported

US (Virginia)

Scope: Global

Capability Support Capability Support

Input Modality

Text

Output Modality

Text

Model Experience

Supported

Function Calling

Unsupported

Structured Outputs

Supported

Web Search

Supported

Prefix Completion

Supported

Context Caching

Unsupported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Context Limits

Parameter Value Parameter Value

Max Input Length

258048

Max Output Length

32768

Context Window

262144

Pricing

This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.

China (Beijing)

Input<=32k

Billing Item Price (CNY) Unit

Input

6

Per 1M tokens

Output

24

Per 1M tokens

32k<Input<=128k

Billing Item Price (CNY) Unit

Input

10

Per 1M tokens

Output

40

Per 1M tokens

128k<Input<=256k

Billing Item Price (CNY) Unit

Input

15

Per 1M tokens

Output

60

Per 1M tokens

Singapore

Scope: International

Input<=32k

Billing Item Price (CNY) Unit

Input

8.807

Per 1M tokens

Output

44.035

Per 1M tokens

32k<Input<=128k

Billing Item Price (CNY) Unit

Input

17.614

Per 1M tokens

Output

88.071

Per 1M tokens

128k<Input<=256k

Billing Item Price (CNY) Unit

Input

22.018

Per 1M tokens

Output

110.089

Per 1M tokens

Germany (Frankfurt)

Scope: Global

Input<=32k

Billing Item Price (CNY) Unit

Input

6

Per 1M tokens

Output

24

Per 1M tokens

32k<Input<=128k

Billing Item Price (CNY) Unit

Input

10

Per 1M tokens

Output

40

Per 1M tokens

128k<Input<=256k

Billing Item Price (CNY) Unit

Input

15

Per 1M tokens

Output

60

Per 1M tokens

US (Virginia)

Scope: Global

Input<=32k

Billing Item Price (CNY) Unit

Input

6

Per 1M tokens

Output

24

Per 1M tokens

32k<Input<=128k

Billing Item Price (CNY) Unit

Input

10

Per 1M tokens

Output

40

Per 1M tokens

128k<Input<=256k

Billing Item Price (CNY) Unit

Input

15

Per 1M tokens

Output

60

Per 1M tokens

Rate Limits

China (Beijing)

Parameter Value

RPM (Requests Per Minute)

60

TPM (Tokens Per Minute)

100,000

Singapore

Scope: International

Parameter Value

RPM (Requests Per Minute)

60

TPM (Tokens Per Minute)

100,000

Germany (Frankfurt)

Scope: Global

Parameter Value

RPM (Requests Per Minute)

60

TPM (Tokens Per Minute)

100,000

US (Virginia)

Scope: Global

Parameter Value

RPM (Requests Per Minute)

60

TPM (Tokens Per Minute)

100,000