The Qwen 3 series Max model has undergone specialized upgrades in agent programming and tool invocation compared to the preview version. The officially released model this time has achieved state-of-the-art (SOTA) performance in its field and is better suited to meet the demands of agents operating in more complex scenarios.This model version is functionally equivalent to the snapshot model qwen3-max-2026-01-23.
Inference Service Provider
The inference service provider for qwen3-max is Alibaba Cloud Model Studio.
Model Capabilities
China (Beijing)
| Capability | Support | Capability | Support |
|---|---|---|---|
Input Modality |
Text |
Output Modality |
Text |
Model Experience |
Function Calling |
||
Structured Outputs |
Web Search |
||
Prefix Completion |
Context Caching |
||
Batch Inference |
Fine-tuning |
Singapore
Scope: International
| Capability | Support | Capability | Support |
|---|---|---|---|
Input Modality |
Text |
Output Modality |
Text |
Model Experience |
Function Calling |
||
Structured Outputs |
Web Search |
||
Prefix Completion |
Context Caching |
||
Batch Inference |
Fine-tuning |
Germany (Frankfurt)
Scope: EU, Global
| Capability | Support | Capability | Support |
|---|---|---|---|
Input Modality |
Text |
Output Modality |
Text |
Model Experience |
Function Calling |
||
Structured Outputs |
Web Search |
||
Prefix Completion |
Context Caching |
||
Batch Inference |
Fine-tuning |
US (Virginia)
Scope: Global
| Capability | Support | Capability | Support |
|---|---|---|---|
Input Modality |
Text |
Output Modality |
Text |
Model Experience |
Function Calling |
||
Structured Outputs |
Web Search |
||
Prefix Completion |
Context Caching |
||
Batch Inference |
Fine-tuning |
Context Limits
| Parameter | Value | Parameter | Value |
|---|---|---|---|
Max Input Length |
258048 |
Max Output Length |
65536 |
Context Window |
262144 |
Max Input Length (Thinking Mode) |
258048 |
Max Output Length (Thinking Mode) |
32768 |
Max Chain-of-Thought Length |
81920 |
Pricing
This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.
China (Beijing)
Singapore
Scope: International
Germany (Frankfurt)
Scope: EU
Scope: Global
US (Virginia)
Scope: Global
Rate Limits
China (Beijing)
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) |
30000 |
TPM (Tokens Per Minute) |
5,000,000 |
Singapore
Scope: International
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) |
600 |
TPM (Tokens Per Minute) |
1,000,000 |
Germany (Frankfurt)
Scope: EU
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) |
600 |
TPM (Tokens Per Minute) |
1,000,000 |
Scope: Global
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) |
600 |
TPM (Tokens Per Minute) |
1,000,000 |
US (Virginia)
Scope: Global
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) |
600 |
TPM (Tokens Per Minute) |
1,000,000 |
Snapshot Versions
qwen3-max-preview
A preview version of the Max model in the Qwen 3 series, achieving an effective integration of thinking and non-thinking modes. In thinking mode, there is a significant enhancement in capabilities such as intelligent agent programming, common-sense reasoning, and reasoning across mathematics, science, and general domains.
Inference Service Provider
The inference service provider for qwen3-max-preview is Alibaba Cloud Model Studio.
Model Capabilities
China (Beijing)
| Capability | Support | Capability | Support |
|---|---|---|---|
Input Modality |
Text |
Output Modality |
Text |
Model Experience |
Function Calling |
||
Structured Outputs |
Web Search |
||
Prefix Completion |
Context Caching |
||
Batch Inference |
Fine-tuning |
Singapore
Scope: International
| Capability | Support | Capability | Support |
|---|---|---|---|
Input Modality |
Text |
Output Modality |
Text |
Model Experience |
Function Calling |
||
Structured Outputs |
Web Search |
||
Prefix Completion |
Context Caching |
||
Batch Inference |
Fine-tuning |
Germany (Frankfurt)
Scope: Global
| Capability | Support | Capability | Support |
|---|---|---|---|
Input Modality |
Text |
Output Modality |
Text |
Model Experience |
Function Calling |
||
Structured Outputs |
Web Search |
||
Prefix Completion |
Context Caching |
||
Batch Inference |
Fine-tuning |
US (Virginia)
Scope: Global
| Capability | Support | Capability | Support |
|---|---|---|---|
Input Modality |
Text |
Output Modality |
Text |
Model Experience |
Function Calling |
||
Structured Outputs |
Web Search |
||
Prefix Completion |
Context Caching |
||
Batch Inference |
Fine-tuning |
Context Limits
| Parameter | Value | Parameter | Value |
|---|---|---|---|
Max Input Length |
258048 |
Max Output Length |
65536 |
Context Window |
262144 |
Max Input Length (Thinking Mode) |
258048 |
Max Output Length (Thinking Mode) |
32768 |
Max Chain-of-Thought Length |
81920 |
Pricing
This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.
China (Beijing)
Singapore
Scope: International
Germany (Frankfurt)
Scope: Global
US (Virginia)
Scope: Global
Rate Limits
China (Beijing)
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) |
600 |
TPM (Tokens Per Minute) |
1,000,000 |
Singapore
Scope: International
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) |
600 |
TPM (Tokens Per Minute) |
1,000,000 |
Germany (Frankfurt)
Scope: Global
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) |
600 |
TPM (Tokens Per Minute) |
1,000,000 |
US (Virginia)
Scope: Global
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) |
600 |
TPM (Tokens Per Minute) |
1,000,000 |
qwen3-max-2026-01-23
Compared with the snapshot as of September 23, 2025, the Qwen-3 series Max model in this release achieves an effective integration of thinking and non-thinking modes, resulting in a comprehensive and substantial improvement in the model’s overall performance. In thinking mode, the model simultaneously supports web search, web information extraction, and a code interpreter tool, enabling it to tackle more complex and challenging problems with greater accuracy by leveraging external tools while engaging in slow, deliberative reasoning. This version is based on a snapshot taken on January 23, 2026.
Inference Service Provider
The inference service provider for qwen3-max-2026-01-23 is Alibaba Cloud Model Studio.
Model Capabilities
China (Beijing)
| Capability | Support | Capability | Support |
|---|---|---|---|
Input Modality |
Text |
Output Modality |
Text |
Model Experience |
Function Calling |
||
Structured Outputs |
Web Search |
||
Prefix Completion |
Context Caching |
||
Batch Inference |
Fine-tuning |
Singapore
Scope: International
| Capability | Support | Capability | Support |
|---|---|---|---|
Input Modality |
Text |
Output Modality |
Text |
Model Experience |
Function Calling |
||
Structured Outputs |
Web Search |
||
Prefix Completion |
Context Caching |
||
Batch Inference |
Fine-tuning |
Germany (Frankfurt)
Scope: EU
| Capability | Support | Capability | Support |
|---|---|---|---|
Input Modality |
Text |
Output Modality |
Text |
Model Experience |
Function Calling |
||
Structured Outputs |
Web Search |
||
Prefix Completion |
Context Caching |
||
Batch Inference |
Fine-tuning |
Context Limits
| Parameter | Value | Parameter | Value |
|---|---|---|---|
Max Input Length |
258048 |
Max Output Length |
65536 |
Context Window |
262144 |
Max Input Length (Thinking Mode) |
258048 |
Max Output Length (Thinking Mode) |
32768 |
Max Chain-of-Thought Length |
81920 |
Pricing
This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.
China (Beijing)
Singapore
Scope: International
Germany (Frankfurt)
Scope: EU
Rate Limits
China (Beijing)
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) |
600 |
TPM (Tokens Per Minute) |
1,000,000 |
Singapore
Scope: International
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) |
600 |
TPM (Tokens Per Minute) |
1,000,000 |
Germany (Frankfurt)
Scope: EU
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) |
60 |
TPM (Tokens Per Minute) |
100,000 |
qwen3-max-2025-09-23
The Qwen 3 series Max model has undergone specialized upgrades in agent programming and tool invocation compared to the preview version. The officially released model this time has achieved state-of-the-art (SOTA) performance in its field and is better suited to meet the demands of agents operating in more complex scenarios.This version is a snapshot as of September 23, 2025.
Inference Service Provider
The inference service provider for qwen3-max-2025-09-23 is Alibaba Cloud Model Studio.
Model Capabilities
China (Beijing)
| Capability | Support | Capability | Support |
|---|---|---|---|
Input Modality |
Text |
Output Modality |
Text |
Model Experience |
Function Calling |
||
Structured Outputs |
Web Search |
||
Prefix Completion |
Context Caching |
||
Batch Inference |
Fine-tuning |
Singapore
Scope: International
| Capability | Support | Capability | Support |
|---|---|---|---|
Input Modality |
Text |
Output Modality |
Text |
Model Experience |
Function Calling |
||
Structured Outputs |
Web Search |
||
Prefix Completion |
Context Caching |
||
Batch Inference |
Fine-tuning |
Germany (Frankfurt)
Scope: Global
| Capability | Support | Capability | Support |
|---|---|---|---|
Input Modality |
Text |
Output Modality |
Text |
Model Experience |
Function Calling |
||
Structured Outputs |
Web Search |
||
Prefix Completion |
Context Caching |
||
Batch Inference |
Fine-tuning |
US (Virginia)
Scope: Global
| Capability | Support | Capability | Support |
|---|---|---|---|
Input Modality |
Text |
Output Modality |
Text |
Model Experience |
Function Calling |
||
Structured Outputs |
Web Search |
||
Prefix Completion |
Context Caching |
||
Batch Inference |
Fine-tuning |
Context Limits
| Parameter | Value | Parameter | Value |
|---|---|---|---|
Max Input Length |
258048 |
Max Output Length |
32768 |
Context Window |
262144 |
Pricing
This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.
China (Beijing)
Singapore
Scope: International
Germany (Frankfurt)
Scope: Global
US (Virginia)
Scope: Global
Rate Limits
China (Beijing)
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) |
60 |
TPM (Tokens Per Minute) |
100,000 |
Singapore
Scope: International
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) |
60 |
TPM (Tokens Per Minute) |
100,000 |
Germany (Frankfurt)
Scope: Global
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) |
60 |
TPM (Tokens Per Minute) |
100,000 |
US (Virginia)
Scope: Global
| Parameter | Value |
|---|---|
RPM (Requests Per Minute) |
60 |
TPM (Tokens Per Minute) |
100,000 |