AI inference fees

Updated at:

The MaxCompute AI inference is a new, out-of-the-box, pay-as-you-go feature that lets you use large models for data processing or offline inference. This topic describes the billing rules for the AI inference.

Overview

The MaxCompute AI inference service provides large model inference, letting you call out-of-the-box models directly from SQL and MaxFrame jobs using a built-in AI function. You are billed on a pay-as-you-go basis according to the total number of input and output tokens.

  • Supported regions: The model inference service is available only in the following regions.

    • Chinese mainland: China (Beijing), China (Shanghai), China (Hangzhou), China (Shenzhen), and China (Ulanqab)

    • International: Singapore

  • Supported models: To use the following models, you must first enable the model inference service in the target region, such as China (Beijing).

    Model

    Type

    Description

    qwen3.8-max

    Supports multimodal data input

    A flagship Mixture of Experts (MoE) model from the Qwen 3.8 series with 2.4 trillion parameters. It features significantly enhanced programming and office productivity capabilities, and can autonomously program for over ten days to deliver complete projects. This model is ideal for highly complex scenarios such as autonomous programming, advanced office automation, scientific research, and long-duration tasks.

    qwen3.7-max

    Text input only

    A flagship model from the Qwen 3.7 series, featuring comprehensive upgrades in inference, code generation, and multilingual understanding. It is suitable for highly complex tasks.

    qwen3.7-plus

    Supports multimodal data input

    A balanced model from the Qwen 3.7 series that offers a good trade-off between performance and cost. It is suitable for enterprise scenarios such as long-text analysis and multi-turn conversations.

    qwen3.7-flash

    Supports multimodal data input

    A lightweight, high-speed, native vision-language model from the Qwen 3.7 series. It features enhanced multimodal understanding and agent execution capabilities, making it ideal for high-concurrency, low-latency online services and lightweight multimodal tasks.

    qwen3.7-text-embedding

    Text input only

    This unified multilingual text vector model, trained on Qwen 3.7, offers significant performance improvements in text retrieval, clustering, and classification compared to the text-embedding-v4 version.

    qwen3-vl-embedding

    Supports multimodal data input

    A Qwen multimodal vector embedding model that supports vector representations for mixed image and text inputs. It is suitable for cross-modal retrieval and image-text matching scenarios.

    text-embedding-v4

    Text input only

    A high-precision text vector embedding model suitable for semantic search, clustering, and similarity calculation.

    qwen3.6-plus

    Supports multimodal data input

    A balanced model from the Qwen 3.6 series that offers a good trade-off between performance and cost. It is suitable for general business scenarios such as dialogue, summarization, and analysis.

    qwen3.6-flash

    Supports multimodal data input

    A lightweight, high-speed model from the Qwen 3.6 series. It provides fast responses at a low cost, making it ideal for high-concurrency, low-latency online services.

    deepseek-v4-pro

    Text input only

    The fourth-generation flagship inference model from DeepSeek. It excels at high-difficulty tasks such as mathematics, coding, and complex logical reasoning.

    deepseek-v4-flash

    Text input only

    The fourth-generation lightweight model from DeepSeek. It offers fast inference and is highly cost-effective, making it suitable for daily conversations and lightweight inference tasks.

    qwen3.5-397b-a17b

    Supports multimodal data input

    A Mixture of Experts (MoE) model from the Qwen 3.5 series. It features efficient parameter activation, a vast knowledge base, and multi-step reasoning capabilities. This model is designed for high-difficulty logical deduction, full-stack code generation, and in-depth knowledge Q&A.

    qwen3.5-omni-plus

    Supports multimodal data input

    Qwen3.5-Omni is Qwen's latest omni-modal large model. It is suitable for scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal understanding and interaction experience.

    qwen3-asr-flash

    Supports multimodal data input

    A lightweight speech recognition model from the Qwen 3 series. It provides real-time transcription with low latency and high concurrency, making it suitable for real-time audio processing scenarios such as generating meeting minutes, live-streaming subtitles, and recognizing voice commands.

    tongyi-embedding-vision-plus

    Supports multimodal data input

    Tongyi-Embedding-Vision is a visual multimodal representation model built on a large language model (LLM). It is suitable for a variety of downstream tasks, including image-to-image search, text-to-image search, text-to-video search, video-to-video search, text-to-text search, and text-to-image-and-text search.

    qwen3-max (Sunsetting soon)

    Text input only

    A high-performance large language model from the Qwen series, suitable for complex inference and content generation.

    Important

    The qwen3-max model is sunsetting soon. To ensure service continuity, migrate your related services to a new model as soon as possible. For more information, see Model Decommissioning Policy.

Note

You incur model inference service fees only when you use the public models listed above in MaxFrame Overview and SQL jobs using an AI function. You are charged only for successful jobs; failed jobs do not incur fees.

Billing rules

The model inference service is billed based on token usage. The billing dimensions include:

  • Region

  • Model type

  • Token type: Distinguishes between input and output tokens and indicates the usage-based pricing tier.

Pricing

A single SQL or MaxFrame job may involve multiple model invocations. If the model has pricing tiers, the system calculates the input token count for each invocation separately to determine its billing tier (token type). The system then meters and bills usage separately for each token type.

Important
  • For deployments in China regions, compute resources for model inference are limited to the Chinese mainland, and static data is stored in the selected region.

  • For deployments in International regions, compute resources for model inference are dynamically scheduled worldwide (excluding the Chinese mainland), and static data is stored in the selected region.

qwen3.8-max

Chinese mainland

For deployments in China regions, compute resources for model inference are limited to the Chinese mainland, and static data is stored in the selected region.

Supported regions:

Chinese mainland: China (Beijing), China (Shanghai), China (Hangzhou), China (Shenzhen), China (Ulanqab)

Input tokens

Use case

Token type

Price (per million tokens)

0 < Tokens ≤ 1,048,576

model input

input_token_tier1

CNY 14.4

model input

(implicit cache hit)

input_token_tier1_cached

CNY 2.88

model input

(explicit cache hit)

input_token_tier1_cached_explicit

CNY 1.44

model input

(create explicit cache)

input_token_tier1_create_cache

CNY 18

model output (non-thinking)

output_token_tier1

CNY 43.2

model output (thinking mode)

output_token_tier1_thinking

CNY 43.2

International

For deployments in International regions, compute resources for model inference are dynamically scheduled worldwide (excluding the Chinese mainland), and static data is stored in the selected region.

Supported regions: Singapore

Input tokens

Use case

Token type

Price (per million tokens)

0 < Tokens ≤ 1,048,576

model input

input_token_tier1

CNY 17.9856

model input

(implicit cache hit)

input_token_tier1_cached

CNY 3.59712

model input

(explicit cache hit)

input_token_tier1_cached_explicit

CNY 1.79856

model input

(create explicit cache)

input_token_tier1_create_cache

CNY 22.482

model output (non-thinking)

output_token_tier1

CNY 53.958

model output (thinking mode)

output_token_tier1_thinking

CNY 53.958

Qwen3.7-max

Chinese mainland

For deployments in China regions, compute resources for model inference are limited to the Chinese mainland, and static data is stored in the selected region.

Supported regions:

Chinese mainland: China (Beijing), China (Shanghai), China (Hangzhou), China (Shenzhen), China (Ulanqab)

Input tokens

Use case

Token type

Price (per 1M tokens)

0 < tokens ≤ 1,048,576

Model input

input_token_tier1

CNY 14.4

Model input

(implicit cache hit)

input_token_tier1_cached

CNY 2.88

Model input

(explicit cache hit)

input_token_tier1_cached_explicit

CNY 1.44

Model input

(create explicit cache)

input_token_tier1_create_cache

CNY 18

Model output

output_token_tier1

CNY 43.2

Model output (thinking mode)

output_token_tier1_thinking

CNY 43.2

International

For deployments in International regions, compute resources for model inference are dynamically scheduled worldwide (excluding the Chinese mainland), and static data is stored in the selected region.

Supported region: Singapore

Input tokens

Use case

Token type

Price (per 1M tokens)

0 < tokens ≤ 1,048,576

Model input

input_token_tier1

CNY 22.4832

Model input

(implicit cache hit)

input_token_tier1_cached

CNY 4.49664

Model input

(explicit cache hit)

input_token_tier1_cached_explicit

CNY 2.24832

Model input

(create explicit cache)

input_token_tier1_create_cache

CNY 28.104

Model output

output_token_tier1

CNY 67.4484

Model output (thinking mode)

output_token_tier1_thinking

CNY 67.4484

qwen3.7-plus

Mainland China

For deployments in China regions, compute resources for model inference are limited to the Chinese mainland, and static data is stored in the selected region.

Supported regions:

Chinese mainland: China (Beijing), China (Shanghai), China (Hangzhou), China (Shenzhen), China (Ulanqab)

Tokens per inference

Scenario

Token type

Price / million tokens

0 < Token ≤ 262,144

model input

input_token_tier1

CNY 2.4

model input (implicit cache hit)

(implicit cache hit)

input_token_tier1_cached

CNY 0.48

model input (explicit cache hit)

(explicit cache hit)

input_token_tier1_cached_explicit

CNY 0.24

model input (create explicit cache)

(explicit cache creation)

input_token_tier1_create_cache

CNY 3

model output (non-thinking)

output_token_tier1

CNY 9.6

model output (thinking mode)

output_token_tier1_thinking

CNY 9.6

262,144 < Token ≤ 1,048,576

model input

input_token_tier2

CNY 7.2

model input (implicit cache hit)

(implicit cache hit)

input_token_tier2_cached

CNY 1.44

model input (explicit cache hit)

(Explicit cache hit)

input_token_tier2_cached_explicit

CNY 0.72

model input (create explicit cache)

(Explicit cache creation)

input_token_tier2_create_cache

CNY 9

model output (non-thinking)

output_token_tier2

CNY 28.8

model output (thinking mode)

output_token_tier2_thinking

CNY 28.8

International

For deployments in International regions, compute resources for model inference are dynamically scheduled worldwide (excluding the Chinese mainland), and static data is stored in the selected region.

Supported regions: Singapore

Tokens per inference

Scenario

Token type

Price / million tokens

0 < Token ≤ 262,144

model input

input_token_tier1

CNY 3.5976

model input (implicit cache hit)

(Implicit cache hit)

input_token_tier1_cached

CNY 0.71952

model input (explicit cache hit)

(explicit cache hit)

input_token_tier1_cached_explicit

CNY 0.35976

model input (create explicit cache)

(Explicit cache creation)

input_token_tier1_create_cache

CNY 4.497

model output (non-thinking)

output_token_tier1

CNY 14.3892

model output (thinking mode)

output_token_tier1_thinking

CNY 14.3892

262,144 < Token ≤ 1,048,576

model input

input_token_tier2

CNY 10.7916

model input (implicit cache hit)

(implicit cache hit)

input_token_tier2_cached

CNY 2.15832

model input (explicit cache hit)

(explicit cache hit)

input_token_tier2_cached_explicit

CNY 1.07916

model input (create explicit cache)

(Explicit cache creation)

input_token_tier2_create_cache

CNY 13.4895

model output (non-thinking)

output_token_tier2

CNY 43.1664

model output (thinking mode)

output_token_tier2_thinking

CNY 43.1664

qwen3.7-flash

Chinese mainland

For deployments in China regions, compute resources for model inference are limited to the Chinese mainland, and static data is stored in the selected region.

Supported regions:

Chinese mainland: China (Beijing), China (Shanghai), China (Hangzhou), China (Shenzhen), China (Ulanqab)

Input tokens

Usage scenario

Token type

Price (per million tokens)

0 < tokens ≤ 32,768

model input

input_token_tier1

CNY 0.24

model input

(implicit cache hit)

input_token_tier1_cached

CNY 0.048

model input

(explicit cache hit)

input_token_tier1_cached_explicit

CNY 0.024

model input

(create explicit cache)

input_token_tier1_create_cache

CNY 0.3

model output (non-thinking)

output_token_tier1

CNY 0.96

model output (thinking mode)

output_token_tier1_thinking

CNY 0.96

32,768 < tokens ≤ 262,144

model input

input_token_tier2

CNY 0.72

model input

(implicit cache hit)

input_token_tier2_cached

CNY 0.144

model input

(explicit cache hit)

input_token_tier2_cached_explicit

CNY 0.072

model input

(create explicit cache)

input_token_tier2_create_cache

CNY 0.9

model output (non-thinking)

output_token_tier2

CNY 2.88

model output (thinking mode)

output_token_tier2_thinking

CNY 2.88

262,144 < tokens ≤ 1,048,576

model input

input_token_tier3

CNY 1.44

model input

(implicit cache hit)

input_token_tier3_cached

CNY 0.288

model input

(explicit cache hit)

input_token_tier3_cached_explicit

CNY 0.144

model input

(create explicit cache)

input_token_tier3_create_cache

CNY 1.8

model output (non-thinking)

output_token_tier3

CNY 5.76

model output (thinking mode)

output_token_tier3_thinking

CNY 5.76

International

For deployments in International regions, compute resources for model inference are dynamically scheduled worldwide (excluding the Chinese mainland), and static data is stored in the selected region.

Supported regions: Singapore

Input tokens

Usage scenario

Token type

Price (per million tokens)

0 < tokens ≤ 32,768

model input

input_token_tier1

CNY 0.27

model input

(implicit cache hit)

input_token_tier1_cached

CNY 0.054

model input

(explicit cache hit)

input_token_tier1_cached_explicit

CNY 0.027

model input

(create explicit cache)

input_token_tier1_create_cache

CNY 0.3375

model output (non-thinking)

output_token_tier1

CNY 1.1688

model output (thinking mode)

output_token_tier1_thinking

CNY 1.1688

32,768 < tokens ≤ 262,144

model input

input_token_tier2

CNY 0.8988

model input

(implicit cache hit)

input_token_tier2_cached

CNY 0.17976

model input

(explicit cache hit)

input_token_tier2_cached_explicit

CNY 0.08988

model input

(create explicit cache)

input_token_tier2_create_cache

CNY 1.1235

model output (non-thinking)

output_token_tier2

CNY 3.5976

model output (thinking mode)

output_token_tier2_thinking

CNY 3.5976

262,144 < tokens ≤ 1,048,576

model input

input_token_tier3

CNY 1.7988

model input

(implicit cache hit)

input_token_tier3_cached

CNY 0.35976

model input

(explicit cache hit)

input_token_tier3_cached_explicit

CNY 0.17988

model input

(create explicit cache)

input_token_tier3_create_cache

CNY 2.2485

model output (non-thinking)

output_token_tier3

CNY 7.194

model output (thinking mode)

output_token_tier3_thinking

CNY 7.194

qwen3.7-text-embedding

The qwen3.7-text-embedding model is billed for input tokens only, with no charges for output tokens. This model has no pricing tiers, and all usage is billed as input_token_tier1.

For deployments in China regions, compute resources for model inference are limited to the Chinese mainland, and static data is stored in the selected region.

Supported regions:

Chinese mainland: China (Beijing), China (Shanghai), China (Hangzhou), China (Shenzhen), China (Ulanqab)

Tier

Use case

Token type

Unit price

No tiers

model input

input_token_tier1

CNY 0.6

qwen3-vl-embedding

For deployments in China regions, compute resources for model inference are limited to the Chinese mainland, and static data is stored in the selected region.

Supported regions:

Chinese mainland: China (Beijing), China (Shanghai), China (Hangzhou), China (Shenzhen), China (Ulanqab)

Region

Type

Use case

Token type

Unit price

Chinese mainland

text

model input

input_token_tier1_text

CNY 0.84

image and video

model input

input_token_tier1_image

CNY 2.16

text-embedding-v4

For text-embedding-v4, you are charged for model input tokens only. There are no charges for model output tokens and no pricing tiers. All usage is billed as input_token_tier1.

Chinese mainland

For deployments in China regions, compute resources for model inference are limited to the Chinese mainland, and static data is stored in the selected region.

Supported regions:

Chinese mainland: China (Beijing), China (Shanghai), China (Hangzhou), China (Shenzhen), China (Ulanqab)

Pricing tier

Use case

Token type

Price per million tokens

No tiers

Model input

input_token_tier1

CNY 0.6

International

For deployments in International regions, compute resources for model inference are dynamically scheduled worldwide (excluding the Chinese mainland), and static data is stored in the selected region.

Supported regions: Singapore

Pricing tier

Use case

Token type

Price per million tokens

No tiers

Model input

input_token_tier1

CNY 0.6168

qwen3.6-plus

Chinese mainland

For deployments in China regions, compute resources for model inference are limited to the Chinese mainland, and static data is stored in the selected region.

Supported regions:

Chinese mainland: China (Beijing), China (Shanghai), China (Hangzhou), China (Shenzhen), China (Ulanqab)

Input tokens

Scenario

Token type

Price (per million tokens)

0 < tokens ≤ 262,144

model input

input_token_tier1

CNY 2.4

model input

(implicit cache hit)

input_token_tier1_cached

CNY 0.48

model input

(explicit cache hit)

input_token_tier1_cached_explicit

CNY 0.24

model input

(create explicit cache)

input_token_tier1_create_cache

CNY 3

model output (non-thinking)

output_token_tier1

CNY 14.4

model output (thinking mode)

output_token_tier1_thinking

CNY 14.4

262,144 < tokens ≤ 1,048,576

model input

input_token_tier2

CNY 9.6

model input

(implicit cache hit)

input_token_tier2_cached

CNY 1.92

model input

(explicit cache hit)

input_token_tier2_cached_explicit

CNY 0.96

model input

(create explicit cache)

input_token_tier2_create_cache

CNY 12

model output (non-thinking)

output_token_tier2

CNY 57.6

model output (thinking mode)

output_token_tier2_thinking

CNY 57.6

International

For deployments in International regions, compute resources for model inference are dynamically scheduled worldwide (excluding the Chinese mainland), and static data is stored in the selected region.

Supported regions: Singapore

Input tokens

Scenario

Token type

Price (per million tokens)

0 < tokens ≤ 262,144

model input

input_token_tier1

CNY 4.49652

model input

(implicit cache hit)

input_token_tier1_cached

CNY 0.899304

model input

(explicit cache hit)

input_token_tier1_cached_explicit

CNY 0.449652

model input

(create explicit cache)

input_token_tier1_create_cache

CNY 5.62065

model output (non-thinking)

output_token_tier1

CNY 26.97912

model output (thinking mode)

output_token_tier1_thinking

CNY 26.97912

262,144 < tokens ≤ 1,048,576

model input

input_token_tier2

CNY 17.98608

model input

(implicit cache hit)

input_token_tier2_cached

CNY 3.597216

model input

(explicit cache hit)

input_token_tier2_cached_explicit

CNY 1.798608

model input

(create explicit cache)

input_token_tier2_create_cache

CNY 22.4826

model output (non-thinking)

output_token_tier2

CNY 53.958

model output (thinking mode)

output_token_tier2_thinking

CNY 53.958

qwen3.6-flash

Chinese mainland

For deployments in China regions, compute resources for model inference are limited to the Chinese mainland, and static data is stored in the selected region.

Supported regions:

Chinese mainland: China (Beijing), China (Shanghai), China (Hangzhou), China (Shenzhen), China (Ulanqab)

Input tokens

Use case

Token type

Price (per million tokens)

0 < tokens ≤ 262,144

model input

input_token_tier1

CNY 1.44

model input

(implicit cache hit)

input_token_tier1_cached

CNY 0.288

model input

(explicit cache hit)

input_token_tier1_cached_explicit

CNY 0.144

model input

(create explicit cache)

input_token_tier1_create_cache

CNY 1.8

model output (non-thinking)

output_token_tier1

CNY 8.64

model output (thinking mode)

output_token_tier1_thinking

CNY 8.64

262,144 < tokens ≤ 1,048,576

model input

input_token_tier2

CNY 5.76

model input

(implicit cache hit)

input_token_tier2_cached

CNY 1.152

model input

(explicit cache hit)

input_token_tier2_cached_explicit

CNY 0.576

model input

(create explicit cache)

input_token_tier2_create_cache

CNY 7.2

model output (non-thinking)

output_token_tier2

CNY 34.56

model output (thinking mode)

output_token_tier2_thinking

CNY 34.56

International

For deployments in International regions, compute resources for model inference are dynamically scheduled worldwide (excluding the Chinese mainland), and static data is stored in the selected region.

Supported regions: Singapore

Input tokens

Use case

Token type

Price (per million tokens)

0 < tokens ≤ 262,144

model input

input_token_tier1

CNY 2.24826

model input

(implicit cache hit)

input_token_tier1_cached

CNY 0.449652

model input

(explicit cache hit)

input_token_tier1_cached_explicit

CNY 0.224826

model input

(create explicit cache)

input_token_tier1_create_cache

CNY 2.810325

model output (non-thinking)

output_token_tier1

CNY 13.48956

model output (thinking mode)

output_token_tier1_thinking

CNY 13.48956

262,144 < tokens ≤ 1,048,576

model input

input_token_tier2

CNY 8.99304

model input

(implicit cache hit)

input_token_tier2_cached

CNY 1.798608

model input

(explicit cache hit)

input_token_tier2_cached_explicit

CNY 0.899304

model input

(create explicit cache)

input_token_tier2_create_cache

CNY 11.2413

model output (non-thinking)

output_token_tier2

CNY 35.97096

model output (thinking mode)

output_token_tier2_thinking

CNY 35.97096

deepseek-v4-pro

Chinese mainland

For deployments in China regions, compute resources for model inference are limited to the Chinese mainland, and static data is stored in the selected region.

Supported regions:

Chinese mainland: China (Beijing), China (Shanghai), China (Hangzhou), China (Shenzhen), China (Ulanqab)

Pricing tier

Use case

Token type

Unit price

No tiers

model input

input_token_tier1

CNY 14.4

model input

(implicit cache hit)

input_token_tier1_cached

CNY 2.88

model input

(explicit cache hit)

input_token_tier1_cached_explicit

CNY 1.44

model input

(create explicit cache)

input_token_tier1_create_cache

CNY 18

model output (non-thinking)

output_token_tier1

CNY 28.8

model output (thinking mode)

output_token_tier1_thinking

CNY 28.8

International

For deployments in International regions, compute resources for model inference are dynamically scheduled worldwide (excluding the Chinese mainland), and static data is stored in the selected region.

Supported regions: Singapore

Input tokens per inference

Use case

Token type

Unit price

0 < tokens ≤ 1,048,576

model input

input_token_tier1

CNY 21.5832

model input

(implicit cache hit)

input_token_tier1_cached

CNY 4.31664

model input

(explicit cache hit)

input_token_tier1_cached_explicit

CNY 2.15832

model input

(create explicit cache)

input_token_tier1_create_cache

CNY 26.979

model output (non-thinking)

output_token_tier1

CNY 43.1664

model output (thinking mode)

output_token_tier1_thinking

CNY 43.1664

deepseek-v4-flash

Chinese mainland

For deployments in China regions, compute resources for model inference are limited to the Chinese mainland, and static data is stored in the selected region.

Supported regions:

Chinese mainland: China (Beijing), China (Shanghai), China (Hangzhou), China (Shenzhen), China (Ulanqab)

Pricing tier

Use case

Token type

Unit price

No tiers

model input

input_token_tier1

CNY 1.2

model input

(implicit cache hit)

input_token_tier1_cached

CNY 0.24

model input

(explicit cache hit)

input_token_tier1_cached_explicit

CNY 0.12

model input

(explicit cache creation)

input_token_tier1_create_cache

CNY 1.5

model output (non-thinking)

output_token_tier1

CNY 2.4

model output (thinking mode)

output_token_tier1_thinking

CNY 2.4

International

For deployments in International regions, compute resources for model inference are dynamically scheduled worldwide (excluding the Chinese mainland), and static data is stored in the selected region.

Supported region: Singapore

Pricing tier

Use case

Token type

Unit price

0 < tokens ≤ 1,048,576

model input

input_token_tier1

CNY 1.7988

model input

(implicit cache hit)

input_token_tier1_cached

CNY 0.35976

model input

(explicit cache hit)

input_token_tier1_cached_explicit

CNY 0.17988

model input

(explicit cache creation)

input_token_tier1_create_cache

CNY 2.2485

model output (non-thinking)

output_token_tier1

CNY 3.5976

model output (thinking mode)

output_token_tier1_thinking

CNY 3.5976

qwen3.5-397b-a17b

For deployments in China regions, compute resources for model inference are limited to the Chinese mainland, and static data is stored in the selected region.

Supported regions:

Chinese mainland: China (Beijing), China (Shanghai), China (Hangzhou), China (Shenzhen), China (Ulanqab)

Tokens per inference

Usage scenario

Type

Unit price (per million tokens)

0 < tokens ≤ 131,072

model input

input_token_tier1

CNY 1.44

model input

(implicit cache hit)

input_token_tier1_cached

CNY 0.288

model input

(explicit cache hit)

input_token_tier1_cached_explicit

CNY 0.144

model input

(create explicit cache)

input_token_tier1_create_cache

CNY 1.8

model output (non-thinking)

output_token_tier1

CNY 8.64

model output (thinking mode)

output_token_tier1_thinking

CNY 8.64

131,072 < tokens ≤ 262,144

model input

input_token_tier2

CNY 3.6

model input

(implicit cache hit)

input_token_tier2_cached

CNY 0.72

model input

(explicit cache hit)

input_token_tier2_cached_explicit

CNY 0.36

model input

(create explicit cache)

input_token_tier2_create_cache

CNY 4.5

model output (non-thinking)

output_token_tier2

CNY 21.6

model output (thinking mode)

output_token_tier2_thinking

CNY 21.6

qwen3.5-omni-plus

Chinese mainland

For deployments in China regions, compute resources for model inference are limited to the Chinese mainland, and static data is stored in the selected region.

Supported regions:

Chinese mainland: China (Beijing), China (Shanghai), China (Hangzhou), China (Shenzhen), China (Ulanqab)

Price tier

Use case

Token type

Unit price

Text

model input

input_token_tier1_text

CNY 8.4

Image

model input

input_token_tier1_image

CNY 8.4

Video

model input

input_token_tier1_video

CNY 8.4

Audio

model input

input_token_tier1_audio

CNY 63.6

Text

model output

output_token_tier1_text

CNY 48

Audio

model output

output_token_tier1_audio

CNY 255.6

International

For deployments in International regions, compute resources for model inference are dynamically scheduled worldwide (excluding the Chinese mainland), and static data is stored in the selected region.

Supported region: Singapore

Price tier

Use case

Token type

Unit price

Text

model input

input_token_tier1_text

CNY 12.588

Image

model input

input_token_tier1_image

CNY 12.588

Video

model input

input_token_tier1_video

CNY 12.588

Audio

model input

input_token_tier1_audio

CNY 98.928

Text

model output

output_token_tier1_text

CNY 74.64

Audio

model output

output_token_tier1_audio

CNY 395.688

qwen3-asr-flash

Billing for qwen3-asr-flash is based only on model input tokens. There are no charges for model output and no pricing tiers. All usage is billed under input_token_tier1.

Chinese mainland

For deployments in China regions, compute resources for model inference are limited to the Chinese mainland, and static data is stored in the selected region.

Supported regions:

Chinese mainland: China (Beijing), China (Shanghai), China (Hangzhou), China (Shenzhen), China (Ulanqab)

Tier

Use case

Token type

Unit price

No tiers

model input

input_token_tier1

CNY 10.56

International

For deployments in International regions, compute resources for model inference are dynamically scheduled worldwide (excluding the Chinese mainland), and static data is stored in the selected region.

Supported regions: Singapore

Tier

Use case

Token type

Unit price

No tiers

model input

input_token_tier1

CNY 12.48

tongyi-embedding-vision-plus

For deployments in China regions, compute resources for model inference are limited to the Chinese mainland, and static data is stored in the selected region.

Supported regions:

Chinese mainland: China (Beijing), China (Shanghai), China (Hangzhou), China (Shenzhen), China (Ulanqab)

Pricing tier

Use case

Token type

Unit price (per million tokens)

Text

Model input

input_token_tier1_text

CNY 0.6

Image & video

Model input

input_token_tier1_image

CNY 0.6

qwen3-max (sunsetting soon)

Chinese mainland

For deployments in China regions, compute resources for model inference are limited to the Chinese mainland, and static data is stored in the selected region.

Supported regions:

Chinese mainland: China (Beijing), China (Shanghai), China (Hangzhou), China (Shenzhen), China (Ulanqab)

Tokens per inference

Scenario

Token type

Price (CNY per million tokens)

0 < tokens ≤ 32,768

model input

input_token_tier1

3

model input

(implicit cache hit)

input_token_tier1_cached

0.6

model input

(explicit cache hit)

input_token_tier1_cached_explicit

0.3

model input

(create explicit cache)

input_token_tier1_create_cache

3.75

model output (standard)

output_token_tier1

12

model output (thinking mode)

output_token_tier1_thinking

12

32,768 < tokens ≤ 131,072

model input

input_token_tier2

4.8

model input

(implicit cache hit)

input_token_tier2_cached

0.96

model input

(explicit cache hit)

input_token_tier2_cached_explicit

0.48

model input

(create explicit cache)

input_token_tier2_create_cache

6

model output (standard)

output_token_tier2

19.2

model output (thinking mode)

output_token_tier2_thinking

19.2

131,072 < tokens ≤ 258,048

model input

input_token_tier3

8.4

model input

(implicit cache hit)

input_token_tier3_cached

1.68

model input

(explicit cache hit)

input_token_tier3_cached_explicit

0.84

model input

(create explicit cache)

input_token_tier3_create_cache

10.5

model output (standard)

output_token_tier3

33.6

model output (thinking mode)

output_token_tier3_thinking

33.6

International

For deployments in International regions, compute resources for model inference are dynamically scheduled worldwide (excluding the Chinese mainland), and static data is stored in the selected region.

Supported region: Singapore

Tokens per inference

Scenario

Token type

Price (CNY per million tokens)

0 < tokens ≤ 32,768

model input

input_token_tier1

10.5684

model input

(implicit cache hit)

input_token_tier1_cached

2.11368

model input

(explicit cache hit)

input_token_tier1_cached_explicit

1.05684

model input

(create explicit cache)

input_token_tier1_create_cache

13.2105

model output (standard)

output_token_tier1

52.842

model output (thinking mode)

output_token_tier1_thinking

52.842

32,768 < tokens ≤ 131,072

model input

input_token_tier2

21.1368

model input

(implicit cache hit)

input_token_tier2_cached

4.22736

model input

(explicit cache hit)

input_token_tier2_cached_explicit

2.11368

model input

(create explicit cache)

input_token_tier2_create_cache

26.421

model output (standard)

output_token_tier2

105.6852

model output (thinking mode)

output_token_tier2_thinking

105.6852

131,072 < tokens ≤ 258,048

model input

input_token_tier3

26.4216

model input

(implicit cache hit)

input_token_tier3_cached

5.28432

model input

(explicit cache hit)

input_token_tier3_cached_explicit

2.64216

model input

(create explicit cache)

input_token_tier3_create_cache

33.027

model output (standard)

output_token_tier3

132.1068

model output (thinking mode)

output_token_tier3_thinking

132.1068

Billing

  • Billing frequency: Billed hourly.

    Due to data aggregation delays, inference fees may take several hours to appear on your bill after a job is complete. The information in Alibaba Cloud Billing Management is final.

  • View your bills

    1. Log on to the Alibaba Cloud Billing Management console.

    2. In the left navigation pane, choose Billing > Bill Details.

    3. On the Bill Details page, set Product Name to MaxCompute and Product Name to MaxCompute model computing service.

FAQ

  • Q: How do I estimate inference costs?
    A: You can estimate costs based on the pricing tables and your expected input and output token lengths. For example, a job with 1,000 inference calls using the qwen3-max model, where each call averages 2,000 input and 1,000 output tokens, costs approximately CNY 18 for the model compute service.

  • Q: Is a free tier or trial available?
    A: The model compute service is offered only on a pay-as-you-go basis and does not have a free tier. We recommend testing with a small dataset first to validate performance and estimate costs.

  • Q: Can I set a spending limit to prevent overages?
    A: No, the service does not support spending or usage limits. To control costs, monitor your usage by reviewing your bills.

Related documentation