Token-based deployment

Updated at:

A deployment mode billed by model token usage, supporting only LoRA fine-tuned models. Pay only for what you use. Suitable for post-tuning model validation and low-cost scenarios with low concurrency and latency requirements.

Overview

Billed by token usage (no usage, no charge), this mode supports only some LoRA efficient fine-tuned models. It is suitable for post-tuning model validation and low-cost scenarios with low concurrency and latency requirements. Throughput, concurrency, and generation speed are preset by the platform and cannot be adjusted.

NoteThe billing method cannot be changed after the service is created. To switch, you must take the deployed model offline and redeploy it. See Model deployment.

Billing rules

Fee = Model input token count × Model input unit price + Model output token count × Model output unit price (minimum billing unit: 1 token)

Billing by model token usage is supported only after you complete efficient SFT training (that is, LoRA efficient fine-tuning; the plan parameter is set to lora for API deployment) on the following base models and obtain a custom model.

Supported models and pricing

Beijing

Base Model

Model Code

Input

CNY/Million Tokens

Output

CNY/Million Tokens

Qwen3.6-27B

qwen3.6-27b

<256K ¥3

<256K ¥18

Qwen3.5-27B

qwen3.5-27b

<128K ¥0.6

128K-256K ¥1.8

<128K ¥4.8

128K-256K ¥14.4

Qwen3-32B

qwen3-32b

Non-thinking mode:¥2

Thinking mode:¥2

Non-thinking mode:¥8

Thinking mode:¥20

Qwen3-14B

qwen3-14b

Non-thinking mode:¥1

Thinking mode:¥1

Non-thinking mode:¥4

Thinking mode:¥10

Qwen3-8B

qwen3-8b

Non-thinking mode:¥0.5

Thinking mode:¥0.5

Non-thinking mode:¥2

Thinking mode:¥5

Qwen3-VL-8B-Instruct

qwen3-vl-8b-instruct

¥0.5

¥2

Qwen3-4B-Instruct-2507

qwen3-4b-instruct-2507

Non-thinking mode:¥0.3

Thinking mode:¥0.3

Non-thinking mode:¥1.2

Thinking mode:¥3

Qwen2.5-72B-Instruct (Open Source)

qwen2.5-72b-instruct

¥4

¥12

Qwen2.5-VL-72B-Instruct

qwen2.5-vl-72b-instruct

¥16

¥48

Qwen2.5-32B-Instruct (Open Source)

qwen2.5-32b-instruct

¥2

¥6

Qwen2.5-VL-32B-Instruct

qwen2.5-vl-32b-instruct

¥8

¥24

Qwen2.5-14B-Instruct (Open Source)

qwen2.5-14b-instruct

¥1

¥3

Qwen2.5-7B-Instruct (Open Source)

qwen2.5-7b-instruct

¥0.5

¥1

Qwen2.5-VL-7B-Instruct

qwen2.5-vl-7b-instruct

¥2

¥5

Singapore

Base Model

Model Code

Input

CNY/Million Tokens

Output

CNY/Million Tokens

Qwen3.6-27B

qwen3.6-27b

<256K ¥4.497

<256K ¥26.979

Qwen3.5-27B

qwen3.5-27b

¥2.202

¥17.614

Qwen3-32B

qwen3-32b

Non-thinking mode:¥1.174

Thinking mode:¥1.174

Non-thinking mode:¥4.697

Thinking mode:¥4.697

Qwen3-14B

qwen3-14b

Non-thinking mode:¥2.569

Thinking mode:¥2.569

Non-thinking mode:¥10.275

Thinking mode:¥30.825

Qwen3-8B

qwen3-8b

Non-thinking mode:¥1.321

Thinking mode:¥1.321

Non-thinking mode:¥5.137

Thinking mode:¥15.412

Qwen3-VL-8B-Instruct

qwen3-vl-8b-instruct

¥1.321

¥5.137

Qwen3-4B-Instruct-2507

qwen3-4b-instruct-2507

Non-thinking mode:¥0.807

Thinking mode:¥0.807

Non-thinking mode:¥3.082

Thinking mode:¥9.247

Qwen2.5-72B-Instruct (Open Source)

qwen2.5-72b-instruct

¥10.275

¥41.1

Qwen2.5-VL-72B-Instruct

qwen2.5-vl-72b-instruct

¥20.55

¥61.65

Qwen2.5-32B-Instruct (Open Source)

qwen2.5-32b-instruct

¥5.137

¥20.55

Qwen2.5-VL-32B-Instruct

qwen2.5-vl-32b-instruct

¥10.275

¥30.825

Qwen2.5-14B-Instruct (Open Source)

qwen2.5-14b-instruct

¥2.569

¥10.275

Qwen2.5-7B-Instruct (Open Source)

qwen2.5-7b-instruct

¥1.284

¥5.137

Qwen2.5-VL-7B-Instruct

qwen2.5-vl-7b-instruct

¥2.569

¥7.706

LoRA deployment

Token-based deployment supports only LoRA fine-tuned models. Set plan to lora when creating via API. The capacity parameter has no effect but is required; to scale, submit an application form on the Model Studio console.

For the complete example of creating a deployment via API, see API deployment guide.

Scaling

For token-usage-based deployments, scaling requires submitting an application form on the console and waiting for manual review. Self-service scaling is not supported.

FAQ

What happens if the deployment is not used for a month?

A token-usage-based deployment is automatically released if not used within a month.

How do I switch to another billing method?

You can only release the existing resources and create new ones with the desired billing method. See Model deployment.