Token-based deployment
A deployment mode billed by model token usage, supporting only LoRA fine-tuned models. Pay only for what you use. Suitable for post-tuning model validation and low-cost scenarios with low concurrency and latency requirements.
Overview
Billed by token usage (no usage, no charge), this mode supports only some LoRA efficient fine-tuned models. It is suitable for post-tuning model validation and low-cost scenarios with low concurrency and latency requirements. Throughput, concurrency, and generation speed are preset by the platform and cannot be adjusted.
NoteThe billing method cannot be changed after the service is created. To switch, you must take the deployed model offline and redeploy it. See Model deployment.
Billing rules
Fee = Model input token count × Model input unit price + Model output token count × Model output unit price (minimum billing unit: 1 token)
Billing by model token usage is supported only after you complete efficient SFT training (that is, LoRA efficient fine-tuning; the plan parameter is set to lora for API deployment) on the following base models and obtain a custom model.
Supported models and pricing
Beijing
Base Model | Model Code | Input CNY/Million Tokens | Output CNY/Million Tokens |
|---|---|---|---|
Qwen3.6-27B | qwen3.6-27b | <256K ¥3 | <256K ¥18 |
Qwen3.5-27B | qwen3.5-27b | <128K ¥0.6 128K-256K ¥1.8 | <128K ¥4.8 128K-256K ¥14.4 |
Qwen3-32B | qwen3-32b | Non-thinking mode:¥2 Thinking mode:¥2 | Non-thinking mode:¥8 Thinking mode:¥20 |
Qwen3-14B | qwen3-14b | Non-thinking mode:¥1 Thinking mode:¥1 | Non-thinking mode:¥4 Thinking mode:¥10 |
Qwen3-8B | qwen3-8b | Non-thinking mode:¥0.5 Thinking mode:¥0.5 | Non-thinking mode:¥2 Thinking mode:¥5 |
Qwen3-VL-8B-Instruct | qwen3-vl-8b-instruct | ¥0.5 | ¥2 |
Qwen3-4B-Instruct-2507 | qwen3-4b-instruct-2507 | Non-thinking mode:¥0.3 Thinking mode:¥0.3 | Non-thinking mode:¥1.2 Thinking mode:¥3 |
Qwen2.5-72B-Instruct (Open Source) | qwen2.5-72b-instruct | ¥4 | ¥12 |
Qwen2.5-VL-72B-Instruct | qwen2.5-vl-72b-instruct | ¥16 | ¥48 |
Qwen2.5-32B-Instruct (Open Source) | qwen2.5-32b-instruct | ¥2 | ¥6 |
Qwen2.5-VL-32B-Instruct | qwen2.5-vl-32b-instruct | ¥8 | ¥24 |
Qwen2.5-14B-Instruct (Open Source) | qwen2.5-14b-instruct | ¥1 | ¥3 |
Qwen2.5-7B-Instruct (Open Source) | qwen2.5-7b-instruct | ¥0.5 | ¥1 |
Qwen2.5-VL-7B-Instruct | qwen2.5-vl-7b-instruct | ¥2 | ¥5 |
Singapore
Base Model | Model Code | Input CNY/Million Tokens | Output CNY/Million Tokens |
|---|---|---|---|
Qwen3.6-27B | qwen3.6-27b | <256K ¥4.497 | <256K ¥26.979 |
Qwen3.5-27B | qwen3.5-27b | ¥2.202 | ¥17.614 |
Qwen3-32B | qwen3-32b | Non-thinking mode:¥1.174 Thinking mode:¥1.174 | Non-thinking mode:¥4.697 Thinking mode:¥4.697 |
Qwen3-14B | qwen3-14b | Non-thinking mode:¥2.569 Thinking mode:¥2.569 | Non-thinking mode:¥10.275 Thinking mode:¥30.825 |
Qwen3-8B | qwen3-8b | Non-thinking mode:¥1.321 Thinking mode:¥1.321 | Non-thinking mode:¥5.137 Thinking mode:¥15.412 |
Qwen3-VL-8B-Instruct | qwen3-vl-8b-instruct | ¥1.321 | ¥5.137 |
Qwen3-4B-Instruct-2507 | qwen3-4b-instruct-2507 | Non-thinking mode:¥0.807 Thinking mode:¥0.807 | Non-thinking mode:¥3.082 Thinking mode:¥9.247 |
Qwen2.5-72B-Instruct (Open Source) | qwen2.5-72b-instruct | ¥10.275 | ¥41.1 |
Qwen2.5-VL-72B-Instruct | qwen2.5-vl-72b-instruct | ¥20.55 | ¥61.65 |
Qwen2.5-32B-Instruct (Open Source) | qwen2.5-32b-instruct | ¥5.137 | ¥20.55 |
Qwen2.5-VL-32B-Instruct | qwen2.5-vl-32b-instruct | ¥10.275 | ¥30.825 |
Qwen2.5-14B-Instruct (Open Source) | qwen2.5-14b-instruct | ¥2.569 | ¥10.275 |
Qwen2.5-7B-Instruct (Open Source) | qwen2.5-7b-instruct | ¥1.284 | ¥5.137 |
Qwen2.5-VL-7B-Instruct | qwen2.5-vl-7b-instruct | ¥2.569 | ¥7.706 |
LoRA deployment
Token-based deployment supports only LoRA fine-tuned models. Set plan to lora when creating via API. The capacity parameter has no effect but is required; to scale, submit an application form on the Model Studio console.
For the complete example of creating a deployment via API, see API deployment guide.
Scaling
For token-usage-based deployments, scaling requires submitting an application form on the console and waiting for manual review. Self-service scaling is not supported.
FAQ
What happens if the deployment is not used for a month?
A token-usage-based deployment is automatically released if not used within a month.
How do I switch to another billing method?
You can only release the existing resources and create new ones with the desired billing method. See Model deployment.