Whether using the platform's pre-configured models or the models you have fine-tuned,you can obtain an independent, resource-dedicated inference service through deployment to meet your business needs for different performance levels such as high concurrency and low latency.
Billing methods
Before deployment, you can view the estimated hourly cost of different models in the Model Deployment console.
The billing method cannot be changed after the service is created. To switch, you must take the deployed model offline and then redeploy it.
|
Provisioned Throughput (PTU, Provisioned Throughput Unit) (High throughput; high performance) |
Model Unit (Custom performance metrics; resource isolation) |
Token-based usage (Pay-as-you-go after fine-tuning/effect validation) |
|||
|
Definition |
A model deployment method that reserves platform resources to guarantee a specific TPM throughput capacity; no rate limiting within the guaranteed quota. |
A model deployment method that configures computing power based on usage duration and the number of Model Units, with dedicated resources. |
A model deployment method that uses the input Tokens and output Tokens generated per call as the usage metering basis. |
||
|
Advantages |
|
|
No charge when not in use. |
||
|
Supported models |
Some pre-configured models |
Some pre-configured models and all fine-tuned models |
Some models fine-tuned with LoRA |
||
|
Use cases |
|
|
Fine-tuned model effect validation |
||
|
Billing diagram |
|
|
|
||
|
Billing method |
By usage duration and provisioned throughput Pay-as-you-go, daily package |
By usage duration and number of Model Units Pay-as-you-go, monthly package |
By model Token usage Pay-as-you-go |
||
|
Scaling method |
Self-service increase/decrease of throughput |
Self-service increase/decrease of Model Units |
Submit an application in the console and wait for manual review. |
||
|
Product constraints |
|
After a prepaid purchase, if you cancel early within the first month, the daily unit price (≈ monthly unit price / 30) will be billed at 1.2 times |
|
||
To view the Token usage and call count history statistics for each call, go to: Model Monitoring.
Billing details
Billing by usage duration (Provisioned Throughput)
Fee = Usage duration × (Input TPM unit price × Input TPM + Output TPM unit price × Output TPM)
Post-paid is calculated hourly: the usage duration unit is hours, and the unit price is taken from the "Continuous 1 hour" column in the table below; prepaid is calculated daily: the usage duration unit is days, and the unit price is taken from the "Continuous 1 day" column in the table below.
-
Prepaid orders take effect in real time after payment, with a validity period of N days ending at 23:59 on day N. If the order is placed after 22:00, the expiration date will be automatically extended by 1 day.
-
After a prepaid order expires, the service will be stopped with a 2-hour delay, and resources will be retained for 14 hours after the stop and then released.
-
Prepaid orders cannot terminate the service early.
-
For post-paid billing, if the account is in arrears, the deployed resources will continue to be retained and billed for 24 hours, during which the service can still be used normally. After 24 hours, the system stops billing, the model deployment enters an arrears state, and the underlying resources will be deleted, but the model deployment task will be retained. After the arrears are paid, the system will reallocate resources and restore usage (fees will continue to accrue after restoration). If you do not want to continue incurring fees, you can delete the model deployment task, and billing will stop after successful deletion.
When the model input exceeds the maximum input Token, the relevant call will automatically switch to the pay-as-you-go mode of the current model; when the purchased TPM is exceeded, it is handled according to the overflow strategy selected at creation ("auto-overflow" switches to pay-as-you-go, "use-only-PTU-capacity" returns 429). At this time, inference performance may degrade and will be subject to the public traffic control of the current snapshot model in the business space, and fees will be charged according to the model invocation (pay-as-you-go) standard.
-
In this case (only under the "auto-overflow" strategy), the API response Header will include:
x-dashscope-ptu-overflow:true. -
For TPM statistics, go to: Model Monitoring (Beijing).
For the specific fee reduction and refund rules in scale-down (downgrade) scenarios, please refer to: Refund rules for configuration downgrades.
PTU deployment supports long-input tiered capacity coefficients and cache discounts; see Provisioned Throughput long input and caching for details.
North China 2 (Beijing)
Qwen
|
Model name |
Model code |
Max input Token |
Post-paid input Per 10K TPM/hour |
Post-paid output Per 1K TPM/hour |
Prepaid input Per 10K TPM/day |
Prepaid output Per 1K TPM/day |
|
Qwen3.8-Max |
qwen3.8-max |
1M |
¥28.8 |
¥8.64 |
¥345.6 |
¥103.68 |
|
Qwen3.7-Flash-2026-07-15 Contact your business manager to activate |
qwen3.7-flash-2026-07-15 |
128K |
¥0.48 |
¥0.19 |
¥5.76 |
¥2.3 |
|
Qwen3.7-Max-2026-05-20 |
qwen3.7-max-2026-05-20 |
256K |
¥28.8 |
¥8.64 |
¥345.6 |
¥103.68 |
|
Qwen3.7-Plus-2026-05-26 |
qwen3.7-plus-2026-05-26 |
256K |
¥4.8 |
¥1.92 |
¥57.6 |
¥23.04 |
|
Qwen3.6-Plus-2026-04-02 |
qwen3.6-plus-2026-04-02 |
128K |
¥4.8 |
¥2.88 |
¥57.6 |
¥34.56 |
|
Qwen3.5-Plus-2026-04-20 |
qwen3.5-plus-2026-04-20 |
128K |
¥1.92 |
¥1.15 |
¥23.04 |
¥13.82 |
|
Qwen3-Max-2025-09-23 |
qwen3-max-2025-09-23 |
128K |
¥7.68 |
¥3.08 |
¥92.16 |
¥36.96 |
|
Qwen-Flash-2025-07-28 |
qwen-flash-2025-07-28 |
128K |
¥0.36 |
¥0.36 |
¥4.32 |
¥4.32 |
|
Qwen-Plus-2025-12-01 |
qwen-plus-2025-12-01 |
128K |
¥1.92 |
Non-thinking: ¥0.48 Thinking: ¥1.92 |
¥23.04 |
Non-thinking: ¥5.76 Thinking: ¥23.04 |
DeepSeek
|
Model Name |
Model Code |
Max Input Token |
Postpaid Input Per 10K TPM/Hour |
Postpaid Output Per 1K TPM/Hour |
Prepaid Input Per 10K TPM/Day |
Prepaid Output Per 1K TPM/Day |
|
DeepSeek-v4-Flash |
deepseek-v4-flash |
256K |
¥3.6 |
¥0.72 |
¥43.2 |
¥8.64 |
|
DeepSeek-v4-Flash-0731 |
deepseek-v4-flash-0731 |
64K |
¥7.2 |
¥1.44 |
¥86.4 |
¥17.28 |
|
DeepSeek-v4-Pro |
deepseek-v4-pro |
256K |
¥43.2 |
¥8.64 |
¥518.4 |
¥103.68 |
|
DeepSeek-v3 |
deepseek-v3 |
64K |
¥7.2 |
¥2.88 |
¥86.4 |
¥34.56 |
Qwen-VL
|
Model Name |
Model Code |
Max Input Token |
Postpaid Input Per 10K TPM/Hour |
Postpaid Output Per 1K TPM/Hour |
Prepaid Input Per 10K TPM/Day |
Prepaid Output Per 1K TPM/Day |
|
Qwen3-VL-Plus-2025-09-23 |
qwen3-vl-plus-2025-09-23 |
128K |
¥2.4 |
¥2.4 |
¥28.8 |
¥28.8 |
GLM
|
Model Name |
Model Code |
Max Input Token |
Postpaid Input Per 10K TPM/Hour |
Postpaid Output Per 1K TPM/Hour |
Prepaid Input Per 10K TPM/Day |
Prepaid Output Per 1K TPM/Day |
|
GLM-5.2 |
glm-5.2 |
1M |
¥28.8 |
¥10.08 |
¥345.6 |
¥120.96 |
Singapore
Qwen
|
Model Name |
Model Code |
Max Input Token |
Postpaid Input Per 10K TPM/Hour |
Postpaid Output Per 1K TPM/Hour |
Prepaid Input Per 10K TPM/Day |
Prepaid Output Per 1K TPM/Day |
|
Qwen3.8-Max |
qwen3.8-max |
1M |
¥35.97 |
¥10.79 |
¥431.7 |
¥129.5 |
|
Qwen3.7-Flash-2026-07-15 Contact your business manager to activate |
qwen3.7-flash-2026-07-15 |
128K |
¥0.54 |
¥0.23 |
¥6.47 |
¥2.81 |
|
Qwen3.7-Max-2026-05-20 |
qwen3.7-max-2026-05-20 |
256K |
¥44.97 |
¥13.49 |
¥539.6 |
¥161.87 |
|
Qwen3.7-Plus-2026-05-26 |
qwen3.7-plus-2026-05-26 |
256K |
¥7.19 |
¥2.88 |
¥86.3 |
¥34.53 |
|
Qwen3.6-Plus-2026-04-02 |
qwen3.6-plus-2026-04-02 |
128K |
¥9 |
¥5.4 |
¥107.9 |
¥64.75 |
|
Qwen3.5-Plus-2026-04-20 |
qwen3.5-plus-2026-04-20 |
128K |
¥7.2 |
¥4.32 |
¥86.3 |
¥51.8 |
DeepSeek
|
Model Name |
Model Code |
Max Input Token |
Postpaid Input Per 10K TPM/Hour |
Postpaid Output Per 1K TPM/Hour |
Prepaid Input Per 10K TPM/Day |
Prepaid Output Per 1K TPM/Day |
|
DeepSeek-v4-Flash |
deepseek-v4-flash |
256K |
¥5.4 |
¥1.08 |
¥64.8 |
¥12.95 |
|
DeepSeek-v4-Flash-0731 |
deepseek-v4-flash-0731 |
64K |
¥10.79 |
¥2.16 |
¥129.5 |
¥25.9 |
|
DeepSeek-v4-Pro |
deepseek-v4-pro |
256K |
¥64.75 |
¥12.95 |
¥777 |
¥155.4 |
Qwen-VL
|
Model Name |
Model Code |
Max Input Token |
Postpaid Input Per 10K TPM/Hour |
Postpaid Output Per 1K TPM/Hour |
Prepaid Input Per 10K TPM/Day |
Prepaid Output Per 1K TPM/Day |
|
Qwen3-VL-Plus-2025-09-23 |
qwen3-vl-plus-2025-09-23 |
128K |
¥3.6 |
¥2.88 |
¥43.2 |
¥34.53 |
GLM
|
Model Name |
Model Code |
Max Input Token |
Postpaid Input Per 10K TPM/Hour |
Postpaid Output Per 1K TPM/Hour |
Prepaid Input Per 10K TPM/Day |
Prepaid Output Per 1K TPM/Day |
|
GLM-5.2 |
glm-5.2 |
1M |
¥37.8 |
¥11.87 |
¥453.3 |
¥142.45 |
Billing by usage duration (Model Unit)
Cost = Usage duration (hours) × Number of Model Units × Model Unit unit price
The "Model Unit unit price" takes the "Hourly unit price" column in the table below for postpaid scenarios; for monthly subscription prepaid billing, the formula becomes Number of monthly subscriptions × Number of Model Units × Monthly unit price.
-
For the first month of a prepaid purchase, if you cancel the subscription early within the first month, the daily unit price (≈ Monthly unit price / 30) will be billed at 1.2 times the rate (less than one day is billed as one day)
The compute resources for the Model Unit postpaid method are first-come, first-served. If the purchase fails, a full refund will be issued.
Text generation
Qwen
|
Model Name |
Model Code |
Model Unit specification |
Hourly unit price (CNY) Minimum billing: minute |
Monthly subscription unit price (CNY) Minimum billing: day |
|
Qwen3.6-35B-A3B |
qwen3.6-35b-a3b |
MU1 x 8 |
¥432 |
¥208,944 |
|
MU2 x 8 |
¥504 |
¥240,288 |
||
|
MU3 x 8 |
¥1,096 |
¥527,752 |
||
|
MU8 x 1 |
¥47 |
¥22,400 |
||
|
MU9 x 1 |
¥51 |
¥24,600 |
||
|
Qwen3.6-27B |
qwen3.6-27b |
MU9 x 1 |
¥51 |
¥24,600 |
|
Qwen3.6-Flash-2026-04-16 |
qwen3.6-flash-2026-04-16 |
MU1 x 2 |
¥108 |
¥52,236 |
|
MU3 x 8 |
¥1,096 |
¥527,752 |
||
|
Qwen3.6-Plus-2026-04-02 |
qwen3.6-plus-2026-04-02 |
MU1 x 8 MU1 x 16 (PD disaggregation mode) |
¥432 PD disaggregation mode: ¥864 |
¥208,944 PD disaggregation mode: ¥417,888 |
|
Qwen3.5-397B-A17B |
qwen3.5-397b-a17b |
MU3 x 8 MU3 x 16 (PD disaggregation mode) |
¥1,096 PD disaggregation mode: ¥2,192 |
¥527,752 PD disaggregation mode: ¥1,055,504 |
|
MU6 x 16 |
¥400 |
¥193,424 |
||
|
Qwen3.5-122B-A10B |
qwen3.5-122b-a10b |
MU1 x 4 |
¥216 |
¥104,472 |
|
MU6 x 16 |
¥400 |
¥193,424 |
||
|
Qwen3.5-35B-A3B |
qwen3.5-35b-a3b |
MU1 x 2 |
¥108 |
¥52,236 |
|
MU2 x 8 |
¥504 |
¥240,288 |
||
|
MU3 x 8 |
¥1,096 |
¥527,752 |
||
|
MU9 x 1 |
¥51 |
¥24,600 |
||
|
Qwen3.5-27B |
qwen3.5-27b |
MU2 x 8 |
¥504 |
¥240,288 |
|
MU3 x 8 |
¥1,096 |
¥527,752 |
||
|
MU8 x 1 |
¥47 |
¥22,400 |
||
|
MU9 x 1 |
¥51 |
¥24,600 |
||
|
Qwen3.5-9B |
qwen3.5-9b |
MU1 x 2 |
¥108 |
¥52,236 |
|
MU2 x 8 |
¥504 |
¥240,288 |
||
|
Qwen3.5-Flash-2026-02-23 |
qwen3.5-flash-2026-02-23 |
MU1 x 2 |
¥108 |
¥52,236 |
|
Qwen3.5-Plus-2026-02-15 |
qwen3.5-plus-2026-02-15 |
MU1 x 8 MU1 x 16 (PD disaggregation mode) |
¥432 PD disaggregation mode: ¥864 |
¥208,944 PD disaggregation mode: ¥417,888 |
|
MU2 x 8 |
¥504 |
¥240,288 |
||
|
MU3 x 8 MU3 x 16 (PD disaggregation mode) |
¥1,096 PD disaggregation mode: ¥2,192 |
¥527,752 PD disaggregation mode: ¥1,055,504 |
||
|
Qwen3-235B-A22B-Instruct-2507 |
qwen3-235b-a22b-instruct-2507 |
MU1 x 4 |
¥216 |
¥104,472 |
|
MU2 x 8 |
¥504 |
¥240,288 |
||
|
MU3 x 8 |
¥1,096 |
¥527,752 |
||
|
Qwen3-32B |
qwen3-32b |
MU6 x 16 |
¥400 |
¥193,424 |
|
Qwen3-30B-A3B-Thinking-2507 |
qwen3-30b-a3b-thinking-2507 |
MU1 x 2 |
¥108 |
¥52,236 |
|
Qwen3-8B |
qwen3-8b |
MU1 x 2 |
¥108 |
¥52,236 |
|
MU2 x 2 |
¥126 |
¥60,072 |
||
|
Qwen3-4B |
qwen3-4b |
MU1 x 2 |
¥108 |
¥52,236 |
|
MU5 x 1 |
¥21 |
¥10,139 |
||
|
Qwen3-Embedding-0.6B |
qwen3-embedding-0.6b |
MU5 x 1 |
¥21 |
¥10,139 |
|
MU6 x 1 |
¥25 |
¥12,089 |
||
|
Qwen3-MoE-Rerank-0.6B |
qwen3-moe-rerank-0.6b |
MU5 x 1 |
¥21 |
¥10,139 |
|
Qwen3-Rerank-0.6B |
qwen3-rerank-0.6b |
MU5 x 1 |
¥21 |
¥10,139 |
|
MU6 x 1 |
¥25 |
¥12,089 |
||
|
Qwen3-Max-2025-09-23 |
qwen3-max-2025-09-23 |
MU2 x 8 |
¥504 |
¥240,288 |
|
MU3 x 8 |
¥1,096 |
¥527,752 |
||
|
Qwen3-Rerank |
qwen3-rerank |
MU5 x 1 |
¥21 |
¥10,139 |
|
Qwen2.5-Open-Source-72B |
qwen2.5-72b-instruct |
MU1 x 8 |
¥432 |
¥208,944 |
|
Qwen2.5-Open-Source-14B |
qwen2.5-14b-instruct |
MU1 x 2 |
¥108 |
¥52,236 |
|
Qwen2.5-Open-Source-7B |
qwen2.5-7b-instruct |
MU1 x 2 |
¥108 |
¥52,236 |
|
MU5 x 1 |
¥21 |
¥10,139 |
||
|
Qwen-Plus-2025-07-28 |
qwen-plus-2025-07-28 |
MU1 x 4 MU1 x 16 (PD disaggregation mode) |
¥216 PD disaggregation mode: ¥864 |
¥104,472 PD disaggregation mode: ¥417,888 |
|
Qwen-Plus-2025-12-01 |
qwen-plus-2025-12-01 |
MU1 x 4 |
¥216 |
¥104,472 |
|
Qwen-Plus-Character-2025-11-06 |
qwen-plus-character-2025-11-06 |
MU1 x 4 |
¥216 |
¥104,472 |
GLM
|
Model name |
Model code |
Model unit specification |
Hourly unit price (CNY) Minimum billing: minute |
Monthly subscription unit price (CNY) Minimum billing: day |
|
GLM-5.1 |
glm-5.1 |
MU2 x 8 |
¥504 |
¥240,288 |
|
MU3 x 16 (PD disaggregation mode) |
PD disaggregation mode: ¥2,192 |
PD disaggregation mode: ¥1,055,504 |
||
|
MU6 x 16 |
¥400 |
¥193,424 |
||
|
GLM-5 |
glm-5 |
MU3 x 16 (PD disaggregation mode) |
PD disaggregation mode: ¥2,192 |
PD disaggregation mode: ¥1,055,504 |
|
GLM-4.7 |
glm-4.7 |
MU6 x 32 (PD disaggregation mode) |
PD disaggregation mode: ¥800 |
PD disaggregation mode: ¥386,848 |
DeepSeek
|
Model name |
Model code |
Model unit specification |
Hourly unit price (CNY) Minimum billing: minute |
Monthly subscription unit price (CNY) Minimum billing: day |
|
DeepSeek-v4-Flash |
deepseek-v4-flash |
MU1 x 8 |
¥432 |
¥208,944 |
|
MU3 x 8 |
¥1,096 |
¥527,752 |
||
|
DeepSeek-v3.2 |
deepseek-v3.2 |
MU2 x 16 (PD disaggregation mode) |
PD disaggregation mode: ¥1,008 |
PD disaggregation mode: ¥480,576 |
More models
|
Model name |
Model code |
Model unit specification |
Hourly unit price (CNY) Minimum billing: minute |
Monthly subscription unit price (CNY) Minimum billing: day |
|
Kimi-K2.5 |
kimi-k2.5 |
MU2 x 8 |
¥504 |
¥240,288 |
Model type:
-
Instruct - The model performs inference in non-thinking mode after deployment.
-
Thinking - The model performs inference in thinking mode after deployment.
Model deployment type:
-
PD disaggregation mode - Reduces first-token latency and increases throughput.
For models deployed in this mode, during inference the first-token computation (Prefill) and the subsequent token computation (Decode) are split across different compute nodes for execution.
Multimodal
Qwen-VL
|
Model name |
Model code |
Model unit specification |
Hourly unit price (CNY) Minimum billing: minute |
Monthly subscription unit price (CNY) Minimum billing: day |
|
Qwen3-VL-235B-A22B-Thinking |
qwen3-vl-235b-a22b-thinking |
MU1 x 8 |
¥432 |
¥208,944 |
|
MU2 x 8 |
¥504 |
¥240,288 |
||
|
MU3 x 8 |
¥1,096 |
¥527,752 |
||
|
Qwen3-VL-32B-Instruct |
qwen3-vl-32b-instruct |
MU2 x 8 |
¥504 |
¥240,288 |
|
MU3 x 8 |
¥1,096 |
¥527,752 |
||
|
Qwen3-VL-8B-Instruct |
qwen3-vl-8b-instruct |
MU1 x 2 |
¥108 |
¥52,236 |
|
MU5 x 1 |
¥21 |
¥10,139 |
||
|
Qwen3-VL-4B-Instruct |
qwen3-vl-4b-instruct |
MU1 x 2 |
¥108 |
¥52,236 |
|
Qwen3-VL-2B-Instruct |
qwen3-vl-2b-instruct |
MU5 x 1 |
¥21 |
¥10,139 |
|
Qwen3-VL-Embedding-2B |
qwen3-vl-embedding-2b |
MU5 x 1 |
¥21 |
¥10,139 |
|
Qwen3-VL-Flash-2025-10-15 |
qwen3-vl-flash-2025-10-15 |
MU1 x 4 |
¥216 |
¥104,472 |
|
Qwen3-VL-Plus-2025-09-23 |
qwen3-vl-plus-2025-09-23 |
MU1 x 4 |
¥216 |
¥104,472 |
|
QwenVL-Max-2025-08-13 |
qwen-vl-max-2025-08-13 |
MU6 x 4 |
¥100 |
¥48,356 |
Qwen Omni
|
Model Name |
Model Code |
Model Unit Specification |
Hourly Unit Price (CNY) Minimum Billing: Minute |
Monthly Subscription Unit Price (CNY) Minimum Billing: Day |
|
Qwen3.5-Omni-Flash |
qwen3.5-omni-flash |
MU8 x 1 |
¥47 |
¥22,400 |
|
MU9 x 1 |
¥51 |
¥24,600 |
Model Type:
-
Instruct - The model performs inference in non-thinking mode after deployment.
-
Thinking - The model performs inference in thinking mode after deployment.
-
Instruct/Thinking - You can choose whether to enable thinking mode when deploying the model.
Speech Synthesis
CosyVoice
|
Model Name |
Model Code |
Model Unit Specification |
Hourly Unit Price (CNY) |
Monthly Subscription Unit Price (CNY) |
|
cosyvoice-v3-flash |
cosyvoice-v3-flash |
MU5 |
¥21 |
¥10,139 |
By model token usage
Fee = Model input token count × Model input unit price + Model output token count × Model output unit price (minimum billing unit: 1 token)
-
Billing by model token usage is supported only after you complete efficient SFT training (that is, LoRA efficient fine-tuning; the plan parameter is set to lora for API deployment) on the following base models and obtain a custom model.
Beijing
|
Base model |
Model code |
Input CNY/Million tokens |
Output CNY/Million tokens |
|
Qwen3.5-27B In preview |
qwen3.5-27b |
<128K ¥0.6 128K-256K ¥1.8 |
<128K ¥4.8 128K-256K ¥14.4 |
|
Qwen3-32B |
qwen3-32b |
Non-thinking mode: ¥2 Thinking mode: ¥2 |
Non-thinking mode: ¥8 Thinking mode: ¥20 |
|
Qwen3-14B |
qwen3-14b |
Non-thinking mode: ¥1 Thinking mode: ¥1 |
Non-thinking mode: ¥4 Thinking mode: ¥10 |
|
Qwen3-8B |
qwen3-8b |
Non-thinking mode: ¥0.5 Thinking mode: ¥0.5 |
Non-thinking mode: ¥2 Thinking mode: ¥5 |
|
Qwen3-VL-8B-Instruct |
qwen3-vl-8b-instruct |
¥0.5 |
¥2 |
Singapore
|
Base model |
Model code |
Input CNY/Million tokens |
Output CNY/Million tokens |
|
Qwen3-14B |
qwen3-14b |
Non-thinking mode: ¥2.569 Thinking mode: ¥2.569 |
Non-thinking mode: ¥10.275 Thinking mode: ¥30.825 |
To deploy more models, refer to thissolution and select the most suitable deployment plan based on your business requirements.
Deployment methods
You can deploy models on the console. Refer to the following steps:
If you are prompted with insufficient permissions, refer to:What should I do if "insufficient permissions" is prompted during deployment?
|
|
|
|
Important
Fees will be incurred after the model is successfully deployed. |
Deployment configuration
Model Unit
|
Configuration item |
Configuration details |
|
Service name |
A custom name for the deployment service. |
|
Model |
Select the model to deploy, including platform preset models and fine-tuned models. |
|
Model unit type |
Select the deployment specification. Different specifications correspond to different computing power and performance. |
|
Replica count |
Set the initial number of deployment replicas, which affects the concurrent processing capability of the service. |
|
Deployment template |
Select a deployment template (for example, "single-node deployment"). Different templates correspond to different resource configuration schemes. Available only in the model unit billing mode. |
|
Model inference mode |
For some models, when deployed inModel Unit mode, you can configure the inference mode, maximum context, and more.
|
|
Maximum context |
TheModel Unit deployment mode of some models supports this setting. The maximum context length depends on the model type. |
|
Service throttling |
TheModel Unit deployment mode of some models supports this setting, which can limit the RPM and TPM of model calls. |
Deployment list page
After successful deployment, you can view and manage all deployment services on the deployment list page. The list page contains the following information:
-
Service name: The name of the deployment service. Click to view deployment details.
-
Model name: The model used for deployment.
-
Model Code: The unique identifier generated after the model is successfully deployed, used to specify the model when calling the API.
-
Deployment status/Event status: Includes Pending deployment, Deploying, Running, Deployment failed, Going offline, Service paused, Stopped, Deleting, Subscription suspended/Overdue payment suspended, Resuming service, Running (Changing configuration), Running (Change failed), and other statuses.
-
Billing method: The billing method of the current deployment service.
-
Deployment details: Configuration information such as model unit type and replica count.
-
Throttling details: Displays the throttling configuration of the current deployment service, such as RPM (requests per minute) and TPM (tokens per minute).
-
Service time: Displays the creation time and expiration time of the deployment service.
-
Operation: Depending on the deployment status and billing method, you can perform operations such as Update, Monitor, Scale, Renew, Take offline, Delete, and Try.
Post-deployment calls
After the model is successfully deployed, you can call it through OpenAI-compatible, Dashscope, and Assistant SDK.
When calling a successfully deployed model, the value of model should be the model code generated after successful deployment. Go to themodel deployment console (Beijing) to obtain the Model Code.

The following sample code calls the fine-tuned qwen3-8b model as an example:
Model features (whether non-streaming output, structured output, etc. are supported) are consistent with themodel before fine-tuning.
For deep thinking models that have been fine-tuned, whether to enable deep thinking during calls is recommended to be consistent with the fine-tuning data format:
-
If the fine-tuning data contains deep thinking, it is recommended to enable the
enable_thinkingparameter when calling. -
If the fine-tuning data does not contain deep thinking, it is not recommended to enable the
enable_thinkingparameter when calling.
For GLM-5.2 models deployed with the preset throughput deployment method, the thinking_budget parameter (which limits the thinking length) does not take effect when called.
DashScope
import os
import dashscope
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Who are you?"},
]
response = dashscope.Generation.call(
# If you have not configured environment variables, replace the next line with your Bailian API Key: api_key="sk-xxx",
api_key=os.getenv("DASHSCOPE_API_KEY"),
model="qwen3-14b-xxx-xxx", # Please replace with the code returned after the model is successfully deployed
messages=messages,
result_format="message",
enable_thinking=False,
)
print(response)
OpenAI-compatible interface
import os
from openai import OpenAI
client = OpenAI(
# If you have not configured environment variables, replace the next line with your Bailian API Key: api_key="sk-xxx",
api_key=os.getenv('DASHSCOPE_API_KEY'),
base_url="https://dashscope.aliyuncs.com/compatible-mode/v1",
)
completion = client.chat.completions.create(
model="qwen3-14b-xxx-xxx", # Please replace with the code returned after the model is successfully deployed
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Who are you?"},
],
extra_body={"enable_thinking": False},
)
print(completion)
Scaling a Deployment Service
-
Preset throughput (billed by duration): Click the Scaling button to self-service, manually adjust the number of instances. For detailed fee reduction and refund rules, please refer to: Refund rules for configuration downgrades.
-
Model unit (billed by duration): Click the Scaling button to self-service, manually adjust the number of instances.
-
Billed by Token usage: Click the Scale Out button, fill in and submit the scale-out application form, and wait for manual review.
In addition, you can configure an auto-scaling policy (including scaling thresholds, minimum/maximum replica count, scheduled scaling, etc.) via the scaling configuration button in the operation column.
Take a Deployment Service Offline
Go to the Model Deployment console, find the deployment service you want to stop, and click the corresponding operation according to the billing type:
-
Model unit prepaid: Click Deactivate and confirm.
-
Postpaid: Click Delete and confirm.
No further billing will be incurred after the operation is completed.

Other Operations
In addition to going offline, the operation column on the deployment list page also supports the following operations:
-
Update: Update the model version of the deployed service, supporting full update or batch update (canary release).
-
Delete: Pay-as-you-go services can be deleted directly to stop billing.
-
Renew: Prepaid services can be renewed to extend the service time, and auto-renewal is supported.
-
Buy capacity package: Purchase a capacity package for the preset throughput deployment.
FAQ
Can I upload and deploy my own models?
You can import some open-source models in the My Models console (Beijing); for the detailed supported list, please refer to: Model import.
In addition, Alibaba Cloud Artificial Intelligence Platform PAI provides the ability to deploy your own models. You can refer to PAI-LLM Large Language Model Deployment to learn about the deployment method.
What should I do if "insufficient permissions" is prompted during deployment?
-
If "Missing permission for this module" is displayed, please ensure that your account has the Model Deployment - Operation permission on the permission management page of the business space.

If you cannot operate normally, please contact your organization or IT administrator to add the relevant permissions or check the permission issues on your behalf.
-
If the error "xx business space does not have permission to deploy the xx model" is reported during deployment, please go to the Business Space Management page of Model Studio to add the deployment permission of the corresponding model for the corresponding business space.
API call error:
Workspace xxx does not have deployment privilege for model xxxx.

If insufficient permissions are prompted, please contact your organization or IT administrator to add the relevant permissions or operate on your behalf.
How do I switch to other billing methods?
You can only release the original resources and then create new resources using the desired billing method.
It is recommended to switch according to the following steps:
-
Deploy new resources using the desired billing method.
-
Switch the API and test the service availability.
-
Take offline and release the original resources.







