Model deployment

更新时间:
复制 MD 格式

Whether using the platform's pre-configured models or the models you have fine-tuned,you can obtain an independent, resource-dedicated inference service through deployment to meet your business needs for different performance levels such as high concurrency and low latency.

Billing methods

Before deployment, you can view the estimated hourly cost of different models in the Model Deployment console.
Note

The billing method cannot be changed after the service is created. To switch, you must take the deployed model offline and then redeploy it.

Provisioned Throughput (PTU, Provisioned Throughput Unit)

(High throughput; high performance)

Model Unit

(Custom performance metrics; resource isolation)

Token-based usage

(Pay-as-you-go after fine-tuning/effect validation)

Definition

A model deployment method that reserves platform resources to guarantee a specific TPM throughput capacity; no rate limiting within the guaranteed quota.

A model deployment method that configures computing power based on usage duration and the number of Model Units, with dedicated resources.

A model deployment method that uses the input Tokens and output Tokens generated per call as the usage metering basis.

Advantages

  1. Provides stable throughput capacity, lower latency, and stronger resource certainty for high-load production environments.

  2. Compared with Token-based billing, TPS (Tokens generated per second) typically increases by approximately 1.5 to 2.0 times.

  3. Supports auto-renewal settings.

  1. Performance metrics such as latency/throughput can be customized.

  2. Supports auto-renewal settings.

  3. Supports PD disaggregated computing mode.

No charge when not in use.

Supported models

Some pre-configured models

Some pre-configured models and all fine-tuned models

Some models fine-tuned with LoRA

Use cases

  1. Intelligent customer service for banking apps (stable traffic, requires guaranteed concurrent experience).

  2. Real-time content moderation for social platforms (requires stable processing of predictable pipeline tasks).

  3. Public cloud translation API (provides baseline service guarantees for standard package users).

  1. E-commerce exclusive fine-tuned large models (deploy private models, manually scale up during major promotions).

  2. Pharmaceutical company molecular screening models (require dedicated resources for long-running tasks).

  3. Autonomous driving simulation (requires long-duration continuous computing).

Fine-tuned model effect validation

Billing diagram

image

image

image

Billing method

By usage duration and provisioned throughput

Pay-as-you-go, daily package

By usage duration and number of Model Units

Pay-as-you-go, monthly package

By model Token usage

Pay-as-you-go

Scaling method

Self-service increase/decrease of throughput

Self-service increase/decrease of Model Units

Submit an application in the console and wait for manual review.

Product constraints

  1. Prepaid billed daily. No early refund available.

  2. If usage within a unit time exceeds the purchased throughput, it is handled according to the overflow strategy selected at creation: auto-overflow switches to model invocation pay-as-you-go billing for that model, while using-only-PTU-capacity returns 429.

After a prepaid purchase, if you cancel early within the first month, the daily unit price (≈ monthly unit price / 30) will be billed at 1.2 times

  1. Only supports some models after efficient fine-tuning (LoRA).

  2. Will be automatically released if not used within one month.

To view the Token usage and call count history statistics for each call, go to: Model Monitoring.

Billing details

Billing by usage duration (Provisioned Throughput)

Fee = Usage duration × (Input TPM unit price × Input TPM + Output TPM unit price × Output TPM)

Post-paid is calculated hourly: the usage duration unit is hours, and the unit price is taken from the "Continuous 1 hour" column in the table below; prepaid is calculated daily: the usage duration unit is days, and the unit price is taken from the "Continuous 1 day" column in the table below.

  • Prepaid orders take effect in real time after payment, with a validity period of N days ending at 23:59 on day N. If the order is placed after 22:00, the expiration date will be automatically extended by 1 day.

  • After a prepaid order expires, the service will be stopped with a 2-hour delay, and resources will be retained for 14 hours after the stop and then released.

  • Prepaid orders cannot terminate the service early.

  • For post-paid billing, if the account is in arrears, the deployed resources will continue to be retained and billed for 24 hours, during which the service can still be used normally. After 24 hours, the system stops billing, the model deployment enters an arrears state, and the underlying resources will be deleted, but the model deployment task will be retained. After the arrears are paid, the system will reallocate resources and restore usage (fees will continue to accrue after restoration). If you do not want to continue incurring fees, you can delete the model deployment task, and billing will stop after successful deletion.

When the model input exceeds the maximum input Token, the relevant call will automatically switch to the pay-as-you-go mode of the current model; when the purchased TPM is exceeded, it is handled according to the overflow strategy selected at creation ("auto-overflow" switches to pay-as-you-go, "use-only-PTU-capacity" returns 429). At this time, inference performance may degrade and will be subject to the public traffic control of the current snapshot model in the business space, and fees will be charged according to the model invocation (pay-as-you-go) standard.

  • In this case (only under the "auto-overflow" strategy), the API response Header will include: x-dashscope-ptu-overflow:true.

  • For TPM statistics, go to: Model Monitoring (Beijing).

For the specific fee reduction and refund rules in scale-down (downgrade) scenarios, please refer to: Refund rules for configuration downgrades.

Note

PTU deployment supports long-input tiered capacity coefficients and cache discounts; see Provisioned Throughput long input and caching for details.

North China 2 (Beijing)

Qwen

Model name

Model code

Max input Token

Post-paid input

Per 10K TPM/hour

Post-paid output

Per 1K TPM/hour

Prepaid input

Per 10K TPM/day

Prepaid output

Per 1K TPM/day

Qwen3.8-Max

qwen3.8-max

1M

¥28.8

¥8.64

¥345.6

¥103.68

Qwen3.7-Flash-2026-07-15 Contact your business manager to activate

qwen3.7-flash-2026-07-15

128K

¥0.48

¥0.19

¥5.76

¥2.3

Qwen3.7-Max-2026-05-20

qwen3.7-max-2026-05-20

256K

¥28.8

¥8.64

¥345.6

¥103.68

Qwen3.7-Plus-2026-05-26

qwen3.7-plus-2026-05-26

256K

¥4.8

¥1.92

¥57.6

¥23.04

Qwen3.6-Plus-2026-04-02

qwen3.6-plus-2026-04-02

128K

¥4.8

¥2.88

¥57.6

¥34.56

Qwen3.5-Plus-2026-04-20

qwen3.5-plus-2026-04-20

128K

¥1.92

¥1.15

¥23.04

¥13.82

Qwen3-Max-2025-09-23

qwen3-max-2025-09-23

128K

¥7.68

¥3.08

¥92.16

¥36.96

Qwen-Flash-2025-07-28

qwen-flash-2025-07-28

128K

¥0.36

¥0.36

¥4.32

¥4.32

Qwen-Plus-2025-12-01

qwen-plus-2025-12-01

128K

¥1.92

Non-thinking: ¥0.48

Thinking: ¥1.92

¥23.04

Non-thinking: ¥5.76

Thinking: ¥23.04

DeepSeek

Model Name

Model Code

Max Input Token

Postpaid Input

Per 10K TPM/Hour

Postpaid Output

Per 1K TPM/Hour

Prepaid Input

Per 10K TPM/Day

Prepaid Output

Per 1K TPM/Day

DeepSeek-v4-Flash

deepseek-v4-flash

256K

¥3.6

¥0.72

¥43.2

¥8.64

DeepSeek-v4-Flash-0731

deepseek-v4-flash-0731

64K

¥7.2

¥1.44

¥86.4

¥17.28

DeepSeek-v4-Pro

deepseek-v4-pro

256K

¥43.2

¥8.64

¥518.4

¥103.68

DeepSeek-v3

deepseek-v3

64K

¥7.2

¥2.88

¥86.4

¥34.56

Qwen-VL

Model Name

Model Code

Max Input Token

Postpaid Input

Per 10K TPM/Hour

Postpaid Output

Per 1K TPM/Hour

Prepaid Input

Per 10K TPM/Day

Prepaid Output

Per 1K TPM/Day

Qwen3-VL-Plus-2025-09-23

qwen3-vl-plus-2025-09-23

128K

¥2.4

¥2.4

¥28.8

¥28.8

GLM

Model Name

Model Code

Max Input Token

Postpaid Input

Per 10K TPM/Hour

Postpaid Output

Per 1K TPM/Hour

Prepaid Input

Per 10K TPM/Day

Prepaid Output

Per 1K TPM/Day

GLM-5.2

glm-5.2

1M

¥28.8

¥10.08

¥345.6

¥120.96

Singapore

Qwen

Model Name

Model Code

Max Input Token

Postpaid Input

Per 10K TPM/Hour

Postpaid Output

Per 1K TPM/Hour

Prepaid Input

Per 10K TPM/Day

Prepaid Output

Per 1K TPM/Day

Qwen3.8-Max

qwen3.8-max

1M

¥35.97

¥10.79

¥431.7

¥129.5

Qwen3.7-Flash-2026-07-15 Contact your business manager to activate

qwen3.7-flash-2026-07-15

128K

¥0.54

¥0.23

¥6.47

¥2.81

Qwen3.7-Max-2026-05-20

qwen3.7-max-2026-05-20

256K

¥44.97

¥13.49

¥539.6

¥161.87

Qwen3.7-Plus-2026-05-26

qwen3.7-plus-2026-05-26

256K

¥7.19

¥2.88

¥86.3

¥34.53

Qwen3.6-Plus-2026-04-02

qwen3.6-plus-2026-04-02

128K

¥9

¥5.4

¥107.9

¥64.75

Qwen3.5-Plus-2026-04-20

qwen3.5-plus-2026-04-20

128K

¥7.2

¥4.32

¥86.3

¥51.8

DeepSeek

Model Name

Model Code

Max Input Token

Postpaid Input

Per 10K TPM/Hour

Postpaid Output

Per 1K TPM/Hour

Prepaid Input

Per 10K TPM/Day

Prepaid Output

Per 1K TPM/Day

DeepSeek-v4-Flash

deepseek-v4-flash

256K

¥5.4

¥1.08

¥64.8

¥12.95

DeepSeek-v4-Flash-0731

deepseek-v4-flash-0731

64K

¥10.79

¥2.16

¥129.5

¥25.9

DeepSeek-v4-Pro

deepseek-v4-pro

256K

¥64.75

¥12.95

¥777

¥155.4

Qwen-VL

Model Name

Model Code

Max Input Token

Postpaid Input

Per 10K TPM/Hour

Postpaid Output

Per 1K TPM/Hour

Prepaid Input

Per 10K TPM/Day

Prepaid Output

Per 1K TPM/Day

Qwen3-VL-Plus-2025-09-23

qwen3-vl-plus-2025-09-23

128K

¥3.6

¥2.88

¥43.2

¥34.53

GLM

Model Name

Model Code

Max Input Token

Postpaid Input

Per 10K TPM/Hour

Postpaid Output

Per 1K TPM/Hour

Prepaid Input

Per 10K TPM/Day

Prepaid Output

Per 1K TPM/Day

GLM-5.2

glm-5.2

1M

¥37.8

¥11.87

¥453.3

¥142.45

Billing by usage duration (Model Unit)

Cost = Usage duration (hours) × Number of Model Units × Model Unit unit price

The "Model Unit unit price" takes the "Hourly unit price" column in the table below for postpaid scenarios; for monthly subscription prepaid billing, the formula becomes Number of monthly subscriptions × Number of Model Units × Monthly unit price.

  • For the first month of a prepaid purchase, if you cancel the subscription early within the first month, the daily unit price (≈ Monthly unit price / 30) will be billed at 1.2 times the rate (less than one day is billed as one day)

Note

The compute resources for the Model Unit postpaid method are first-come, first-served. If the purchase fails, a full refund will be issued.

Text generation

Qwen

Model Name

Model Code

Model Unit specification

Hourly unit price (CNY)

Minimum billing: minute

Monthly subscription unit price (CNY)

Minimum billing: day

Qwen3.6-35B-A3B

qwen3.6-35b-a3b

MU1 x 8

¥432

¥208,944

MU2 x 8

¥504

¥240,288

MU3 x 8

¥1,096

¥527,752

MU8 x 1

¥47

¥22,400

MU9 x 1

¥51

¥24,600

Qwen3.6-27B

qwen3.6-27b

MU9 x 1

¥51

¥24,600

Qwen3.6-Flash-2026-04-16

qwen3.6-flash-2026-04-16

MU1 x 2

¥108

¥52,236

MU3 x 8

¥1,096

¥527,752

Qwen3.6-Plus-2026-04-02

qwen3.6-plus-2026-04-02

MU1 x 8

MU1 x 16 (PD disaggregation mode)

¥432

PD disaggregation mode: ¥864

¥208,944

PD disaggregation mode: ¥417,888

Qwen3.5-397B-A17B

qwen3.5-397b-a17b

MU3 x 8

MU3 x 16 (PD disaggregation mode)

¥1,096

PD disaggregation mode: ¥2,192

¥527,752

PD disaggregation mode: ¥1,055,504

MU6 x 16

¥400

¥193,424

Qwen3.5-122B-A10B

qwen3.5-122b-a10b

MU1 x 4

¥216

¥104,472

MU6 x 16

¥400

¥193,424

Qwen3.5-35B-A3B

qwen3.5-35b-a3b

MU1 x 2

¥108

¥52,236

MU2 x 8

¥504

¥240,288

MU3 x 8

¥1,096

¥527,752

MU9 x 1

¥51

¥24,600

Qwen3.5-27B

qwen3.5-27b

MU2 x 8

¥504

¥240,288

MU3 x 8

¥1,096

¥527,752

MU8 x 1

¥47

¥22,400

MU9 x 1

¥51

¥24,600

Qwen3.5-9B

qwen3.5-9b

MU1 x 2

¥108

¥52,236

MU2 x 8

¥504

¥240,288

Qwen3.5-Flash-2026-02-23

qwen3.5-flash-2026-02-23

MU1 x 2

¥108

¥52,236

Qwen3.5-Plus-2026-02-15

qwen3.5-plus-2026-02-15

MU1 x 8

MU1 x 16 (PD disaggregation mode)

¥432

PD disaggregation mode: ¥864

¥208,944

PD disaggregation mode: ¥417,888

MU2 x 8

¥504

¥240,288

MU3 x 8

MU3 x 16 (PD disaggregation mode)

¥1,096

PD disaggregation mode: ¥2,192

¥527,752

PD disaggregation mode: ¥1,055,504

Qwen3-235B-A22B-Instruct-2507

qwen3-235b-a22b-instruct-2507

MU1 x 4

¥216

¥104,472

MU2 x 8

¥504

¥240,288

MU3 x 8

¥1,096

¥527,752

Qwen3-32B

qwen3-32b

MU6 x 16

¥400

¥193,424

Qwen3-30B-A3B-Thinking-2507

qwen3-30b-a3b-thinking-2507

MU1 x 2

¥108

¥52,236

Qwen3-8B

qwen3-8b

MU1 x 2

¥108

¥52,236

MU2 x 2

¥126

¥60,072

Qwen3-4B

qwen3-4b

MU1 x 2

¥108

¥52,236

MU5 x 1

¥21

¥10,139

Qwen3-Embedding-0.6B

qwen3-embedding-0.6b

MU5 x 1

¥21

¥10,139

MU6 x 1

¥25

¥12,089

Qwen3-MoE-Rerank-0.6B

qwen3-moe-rerank-0.6b

MU5 x 1

¥21

¥10,139

Qwen3-Rerank-0.6B

qwen3-rerank-0.6b

MU5 x 1

¥21

¥10,139

MU6 x 1

¥25

¥12,089

Qwen3-Max-2025-09-23

qwen3-max-2025-09-23

MU2 x 8

¥504

¥240,288

MU3 x 8

¥1,096

¥527,752

Qwen3-Rerank

qwen3-rerank

MU5 x 1

¥21

¥10,139

Qwen2.5-Open-Source-72B

qwen2.5-72b-instruct

MU1 x 8

¥432

¥208,944

Qwen2.5-Open-Source-14B

qwen2.5-14b-instruct

MU1 x 2

¥108

¥52,236

Qwen2.5-Open-Source-7B

qwen2.5-7b-instruct

MU1 x 2

¥108

¥52,236

MU5 x 1

¥21

¥10,139

Qwen-Plus-2025-07-28

qwen-plus-2025-07-28

MU1 x 4

MU1 x 16 (PD disaggregation mode)

¥216

PD disaggregation mode: ¥864

¥104,472

PD disaggregation mode: ¥417,888

Qwen-Plus-2025-12-01

qwen-plus-2025-12-01

MU1 x 4

¥216

¥104,472

Qwen-Plus-Character-2025-11-06

qwen-plus-character-2025-11-06

MU1 x 4

¥216

¥104,472

GLM

Model name

Model code

Model unit specification

Hourly unit price (CNY)

Minimum billing: minute

Monthly subscription unit price (CNY)

Minimum billing: day

GLM-5.1

glm-5.1

MU2 x 8

¥504

¥240,288

MU3 x 16 (PD disaggregation mode)

PD disaggregation mode: ¥2,192

PD disaggregation mode: ¥1,055,504

MU6 x 16

¥400

¥193,424

GLM-5

glm-5

MU3 x 16 (PD disaggregation mode)

PD disaggregation mode: ¥2,192

PD disaggregation mode: ¥1,055,504

GLM-4.7

glm-4.7

MU6 x 32 (PD disaggregation mode)

PD disaggregation mode: ¥800

PD disaggregation mode: ¥386,848

DeepSeek

Model name

Model code

Model unit specification

Hourly unit price (CNY)

Minimum billing: minute

Monthly subscription unit price (CNY)

Minimum billing: day

DeepSeek-v4-Flash

deepseek-v4-flash

MU1 x 8

¥432

¥208,944

MU3 x 8

¥1,096

¥527,752

DeepSeek-v3.2

deepseek-v3.2

MU2 x 16 (PD disaggregation mode)

PD disaggregation mode: ¥1,008

PD disaggregation mode: ¥480,576

More models

Model name

Model code

Model unit specification

Hourly unit price (CNY)

Minimum billing: minute

Monthly subscription unit price (CNY)

Minimum billing: day

Kimi-K2.5

kimi-k2.5

MU2 x 8

¥504

¥240,288

Model type:

  • Instruct - The model performs inference in non-thinking mode after deployment.

  • Thinking - The model performs inference in thinking mode after deployment.

Model deployment type:

  • PD disaggregation mode - Reduces first-token latency and increases throughput.

    For models deployed in this mode, during inference the first-token computation (Prefill) and the subsequent token computation (Decode) are split across different compute nodes for execution.

Multimodal

Qwen-VL

Model name

Model code

Model unit specification

Hourly unit price (CNY)

Minimum billing: minute

Monthly subscription unit price (CNY)

Minimum billing: day

Qwen3-VL-235B-A22B-Thinking

qwen3-vl-235b-a22b-thinking

MU1 x 8

¥432

¥208,944

MU2 x 8

¥504

¥240,288

MU3 x 8

¥1,096

¥527,752

Qwen3-VL-32B-Instruct

qwen3-vl-32b-instruct

MU2 x 8

¥504

¥240,288

MU3 x 8

¥1,096

¥527,752

Qwen3-VL-8B-Instruct

qwen3-vl-8b-instruct

MU1 x 2

¥108

¥52,236

MU5 x 1

¥21

¥10,139

Qwen3-VL-4B-Instruct

qwen3-vl-4b-instruct

MU1 x 2

¥108

¥52,236

Qwen3-VL-2B-Instruct

qwen3-vl-2b-instruct

MU5 x 1

¥21

¥10,139

Qwen3-VL-Embedding-2B

qwen3-vl-embedding-2b

MU5 x 1

¥21

¥10,139

Qwen3-VL-Flash-2025-10-15

qwen3-vl-flash-2025-10-15

MU1 x 4

¥216

¥104,472

Qwen3-VL-Plus-2025-09-23

qwen3-vl-plus-2025-09-23

MU1 x 4

¥216

¥104,472

QwenVL-Max-2025-08-13

qwen-vl-max-2025-08-13

MU6 x 4

¥100

¥48,356

Qwen Omni

Model Name

Model Code

Model Unit Specification

Hourly Unit Price (CNY)

Minimum Billing: Minute

Monthly Subscription Unit Price (CNY)

Minimum Billing: Day

Qwen3.5-Omni-Flash

qwen3.5-omni-flash

MU8 x 1

¥47

¥22,400

MU9 x 1

¥51

¥24,600

Model Type:

  • Instruct - The model performs inference in non-thinking mode after deployment.

  • Thinking - The model performs inference in thinking mode after deployment.

  • Instruct/Thinking - You can choose whether to enable thinking mode when deploying the model.

Speech Synthesis

CosyVoice

Model Name

Model Code

Model Unit Specification

Hourly Unit Price (CNY)

Monthly Subscription Unit Price (CNY)

cosyvoice-v3-flash

cosyvoice-v3-flash

MU5

¥21

¥10,139

By model token usage

Fee = Model input token count × Model input unit price + Model output token count × Model output unit price (minimum billing unit: 1 token)

  • Billing by model token usage is supported only after you complete efficient SFT training (that is, LoRA efficient fine-tuning; the plan parameter is set to lora for API deployment) on the following base models and obtain a custom model.

Beijing

Base model

Model code

Input

CNY/Million tokens

Output

CNY/Million tokens

Qwen3.5-27B In preview

qwen3.5-27b

<128K ¥0.6

128K-256K ¥1.8

<128K ¥4.8

128K-256K ¥14.4

Qwen3-32B

qwen3-32b

Non-thinking mode: ¥2

Thinking mode: ¥2

Non-thinking mode: ¥8

Thinking mode: ¥20

Qwen3-14B

qwen3-14b

Non-thinking mode: ¥1

Thinking mode: ¥1

Non-thinking mode: ¥4

Thinking mode: ¥10

Qwen3-8B

qwen3-8b

Non-thinking mode: ¥0.5

Thinking mode: ¥0.5

Non-thinking mode: ¥2

Thinking mode: ¥5

Qwen3-VL-8B-Instruct

qwen3-vl-8b-instruct

¥0.5

¥2

Singapore

Base model

Model code

Input

CNY/Million tokens

Output

CNY/Million tokens

Qwen3-14B

qwen3-14b

Non-thinking mode: ¥2.569

Thinking mode: ¥2.569

Non-thinking mode: ¥10.275

Thinking mode: ¥30.825

To deploy more models, refer to thissolution and select the most suitable deployment plan based on your business requirements.

Deployment methods

You can deploy models on the console. Refer to the following steps:

If you are prompted with insufficient permissions, refer to:What should I do if "insufficient permissions" is prompted during deployment?
  1. Go to themodel deployment console.

image

image

  1. Enter the service name, select a model and a billing method, keep other settings as default, and click OK.

    You must completemodel fine-tuning before you can deploy most models.
  1. When the deployment status isRunning, the model has been deployed successfully.

Important

Fees will be incurred after the model is successfully deployed.

Deployment configuration

Model Unit

Configuration item

Configuration details

Service name

A custom name for the deployment service.

Model

Select the model to deploy, including platform preset models and fine-tuned models.

Model unit type

Select the deployment specification. Different specifications correspond to different computing power and performance.

Replica count

Set the initial number of deployment replicas, which affects the concurrent processing capability of the service.

Deployment template

Select a deployment template (for example, "single-node deployment"). Different templates correspond to different resource configuration schemes. Available only in the model unit billing mode.

Model inference mode

For some models, when deployed inModel Unit mode, you can configure the inference mode, maximum context, and more.

  • Instruct - The model performs inference in non-thinking mode after deployment.

  • Thinking - The model performs inference in thinking mode after deployment.

Maximum context

TheModel Unit deployment mode of some models supports this setting. The maximum context length depends on the model type.

Service throttling

TheModel Unit deployment mode of some models supports this setting, which can limit the RPM and TPM of model calls.

Deployment list page

After successful deployment, you can view and manage all deployment services on the deployment list page. The list page contains the following information:

  • Service name: The name of the deployment service. Click to view deployment details.

  • Model name: The model used for deployment.

  • Model Code: The unique identifier generated after the model is successfully deployed, used to specify the model when calling the API.

  • Deployment status/Event status: Includes Pending deployment, Deploying, Running, Deployment failed, Going offline, Service paused, Stopped, Deleting, Subscription suspended/Overdue payment suspended, Resuming service, Running (Changing configuration), Running (Change failed), and other statuses.

  • Billing method: The billing method of the current deployment service.

  • Deployment details: Configuration information such as model unit type and replica count.

  • Throttling details: Displays the throttling configuration of the current deployment service, such as RPM (requests per minute) and TPM (tokens per minute).

  • Service time: Displays the creation time and expiration time of the deployment service.

  • Operation: Depending on the deployment status and billing method, you can perform operations such as Update, Monitor, Scale, Renew, Take offline, Delete, and Try.

Post-deployment calls

After the model is successfully deployed, you can call it through OpenAI-compatible, Dashscope, and Assistant SDK.

When calling a successfully deployed model, the value of model should be the model code generated after successful deployment. Go to themodel deployment console (Beijing) to obtain the Model Code.

image

The following sample code calls the fine-tuned qwen3-8b model as an example:

Note

Model features (whether non-streaming output, structured output, etc. are supported) are consistent with themodel before fine-tuning.

For deep thinking models that have been fine-tuned, whether to enable deep thinking during calls is recommended to be consistent with the fine-tuning data format:

  • If the fine-tuning data contains deep thinking, it is recommended to enable the enable_thinking parameter when calling.

  • If the fine-tuning data does not contain deep thinking, it is not recommended to enable the enable_thinking parameter when calling.

Important

For GLM-5.2 models deployed with the preset throughput deployment method, the thinking_budget parameter (which limits the thinking length) does not take effect when called.

DashScope

import os
import dashscope

messages = [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": "Who are you?"},
]
response = dashscope.Generation.call(
    # If you have not configured environment variables, replace the next line with your Bailian API Key: api_key="sk-xxx",
    api_key=os.getenv("DASHSCOPE_API_KEY"),
    model="qwen3-14b-xxx-xxx",  # Please replace with the code returned after the model is successfully deployed
    messages=messages,
    result_format="message",
    enable_thinking=False,
)
print(response)

OpenAI-compatible interface

import os
from openai import OpenAI


client = OpenAI(
    # If you have not configured environment variables, replace the next line with your Bailian API Key: api_key="sk-xxx",
    api_key=os.getenv('DASHSCOPE_API_KEY'),
    base_url="https://dashscope.aliyuncs.com/compatible-mode/v1",
)

completion = client.chat.completions.create(
    model="qwen3-14b-xxx-xxx",  # Please replace with the code returned after the model is successfully deployed
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Who are you?"},
    ],
    extra_body={"enable_thinking": False},
)
print(completion)

Scaling a Deployment Service

  • Preset throughput (billed by duration): Click the Scaling button to self-service, manually adjust the number of instances. For detailed fee reduction and refund rules, please refer to: Refund rules for configuration downgrades.

  • Model unit (billed by duration): Click the Scaling button to self-service, manually adjust the number of instances.

  • Billed by Token usage: Click the Scale Out button, fill in and submit the scale-out application form, and wait for manual review.

In addition, you can configure an auto-scaling policy (including scaling thresholds, minimum/maximum replica count, scheduled scaling, etc.) via the scaling configuration button in the operation column.

Take a Deployment Service Offline

Go to the Model Deployment console, find the deployment service you want to stop, and click the corresponding operation according to the billing type:

  • Model unit prepaid: Click Deactivate and confirm.

  • Postpaid: Click Delete and confirm.

No further billing will be incurred after the operation is completed.

image

Other Operations

In addition to going offline, the operation column on the deployment list page also supports the following operations:

  • Update: Update the model version of the deployed service, supporting full update or batch update (canary release).

  • Delete: Pay-as-you-go services can be deleted directly to stop billing.

  • Renew: Prepaid services can be renewed to extend the service time, and auto-renewal is supported.

  • Buy capacity package: Purchase a capacity package for the preset throughput deployment.

FAQ

Can I upload and deploy my own models?

You can import some open-source models in the My Models console (Beijing); for the detailed supported list, please refer to: Model import.

In addition, Alibaba Cloud Artificial Intelligence Platform PAI provides the ability to deploy your own models. You can refer to PAI-LLM Large Language Model Deployment to learn about the deployment method.

What should I do if "insufficient permissions" is prompted during deployment?

  1. If "Missing permission for this module" is displayed, please ensure that your account has the Model Deployment - Operation permission on the permission management page of the business space.

    PixPin_2025-11-27_15-09-44

    If you cannot operate normally, please contact your organization or IT administrator to add the relevant permissions or check the permission issues on your behalf.

  2. If the error "xx business space does not have permission to deploy the xx model" is reported during deployment, please go to the Business Space Management page of Model Studio to add the deployment permission of the corresponding model for the corresponding business space.

    API call error: Workspace xxx does not have deployment privilege for model xxxx.

    PixPin_2025-11-27_15-03-57

    PixPin_2025-11-27_15-06-41

    If insufficient permissions are prompted, please contact your organization or IT administrator to add the relevant permissions or operate on your behalf.

How do I switch to other billing methods?

You can only release the original resources and then create new resources using the desired billing method.

It is recommended to switch according to the following steps:

  1. Deploy new resources using the desired billing method.

  2. Switch the API and test the service availability.

  3. Take offline and release the original resources.