New Model Deployment API Reference

Updated at:

This document uses the Qwen model deployment as an example to demonstrate how to use the model deployment feature of Alibaba Cloud Model Studio through HTTP API calls.

ImportantThis topic is applicable only to the China (Beijing) region.

Prerequisites

Get a List of Deployable Models

GET https://dashscope.aliyuncs.com/api/v1/deployments/models

Request Example

Run the following command to list deployable models. Use version=v1.0 to retrieve a full response that includes deployment plans and templates.

curl "https://dashscope.aliyuncs.com/api/v1/deployments/models?page_no=1&page_size=100&version=v1.0&model_source=base" \
    --header "Authorization: Bearer ${DASHSCOPE_API_KEY}" \
    --header 'Content-Type: application/json'

Query fine-tuned models:

curl "https://dashscope.aliyuncs.com/api/v1/deployments/models?page_no=1&page_size=100&version=v1.0&model_source=custom" \
    --header "Authorization: Bearer ${DASHSCOPE_API_KEY}" \
    --header 'Content-Type: application/json'

Request Parameters

Parameter

Type

Passing Parameters

Required

Description

page_no

Number

query

No

Page number. Default is 1.

page_size

Number

query

No

Page size. Default is 50. Maximum is 100. Minimum is 1.

model_source

String

query

No

Model source. base means system models (default). custom means fine-tuned models.

version

String

query

No

API version. Use v1.0. When you use v1.0, the response includes full deployment plan and template information.

Response Example and Parameters

Click to view the details.

After you run the command, the following result is returned:

{
    "request_id": "f7da015c-ea90-4d96-af89-2f8d7604026a",
    "output": {
        "page_no": 1,
        "page_size": 100,
        "total": 5,
        "models": [
            {
                "model_name": "qwen3-8b",
                "plans": [
                    {
                        "plan": "mu",
                        "templates": [
                            {
                                "template_id": "MU1",
                                "template_name": "Single-Node Deployment – Standard Inference",
                                "template_type": "COUPLED",
                                "template_version": "v1",
                                "template_desc": "For standard inference scenarios",
                                "roles": {
                                    "unified": {
                                        "model_unit_spec": "MU1",
                                        "capacity_unit_per_instance": 4
                                    }
                                }
                            },
                            {
                                "template_id": "MU1-PD",
                                "template_name": "PD-Separated Deployment – Standard Inference",
                                "template_type": "SEPERATED",
                                "template_version": "v1",
                                "template_desc": "For PD-separated inference scenarios",
                                "roles": {
                                    "prefill": {
                                        "model_unit_spec": "MU1",
                                        "capacity_unit_per_instance": 4
                                    },
                                    "decode": {
                                        "model_unit_spec": "MU1",
                                        "capacity_unit_per_instance": 4
                                    }
                                }
                            }
                        ]
                    },
                    {
                        "plan": "lora"
                    }
                ]
            }
        ]
    }
}

Response Parameters

Parameter

Type

Description

models

Array

List of deployable models.

models[].model_name

String

Model name.

models[].plans

Array

List of supported deployment plans for this model. Returned when using version=v1.0.

models[].plans[].plan

String

Deployment plan type: mu (model unit), cu (compute unit), ptu_v2 (pre-provisioned throughput v2), or lora (LoRA shared deployment).

models[].plans[].templates

Array

List of deployment templates. Returned when plan=mu.

page_no

Number

Page number of the query.

page_size

Number

Page size of the query.

total

Long

Total number of models that match the query criteria.

Template field description (templates)

Parameter

Type

Description

template_id

String

Template ID. Pass it as the template_id parameter in Create a Model Deployment Task.

template_name

String

Template display name.

template_type

String

Template type: COUPLED (non-PD-separated, uses the capacity parameter) or SEPERATED (PD-separated, uses prefill_capacity and decode_capacity parameters).

template_version

String

Template version.

template_desc

String

Template description.

roles

Object

Node role configuration. COUPLED mode has a unified node. SEPERATED mode has prefill and decode nodes.

Roles Node Fields

Parameter

Type

Description

model_unit_spec

String

Model unit specification.

capacity_unit_per_instance

Number

Capacity units per instance, also called base_capacity. When you create a deployment, capacity must be a multiple of this value.

Create a Model Deployment Task

POST https://dashscope.aliyuncs.com/api/v1/deployments

Request Example

Example 1: Model Unit Deployment (COUPLED Mode)

Use template_id to specify MU1 and set capacity to the number of capacity units:

curl "https://dashscope.aliyuncs.com/api/v1/deployments" \
--header "Authorization: Bearer $DASHSCOPE_API_KEY" \
--header 'Content-Type: application/json' \
--data '{
    "model_name": "qwen3-8b",
    "plan": "mu",
    "template_id": "MU1",
    "capacity": 4
}'

Example 2: Model Unit Deployment (PD-Separated Mode)

Use the MU1-PD template. Set prefill_capacity and decode_capacity:

curl "https://dashscope.aliyuncs.com/api/v1/deployments" \
--header "Authorization: Bearer $DASHSCOPE_API_KEY" \
--header 'Content-Type: application/json' \
--data '{
    "model_name": "qwen3-8b",
    "plan": "mu",
    "template_id": "MU1-PD",
    "prefill_capacity": 4,
    "decode_capacity": 8
}'

Example 3: PTU v2 Deployment

Billed by token usage:

curl "https://dashscope.aliyuncs.com/api/v1/deployments" \
--header "Authorization: Bearer $DASHSCOPE_API_KEY" \
--header 'Content-Type: application/json' \
--data '{
    "model_name": "qwen3-8b",
    "plan": "ptu_v2"
}'

Example 4: LoRA Shared Deployment

Billed by token usage:

curl "https://dashscope.aliyuncs.com/api/v1/deployments" \
--header "Authorization: Bearer $DASHSCOPE_API_KEY" \
--header 'Content-Type: application/json' \
--data '{
    "model_name": "qwen3-8b-ft-202511132025-0260",
    "plan": "lora"
}'

Request Parameters

Parameter

Type

Parameter passing

Required

Description

model_name

String

body

Yes

Name of the model to deploy.

plan

String

body

Yes

Deployment plan: mu (model unit), cu (compute unit), ptu_v2 (pre-provisioned throughput v2), or lora (LoRA shared deployment).

template_id

String

body

This parameter is required.

Template ID. Must match the template_id returned by the Get Deployable Models API. Required when plan is mu. We recommend using this parameter to specify the model unit spec.

deploy_spec (deprecated)

String

body

No

Model unit spec. Deprecated.

capacity

Number

body

This parameter is required.

Target number of capacity units. Required for COUPLED mode (non-PD-separated). Must be a multiple of capacity_unit_per_instance. Ignored but required when billing by token usage (plan is lora).

prefill_capacity

Number

body

This parameter is required.

Prefill capacity. Required for PD-separated mode (template_type is SEPERATED).

decode_capacity

Number

body

This parameter is required.

Decode capacity. Required for PD-separated mode (template_type is SEPERATED).

name

String

body

No

Service display name.

suffix

String

body

No

When you deploy a model, a new model name is generated. The suffix sets the suffix for the new model name. Max length is 8 characters. Must be globally unique. You can omit the suffix for the first deployment of a model. You must set it for later deployments of the same model.

charge_type

String

body

No

Billing method: post_paid (pay-as-you-go, default) or pre_paid (subscription).

pre_paid_info

Object

body

This parameter is required.

Subscription info. Required when charge_type is pre_paid. Includes: auto_renewal (Boolean, enable auto-renewal), duration (Number, subscription duration in months), and auto_renewal_duration (Number, auto-renewal duration in months).

pre_paid_info.auto_renewal

Boolean

body

Yes

Enable auto-renewal.

pre_paid_info.duration

Integer

body

Yes

Subscription duration in months.

pre_paid_info.auto_renewal_duration

Integer

body

Conditionally required

Auto-renewal duration in months. Required when autoRenewal=true.

enable_thinking

Boolean

body

No

Enable thinking mode. Supported by some models.

max_context_length

Number

body

No

Maximum context length. Supported by some models.

quantization_precision

String

body

No

Quantization precision.

qpm_limit

Number

body

No

QPM rate limit. Requests per minute. Supported only by some models with plan = mu.

tpm_limit

Number

body

No

TPM rate limit. Tokens per minute. Supported only by some models with plan = mu.

Supported Models

Model Unit deployment supports features such as rate limiting and inference acceleration. This table shows which features are supported.

Text generation

Qwen

Model name

Model code

Model unit specification

Hourly unit price (CNY)

Minimum billing: minute

Monthly unit price (CNY)

Minimum billing: day

Qwen3.6-35B-A3B

qwen3.6-35b-a3b

MU1 x 8

¥432

¥208,944

MU2 x 8

¥504

¥240,288

MU3 x 8

¥1,096

¥527,752

MU8 x 1

¥47

¥22,400

MU9 x 1

¥51

¥24,600

Qwen3.6-27B

qwen3.6-27b

MU9 x 1

¥51

¥24,600

Qwen3.6-Flash-2026-04-16

qwen3.6-flash-2026-04-16

MU1 x 2

¥108

¥52,236

MU3 x 8

¥1,096

¥527,752

Qwen3.6-Plus-2026-04-02

qwen3.6-plus-2026-04-02

MU1 x 8

MU1 x 16 (PD separation mode)

¥432

PD separation mode: ¥864

¥208,944

PD separation mode: ¥417,888

Qwen3.5-397B-A17B

qwen3.5-397b-a17b

MU3 x 8

MU3 x 16 (PD separation mode)

¥1,096

PD separation mode: ¥2,192

¥527,752

PD separation mode: ¥1,055,504

MU6 x 16

¥400

¥193,424

Qwen3.5-122B-A10B

qwen3.5-122b-a10b

MU1 x 4

¥216

¥104,472

MU6 x 16

¥400

¥193,424

Qwen3.5-35B-A3B

qwen3.5-35b-a3b

MU1 x 2

¥108

¥52,236

MU2 x 8

¥504

¥240,288

MU3 x 8

¥1,096

¥527,752

MU9 x 1

¥51

¥24,600

Qwen3.5-27B

qwen3.5-27b

MU2 x 8

¥504

¥240,288

MU3 x 8

¥1,096

¥527,752

MU8 x 1

¥47

¥22,400

MU9 x 1

¥51

¥24,600

Qwen3.5-9B

qwen3.5-9b

MU1 x 2

¥108

¥52,236

MU2 x 8

¥504

¥240,288

Qwen3.5-Flash-2026-02-23

qwen3.5-flash-2026-02-23

MU1 x 2

¥108

¥52,236

Qwen3.5-Plus-2026-02-15

qwen3.5-plus-2026-02-15

MU1 x 8

MU1 x 16 (PD separation mode)

¥432

PD separation mode: ¥864

¥208,944

PD separation mode: ¥417,888

MU2 x 8

¥504

¥240,288

MU3 x 8

MU3 x 16 (PD separation mode)

¥1,096

PD separation mode: ¥2,192

¥527,752

PD separation mode: ¥1,055,504

Qwen3-235B-A22B-Instruct-2507

qwen3-235b-a22b-instruct-2507

MU1 x 4

¥216

¥104,472

MU2 x 8

¥504

¥240,288

MU3 x 8

¥1,096

¥527,752

Qwen3-32B

qwen3-32b

MU6 x 16

¥400

¥193,424

Qwen3-30B-A3B-Thinking-2507

qwen3-30b-a3b-thinking-2507

MU1 x 2

¥108

¥52,236

Qwen3-8B

qwen3-8b

MU1 x 2

¥108

¥52,236

MU2 x 2

¥126

¥60,072

Qwen3-4B

qwen3-4b

MU1 x 2

¥108

¥52,236

MU5 x 1

¥21

¥10,139

Qwen3-Embedding-0.6B

qwen3-embedding-0.6b

MU5 x 1

¥21

¥10,139

MU6 x 1

¥25

¥12,089

Qwen3-MoE-Rerank-0.6B

qwen3-moe-rerank-0.6b

MU5 x 1

¥21

¥10,139

Qwen3-Rerank-0.6B

qwen3-rerank-0.6b

MU5 x 1

¥21

¥10,139

MU6 x 1

¥25

¥12,089

Qwen3-Max-2025-09-23

qwen3-max-2025-09-23

MU2 x 8

¥504

¥240,288

MU3 x 8

¥1,096

¥527,752

Qwen3-Rerank

qwen3-rerank

MU5 x 1

¥21

¥10,139

Qwen2.5-72B (Open Source)

qwen2.5-72b-instruct

MU1 x 8

¥432

¥208,944

Qwen2.5-14B (Open Source)

qwen2.5-14b-instruct

MU1 x 2

¥108

¥52,236

Qwen2.5-7B (Open Source)

qwen2.5-7b-instruct

MU1 x 2

¥108

¥52,236

MU5 x 1

¥21

¥10,139

Qwen-Plus-2025-07-28

qwen-plus-2025-07-28

MU1 x 4

MU1 x 16 (PD separation mode)

¥216

PD separation mode: ¥864

¥104,472

PD separation mode: ¥417,888

Qwen-Plus-2025-12-01

qwen-plus-2025-12-01

MU1 x 4

¥216

¥104,472

Qwen-Plus-Character-2025-11-06

qwen-plus-character-2025-11-06

MU1 x 4

¥216

¥104,472

GLM

Model name

Model code

Model unit specification

Hourly unit price (CNY)

Minimum billing: minute

Monthly unit price (CNY)

Minimum billing: day

GLM-5.1

glm-5.1

MU2 x 8

¥504

¥240,288

MU3 x 16 (PD separation mode)

PD separation mode: ¥2,192

PD separation mode: ¥1,055,504

MU6 x 16

¥400

¥193,424

GLM-5

glm-5

MU3 x 16 (PD separation mode)

PD separation mode: ¥2,192

PD separation mode: ¥1,055,504

GLM-4.7

glm-4.7

MU6 x 32 (PD separation mode)

PD separation mode: ¥800

PD separation mode: ¥386,848

DeepSeek

Model name

Model code

Model unit specification

Hourly unit price (CNY)

Minimum billing: minute

Monthly unit price (CNY)

Minimum billing: day

DeepSeek-v4-Flash

deepseek-v4-flash

MU1 x 8

¥432

¥208,944

MU3 x 8

¥1,096

¥527,752

DeepSeek-v3.2

deepseek-v3.2

MU2 x 16 (PD separation mode)

PD separation mode: ¥1,008

PD separation mode: ¥480,576

More models

Model name

Model code

Model unit specification

Hourly unit price (CNY)

Minimum billing: minute

Monthly unit price (CNY)

Minimum billing: day

Kimi-K2.5

kimi-k2.5

MU2 x 8

¥504

¥240,288

Model type:

  • Instruct - After deployment, the model performs inference in non-thinking mode.
  • Thinking - After deployment, the model performs inference in thinking mode.

Model deployment type:

  • PD separation mode - Reduces first-token latency and increases throughput. For models deployed in this mode, during inference the first-token computation (Prefill) and subsequent-token computation (Decode) — two computation phases — are split onto different compute nodes.

Multimodal

Qwen-VL

Model name

Model code

Model unit specification

Hourly unit price (CNY)

Minimum billing: minute

Monthly unit price (CNY)

Minimum billing: day

Qwen3-VL-235B-A22B-Thinking

qwen3-vl-235b-a22b-thinking

MU1 x 8

¥432

¥208,944

MU2 x 8

¥504

¥240,288

MU3 x 8

¥1,096

¥527,752

Qwen3-VL-32B-Instruct

qwen3-vl-32b-instruct

MU2 x 8

¥504

¥240,288

MU3 x 8

¥1,096

¥527,752

Qwen3-VL-8B-Instruct

qwen3-vl-8b-instruct

MU1 x 2

¥108

¥52,236

MU5 x 1

¥21

¥10,139

Qwen3-VL-4B-Instruct

qwen3-vl-4b-instruct

MU1 x 2

¥108

¥52,236

Qwen3-VL-2B-Instruct

qwen3-vl-2b-instruct

MU5 x 1

¥21

¥10,139

Qwen3-VL-Embedding-2B

qwen3-vl-embedding-2b

MU5 x 1

¥21

¥10,139

Qwen3-VL-Flash-2025-10-15

qwen3-vl-flash-2025-10-15

MU1 x 4

¥216

¥104,472

Qwen3-VL-Plus-2025-09-23

qwen3-vl-plus-2025-09-23

MU1 x 4

¥216

¥104,472

Qwen-VL-Max-2025-08-13

qwen-vl-max-2025-08-13

MU6 x 4

¥100

¥48,356

Qwen Omni

Model name

Model code

Model unit specification

Hourly unit price (CNY)

Minimum billing: minute

Monthly unit price (CNY)

Minimum billing: day

Qwen3.5-Omni-Flash

qwen3.5-omni-flash

MU8 x 1

¥47

¥22,400

MU9 x 1

¥51

¥24,600

Model type:

  • Instruct - After model deployment, inference is performed in non-thinking mode.
  • Thinking - After model deployment, inference is performed in thinking mode.
  • Instruct/Thinking - You can choose whether to enable thinking mode during model deployment.

Speech synthesis

CosyVoice

Model name

Model code

Model unit specification

Hourly unit price (CNY)

Monthly unit price (CNY)

cosyvoice-v3-flash

cosyvoice-v3-flash

MU5

¥21

¥10,139

Response Example and Parameters

Click to view the details.

After you run the command, the following result is returned:

{
    "request_id": "f2ae64f7-83cc-410c-bc0b-840443f7eb86",
    "output": {
        "deployed_model": "qwen3-8b-20260115-abcd",
        "name": "My Model Service",
        "model_name": "qwen3-8b",
        "base_model": "qwen3-8b",
        "status": "PENDING",
        "capacity": 4,
        "ready_capacity": 0,
        "base_capacity": 4,
        "workspace_id": "llm-8v53e*******",
        "charge_type": "post_paid",
        "plan": "mu",
        "model_unit_spec": "MU1",
        "creator": "16542902******",
        "modifier": "16542902******",
        "qpm_limit": 100000000,
        "tpm_limit": 3600000,
        "mu_capacity": {
            "unified": {
                "base_capacity": 4,
                "model_unit_spec": "MU1",
                "ready_capacity": 0,
                "capacity": 4
            }
        },
        "gmt_create": "2026-01-15T10:00:00",
        "gmt_modified": "2026-01-15T10:00:00"
    }
}

Parameter

Type

Description

request_id

String

ID of this request.

output

Object

Detailed information about this deployment task.

deployed_model

String

Unique identifier for the new model. Pass it in SDK parameters when you call the model.

name

String

Service display name.

model_name

String

Model name used for this deployment task.

base_model

String

Base model ID for the model used in this deployment task.

status

String

Status of the deployment task.

  • PENDING — Creating the deployment task.

  • DEPLOYING — Deploying.

  • RUNNING — Running and ready to process requests.

  • SCALING — Scaling.

  • UPDATING — Updating the deployment task.

  • UPDATING_PAUSED — Update paused.

  • SUSPENDING — Suspending (due to overdue payment).

  • SUSPENDED — Suspended (due to overdue payment).

  • STOPPING — Stopping.

  • STOPPED — Stopped. No longer billed.

  • STARTING — Starting (from STOPPED or FAILED state).

  • FAILED — Failed to create or update the deployment task.

  • DELETING — Deleting the deployment task.

capacity

Number

Target number of capacity units for this deployment task.

ready_capacity

Number

Number of capacity units that are ready and can immediately process requests.

base_capacity

Number

Base capacity per instance.

workspace_id

String

ID of the workspace where this deployment task belongs.

charge_type

String

post_paid: pay-as-you-go. pre_paid: subscription.

plan

String

Billing mode for this deployment task.

model_unit_spec

String

Model unit specification.

gmt_create

String

Time when the deployment task was created.

gmt_modified

String

You can modify the deployment task time.

creator

String

UID of the user who created this deployment task.

modifier

String

UID of the account that last performed an operation on this deployment task.

fail_reason

String

Failure reason. Returned only when status is FAILED.

enable_thinking

Boolean

Enable thinking mode. Supported by some models.

max_context_length

Number

Maximum context length limit.

qpm_limit

Number

QPM rate limit. Requests per minute.

tpm_limit

Number

TPM rate limit. Tokens per minute.

quantization_precision

String

Quantization precision.

Returned only for model unit (mu) deployments.

mu_capacity

Object

MU capacity details. COUPLED mode has a unified node. SEPERATED mode has prefill and decode nodes. Each node includes: base_capacity, model_unit_spec, ready_capacity, and capacity.

pre_paid_info

Object

Subscription information.

pre_paid_info.auto_renewal

Boolean

Enable auto-renewal.

pre_paid_info.duration

Integer

Subscription duration in months.

pre_paid_info.auto_renewal_duration

Integer

Auto-renewal duration in months.

pre_paid_instance_id

String

Subscription instance ID.

pre_paid_gmt_expired

String

Subscription expiration time.

Returned only for pre-provisioned throughput v2 (ptu_v2) deployments.

ptu_capacity

Object

PTU v2 capacity info. Includes:

  • input_tpm_quota (Number, input TPM quota)

  • output_tpm_quota (Number, output TPM quota)

pre_paid_info

Object

Subscription information.

pre_paid_info.auto_renewal

Boolean

Enable auto-renewal.

pre_paid_info.duration

Integer

Subscription duration in months.

pre_paid_info.auto_renewal_duration

Integer

Auto-renewal duration in months.

pre_paid_instance_id

String

Subscription instance ID.

pre_paid_gmt_expired

String

Subscription expiration time.

Query a Model Deployment Task

GET https://dashscope.aliyuncs.com/api/v1/deployments/{deployed_model}

Request Example

curl "https://dashscope.aliyuncs.com/api/v1/deployments/qwen3-8b-20260115-abcd" \
    --header "Authorization: Bearer ${DASHSCOPE_API_KEY}" \
    --header 'Content-Type: application/json'

Request Parameters

Parameter

Type

Passing Parameters

Required

Description

deployed_model

String

path

Yes

Unique identifier for the deployed model. Get it from Create a Model Deployment Task or List Model Deployment Tasks.

Response Example and Parameters

Click to view the details.

After you run the command, the following result is returned:

{
    "request_id": "d8b8c47b-ae32-4c41-b9f4-9a1e2b3c4d5e",
    "output": {
        "deployed_model": "qwen3-8b-20260115-abcd",
        "name": "My Model Service",
        "model_name": "qwen3-8b",
        "base_model": "qwen3-8b",
        "status": "RUNNING",
        "capacity": 4,
        "ready_capacity": 4,
        "base_capacity": 4,
        "workspace_id": "llm-8v53e*******",
        "charge_type": "post_paid",
        "plan": "mu",
        "model_unit_spec": "MU1",
        "creator": "16542902******",
        "modifier": "16542902******",
        "qpm_limit": 100000000,
        "tpm_limit": 3600000,
        "mu_capacity": {
            "unified": {
                "base_capacity": 4,
                "model_unit_spec": "MU1",
                "ready_capacity": 4,
                "capacity": 4
            }
        },
        "gmt_create": "2026-01-15T10:00:00",
        "gmt_modified": "2026-01-15T10:30:00"
    }
}

Response Parameters

The response parameters are the same as those for creating a model deployment task. For more information, see Response Parameters for Create a Model Deployment Task.

List Model Deployment Tasks

GET https://dashscope.aliyuncs.com/api/v1/deployments

Request Example

curl "https://dashscope.aliyuncs.com/api/v1/deployments?page_no=1&page_size=10" \
    --header "Authorization: Bearer ${DASHSCOPE_API_KEY}" \
    --header 'Content-Type: application/json'

Request Parameters

Parameter

Type

Parameter passing

Required

Description

Age

Number

query

No

Page number. Default is 1.

page_size

Number

query

No

Page size. Default is 10. Maximum is 100.

Response Example and Parameters

Click to view the details.

After you run the command, the following result is returned:

{
    "request_id": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
    "output": {
        "page_no": 1,
        "page_size": 10,
        "total": 2,
        "deployments": [
            {
                "deployed_model": "qwen3-8b-20260115-abcd",
                "name": "My Model Service",
                "model_name": "qwen3-8b",
                "base_model": "qwen3-8b",
                "status": "RUNNING",
                "capacity": 4,
                "ready_capacity": 4,
                "base_capacity": 4,
                "workspace_id": "llm-8v53e*******",
                "charge_type": "post_paid",
                "plan": "mu",
                "model_unit_spec": "MU1",
                "creator": "16542902******",
                "modifier": "16542902******",
                "qpm_limit": 100000000,
                "tpm_limit": 3600000,
                "mu_capacity": {
                    "unified": {
                        "base_capacity": 4,
                        "model_unit_spec": "MU1",
                        "ready_capacity": 4,
                        "capacity": 4
                    }
                },
                "gmt_create": "2026-01-15T10:00:00",
                "gmt_modified": "2026-01-15T10:30:00"
            }
        ]
    }
}

Response Parameters

Parameter

Type

Description

page_no

Number

Page number of the query.

page_size

Number

Page size of the query.

total

Long

Total number of deployment tasks that match the query criteria.

deployments

Array

List of deployment tasks. The fields for each element match the response parameters for Create a Model Deployment Task. For details, see the response parameters for Create a Model Deployment Task.

Update a Model Deployment Task

NoteUpdating a model deployment task is equivalent to scaling the deployment service. You can call this API only when the service status is RUNNING or SCALING. If you call this API when the service is in another state, an OperationDenied error is returned.

PUT https://dashscope.aliyuncs.com/api/v1/deployments/{deployed_model}/scale

Request Example

Scale in COUPLED mode:

curl -X PUT "https://dashscope.aliyuncs.com/api/v1/deployments/qwen3-8b-20260115-abcd/scale" \
    --header "Authorization: Bearer ${DASHSCOPE_API_KEY}" \
    --header 'Content-Type: application/json' \
    --data '{
    "capacity": 8
}'

Scale in PD-separated mode:

curl -X PUT "https://dashscope.aliyuncs.com/api/v1/deployments/qwen3-8b-20260115-abcd/scale" \
    --header "Authorization: Bearer ${DASHSCOPE_API_KEY}" \
    --header 'Content-Type: application/json' \
    --data '{
    "prefill_capacity": 4,
    "decode_capacity": 8
}'

Request Parameters

Parameter

Type

Passing parameters

Required

Description

deployed_model

String

path

Yes

Unique identifier for the deployed model.

capacity

Number

body

Conditionally required

Target number of capacity units for scaling. Required for COUPLED mode (non-PD-separated). Must be a multiple of base_capacity.

prefill_capacity

Number

body

Conditionally required

Target prefill capacity. Required for PD-separated mode.

decode_capacity

Number

body

A condition is required.

Target decode capacity. Required for PD-separated mode.

ptu_capacity

Object

body

Conditionally required

PTU v2 capacity info. Includes:

  • input_tpm_quota (Number, input TPM quota)

  • output_tpm_quota (Number, output TPM quota)

Response Example and Parameters

Click to view the details.

After you run the command, the following result is returned:

{
    "request_id": "e5f6a7b8-c9d0-1234-5678-abcdef012345",
    "output": {
        "deployed_model": "qwen3-8b-20260115-abcd",
        "name": "My Model Service",
        "model_name": "qwen3-8b",
        "base_model": "qwen3-8b",
        "status": "SCALING",
        "capacity": 8,
        "ready_capacity": 4,
        "base_capacity": 4,
        "workspace_id": "llm-8v53e*******",
        "charge_type": "post_paid",
        "plan": "mu",
        "model_unit_spec": "MU1",
        "creator": "16542902******",
        "modifier": "16542902******",
        "qpm_limit": 100000000,
        "tpm_limit": 3600000,
        "mu_capacity": {
            "unified": {
                "base_capacity": 4,
                "model_unit_spec": "MU1",
                "ready_capacity": 4,
                "capacity": 8
            }
        },
        "gmt_create": "2026-01-15T10:00:00",
        "gmt_modified": "2026-01-15T11:00:00"
    }
}

Response Parameters

The response parameters are the same as those for creating a model deployment task. For more information, see the response parameters for Create a Model Deployment Task.

Stop a Deployed Model

NoteOnly model unit (mu) deployments support stopping the service. You can call this API only when the service status is RUNNING or SCALING. If you call this API for cu, ptu_v2, or lora deployments, an UnsupportedOperation error is returned. After the service is stopped, it is no longer billed. You can restart the service using the API Details API.

POST https://dashscope.aliyuncs.com/api/v1/deployments/{deployed_model}/stop

Request Example

curl -X POST "https://dashscope.aliyuncs.com/api/v1/deployments/qwen3-8b-20260115-abcd/stop" \
    --header "Authorization: Bearer ${DASHSCOPE_API_KEY}" \
    --header 'Content-Type: application/json'

Request Parameters

Parameter

Type

Passing parameters

Required

Description

deployed_model

String

path

Yes

Unique identifier for the deployed model.

Response Example and Parameters

Click to view the details.

After you run the command, the following result is returned:

{
    "request_id": "f1a2b3c4-d5e6-7890-abcd-ef1234567890",
    "output": {
        "deployed_model": "qwen3-8b-20260115-abcd",
        "name": "My Model Service",
        "model_name": "qwen3-8b",
        "base_model": "qwen3-8b",
        "status": "STOPPING",
        "capacity": 4,
        "ready_capacity": 4,
        "base_capacity": 4,
        "workspace_id": "llm-8v53e*******",
        "charge_type": "post_paid",
        "plan": "mu",
        "model_unit_spec": "MU1",
        "creator": "16542902******",
        "modifier": "16542902******",
        "qpm_limit": 100000000,
        "tpm_limit": 3600000,
        "mu_capacity": {
            "unified": {
                "base_capacity": 4,
                "model_unit_spec": "MU1",
                "ready_capacity": 4,
                "capacity": 4
            }
        },
        "gmt_create": "2026-01-15T10:00:00",
        "gmt_modified": "2026-01-15T12:00:00"
    }
}

Response Parameters

The response parameters are the same as those for creating a model deployment task. For more information, see the response parameters for Create a Model Deployment Task.

Restart a Deployed Model

POST https://dashscope.aliyuncs.com/api/v1/deployments/{deployed_model}/start

Request Example

curl -X POST "https://dashscope.aliyuncs.com/api/v1/deployments/qwen3-8b-20260115-abcd/start" \
    --header "Authorization: Bearer ${DASHSCOPE_API_KEY}" \
    --header 'Content-Type: application/json'

Request Parameters

Parameter

Type

Passing parameters

Required

Description

deployed_model

String

path

Yes

Unique identifier for the deployed model. Get it from Create a Model Deployment Task or List Model Deployment Tasks.

Response Example and Parameters

Click to view the details.

After you run the command, the following result is returned:

After the service starts, the status is STARTING and then changes to RUNNING

{
  "request_id": "xxx-xxx-xxx",
  "output": {
    "deployed_model": "qwen-7b-chat-ft-20240101-abcd",
    "model_name": "qwen-7b-chat-ft",
    "base_model": "qwen-7b-chat",
    "status": "STARTING",
    "capacity": 1,
    "ready_capacity": 0,
    "base_capacity": 1,
    "plan": "mu",
    "model_unit_spec": "MU5",
    "workspace_id": "workspace-xxx",
    "charge_type": "post_paid",
    "creator": "user123",
    "modifier": "user123",
    "qpm_limit": 100000000,
    "tpm_limit": 3600000,
    "mu_capacity": {
      "unified": {
        "base_capacity": 1,
        "model_unit_spec": "",
        "ready_capacity": 0,
        "capacity": 1
      }
    },
    "gmt_create": "2024-01-01T10:00:00",
    "gmt_modified": "2024-01-01T13:00:00"
  }
}

Response Parameters

The response parameters are the same as those for creating a model deployment task. For more information, see the response parameters for Create a Model Deployment Task.

Delete a Model Deployment Task

NoteYou can delete a service only when its status is STOPPED or FAILED. For services in the RUNNING state, you must first call the API Details API to stop them.

DELETE https://dashscope.aliyuncs.com/api/v1/deployments/{deployed_model}

Request Example

curl -X DELETE "https://dashscope.aliyuncs.com/api/v1/deployments/qwen3-8b-20260115-abcd" \
    --header "Authorization: Bearer ${DASHSCOPE_API_KEY}" \
    --header 'Content-Type: application/json'

Request Parameters

Parameter

Type

Parameter passing

Required

Description

deployed_model

String

path

Yes

Unique identifier for the deployed model.

Response Example and Parameters

Click to view the details.

After you run the command, the following result is returned:

{
    "request_id": "g1h2i3j4-k5l6-7890-mnop-qr1234567890",
    "output": {
        "deployed_model": "qwen3-8b-20260115-abcd",
        "name": "My Model Service",
        "model_name": "qwen3-8b",
        "base_model": "qwen3-8b",
        "status": "DELETING",
        "capacity": 4,
        "ready_capacity": 0,
        "base_capacity": 4,
        "workspace_id": "llm-8v53e*******",
        "charge_type": "post_paid",
        "plan": "mu",
        "model_unit_spec": "MU1",
        "creator": "16542902******",
        "modifier": "16542902******",
        "gmt_create": "2026-01-15T10:00:00",
        "gmt_modified": "2026-01-15T14:00:00"
    }
}

Response Parameters

The response parameters are the same as those for creating a model deployment task. For more information, see the response parameters for Create a Model Deployment Task.

Abnormal Response

{
    "request_id": "6fec0a3f-a1df-4f89-8c8f-33cb0db9f5b0",
    "code": "InvalidParameter",
    "message": "capacity is required when template_type is COUPLED"
}

Response Parameters

Parameter

Type

Description

request_id

String

System-generated ID for this request.

code

String

Error code that identifies the error type.

message

String

Error details for troubleshooting.

Common Errors

Error Reason

Description

InvalidParameter

Parameter error. For example, a required parameter is missing, a parameter value is invalid, or template_id and deploy_spec are passed at the same time.

AccessDenied

Insufficient permissions or resource does not exist.

OperationDenied

The operation is not allowed in the current state. For example, scaling a service that is not in the RUNNING state, or deleting a service that is not in the STOPPED or FAILED state.

UnsupportedOperation

The current deployment plan does not support this operation. For example, cu deployments do not support stopping the service.

StockLimit

Insufficient resource stock. Cannot complete scaling or deployment.

InternalError

Internal error.

NotFound

The specified deployment task does not exist.

Conflict

Operation conflict. For example, the service is performing another operation.