DTU Dedicated compute deployment

Updated at:

Dedicated compute deployment offers two options: Dedicated Throughput Unit (DTU) and Model Unit (MU), providing dedicated GPU resources for specified models with performance and throughput guarantees. DTU targets newly released models (billed by input/output TPM × duration); MU is used for existing models (billed by model unit count × duration). This document introduces both options' features, billing, usage flow, and supported models.

Overview

Dedicated compute deployment provides dedicated GPU resources and performance guarantees for specified models, supporting both base models and custom model deployment. Bailian offers two dedicated compute options:

  • DTU (Dedicated Throughput Unit): Billed by Input/Output TPM, providing more direct throughput guarantees. Fully managed inference service with dedicated underlying GPU resources maintained by the platform.
  • Model Unit (MU): Configures compute power based on usage duration and the number of Model Units, with dedicated resources. Supports custom performance metrics and PD disaggregated computing mode. Billing granularity is model unit count × duration.

DTU and MU are differentiated by model applicability: newly released models use DTU (metered by input/output TPM), while existing models continue to use MU (metered by model unit count). Both provide dedicated compute, resource isolation, and fine-tuned model deployment. For new deployments, choose DTU if the model supports it.

Shared capabilities of both options:

  • Dedicated GPU resources; the inference environment is physically isolated from other users.
  • Supports base models and custom models (Bailian fine-tuned models or user-uploaded models).
  • The platform maintains the underlying operations, no need to manage GPUs.
  • No proactive RPM/TPM limits; traffic is bound by actual capacity.

For the general model deployment workflow, see Dedicated Deployment Overview and Provisioned Throughput.

Solution Selection

Bailian offers multiple capacity and billing plans for inference calls. DTU suits scenarios requiring dedicated deployment, fine-tuned model deployment, low latency with high concurrency, or data isolation. The plans are compared in the table below.

For the full comparison of all billing methods including Token pay-as-you-go and PTU, see the main table in Dedicated Deployment Overview.

Comparison Dimension

DTU

PTU

Token Pay-as-you-go

PAI/Lingjun

Feature

Dedicated compute + fine-tuned models

Throughput guarantee + low latency

Elastic and zero-threshold

Custom runtime

Resource isolation

Physical isolation

Logical isolation

Shared pool

Physical isolation

Fine-tuned models

Full-parameter + LoRA

Not supported

LoRA

Supported

Custom framework

Not supported

Not supported

Not supported

Supported

Billing mode

TPM quota × duration

TPM quota × duration

Token usage

GPU × duration

Billing Rules

DTU Billing

DTU is billed separately by input TPM and output TPM, using a pre-paid (monthly) model. Each model has fixed input/output baseline TPM (see table below), and you must purchase in integer multiples of the baseline TPM. Purchase at least 1× baseline TPM for both input and output; for production, 2× or above is recommended for each.

Fees are calculated by the backend based on model, input/output TPM, service region, and purchase duration. Prices for each model are shown in the table below; the starting and scaling prices are both the total for 1× baseline TPM of input and output. Fine-tuned model prices are the same as the corresponding base model, subject to the actual console display.

MU Billing

Cost = Usage duration (hours) × Number of Model Units × Model Unit unit price

The "Model Unit unit price" takes the "Hourly unit price" column in the table below for postpaid scenarios; for monthly subscription prepaid billing, the formula becomes Number of monthly subscriptions × Number of Model Units × Monthly unit price.

  • For the first month of a prepaid purchase, if you cancel the subscription early within the first month, the daily unit price (≈ Monthly unit price / 30) will be billed at 1.2 times the rate (less than one day is billed as one day)

NoteThe compute resources for the Model Unit postpaid method are first-come, first-served. If the purchase fails, a full refund will be issued.

Both MU and DTU use pre-paid billing. Early termination settles the used portion at a 1.2x coefficient for refund. For details, see Refund rules for configuration downgrades.

Supported Models and Pricing

DTU Price Table and Performance Baseline

China (Beijing) 2

Model

Spec

Baseline Input (TPM)

Baseline Output (TPM)

Starting Price (CNY/month)

Scaling Price (CNY/month)

qwen3.7-plus-2026-05-26

Fast

1,372,000

170,000

556,466

556,466

qwen3.6-27b

Standard

273,000

34,000

49,118

49,118

qwen3.5-397b-a17b

Standard

896,000

112,000

1,107,904

1,107,904

qwen3.5-122b-a10b

Fast

7,288,000

904,000

2,226,984

2,226,984

qwen3.5-35b-a3b

Standard

656,000

82,000

109,962

109,962

qwen3.5-27b

Standard

2,432,000

304,000

1,108,688

1,108,688

Efficient

291,000

36,000

48,996

48,996

qwen3.5-4b

Standard

1,072,000

134,000

109,478

109,478

glm-5.2

Fast

1,004,000

120,000

1,895,260

1,895,260

glm-5.1

Standard

256,000

32,000

504,704

504,704

deepseek-v4-flash

Standard

2,240,000

280,000

1,107,400

1,107,400

deepseek-v4-flash-0731

Fast

1,676,000

213,000

277,984

277,984

The performance reference data for each model under standard workload is shown in the table below, subject to the actual console display.

The following performance reference data was measured at a 0% cache hit rate. In actual use, as the cache hit rate increases, model performance improves accordingly.

Model

Spec

Input Length

Output Length

Cache Hit Rate

First-token Latency (ms)

Per-token Latency (ms)

qwen3.7-plus-2026-05-26

Fast

16,000

2,000

0

2,418

15

qwen3.6-27b

Standard

4,000

500

0

1,292

19

qwen3.5-397b-a17b

Standard

4,000

500

0

996

27

qwen3.5-122b-a10b

Fast

16,000

2,000

0

568

8

qwen3.5-35b-a3b

Standard

4,000

500

0

471

10

qwen3.5-27b

Standard

4,000

500

0

703

14

Efficient

4,000

500

0

1,448

23

qwen3.5-4b

Standard

4,000

500

0

552

6

glm-5.2

Fast

16,000

2,000

0

1,558

15

glm-5.1

Standard

4,000

500

0

769

27

deepseek-v4-flash

Standard

4,000

500

0

651

19

deepseek-v4-flash-0731

Fast

16,000

2,000

0

1,079

13

MU Price Table

Text generation

Qwen

Model Name

Model Code

Model Unit Specification

Hourly Unit Price (CNY)

Minimum billing: minutes

Monthly Subscription Unit Price (CNY)

Minimum billing: days

Qwen3.8-27B

qwen3.8-27b

MU9 x 4

¥204

¥98,400

Qwen3.7-Plus-2026-05-26

qwen3.7-plus-2026-05-26

MU2 x 8

¥504

¥240,288

MU3 x 8

¥1,096

¥527,752

Qwen3.6-35B-A3B

qwen3.6-35b-a3b

MU1 x 8

¥432

¥208,944

MU2 x 8

¥504

¥240,288

MU3 x 8

¥1,096

¥527,752

MU8 x 1

¥47

¥22,400

MU9 x 1

¥51

¥24,600

Qwen3.6-27B

qwen3.6-27b

MU9 x 1

¥51

¥24,600

Qwen3.6-Flash-2026-04-16

qwen3.6-flash-2026-04-16

MU1 x 2

¥108

¥52,236

MU3 x 8

¥1,096

¥527,752

Qwen3.6-Plus-2026-04-02

qwen3.6-plus-2026-04-02

MU1 x 8

MU1 x 16(PD disaggregation mode)

¥432

PD disaggregation mode:¥864

¥208,944

PD disaggregation mode:¥417,888

MU2 x 8

¥504

¥240,288

Qwen3.5-397B-A17B

qwen3.5-397b-a17b

MU3 x 8

MU3 x 16(PD disaggregation mode)

¥1,096

PD disaggregation mode:¥2,192

¥527,752

PD disaggregation mode:¥1,055,504

MU6 x 16

¥400

¥193,424

Qwen3.5-122B-A10B

qwen3.5-122b-a10b

MU1 x 4

¥216

¥104,472

MU6 x 16

¥400

¥193,424

Qwen3.5-35B-A3B

qwen3.5-35b-a3b

MU1 x 2

¥108

¥52,236

MU2 x 8

¥504

¥240,288

MU3 x 8

¥1,096

¥527,752

MU9 x 1

¥51

¥24,600

Qwen3.5-27B

qwen3.5-27b

MU2 x 8

¥504

¥240,288

MU3 x 8

¥1,096

¥527,752

MU8 x 1

¥47

¥22,400

MU9 x 1

¥51

¥24,600

Qwen3.5-9B

qwen3.5-9b

MU1 x 2

¥108

¥52,236

MU2 x 8

¥504

¥240,288

Qwen3.5-Flash-2026-02-23

qwen3.5-flash-2026-02-23

MU1 x 2

¥108

¥52,236

Qwen3.5-Plus-2026-02-15

qwen3.5-plus-2026-02-15

MU1 x 8

MU1 x 16(PD disaggregation mode)

¥432

PD disaggregation mode:¥864

¥208,944

PD disaggregation mode:¥417,888

MU2 x 8

¥504

¥240,288

MU3 x 8

MU3 x 16(PD disaggregation mode)

¥1,096

PD disaggregation mode:¥2,192

¥527,752

PD disaggregation mode:¥1,055,504

Qwen3-235B-A22B-Instruct-2507

qwen3-235b-a22b-instruct-2507

MU1 x 4

¥216

¥104,472

MU2 x 8

¥504

¥240,288

MU3 x 8

¥1,096

¥527,752

Qwen3-32B

qwen3-32b

MU6 x 16

¥400

¥193,424

Qwen3-30B-A3B-Thinking-2507

qwen3-30b-a3b-thinking-2507

MU1 x 2

¥108

¥52,236

Qwen3-4B

qwen3-4b

MU1 x 2

¥108

¥52,236

MU5 x 1

¥21

¥10,139

Qwen3-Embedding-0.6B

qwen3-embedding-0.6b

MU5 x 1

¥21

¥10,139

MU6 x 1

¥25

¥12,089

Qwen3-MoE-Rerank-0.6B

qwen3-moe-rerank-0.6b

MU5 x 1

¥21

¥10,139

Qwen3-Rerank-0.6B

qwen3-rerank-0.6b

MU5 x 1

¥21

¥10,139

MU6 x 1

¥25

¥12,089

Qwen3-Max-2025-09-23

qwen3-max-2025-09-23

MU2 x 8

¥504

¥240,288

MU3 x 8

¥1,096

¥527,752

Qwen3-Rerank

qwen3-rerank

MU5 x 1

¥21

¥10,139

Qwen2.5-Open-Source-72B

qwen2.5-72b-instruct

MU1 x 8

¥432

¥208,944

Qwen2.5-Open-Source-14B

qwen2.5-14b-instruct

MU1 x 2

¥108

¥52,236

Qwen2.5-Open-Source-7B

qwen2.5-7b-instruct

MU1 x 2

¥108

¥52,236

MU5 x 1

¥21

¥10,139

Qwen-Plus-2025-07-28

qwen-plus-2025-07-28

MU1 x 4

MU1 x 16(PD disaggregation mode)

¥216

PD disaggregation mode:¥864

¥104,472

PD disaggregation mode:¥417,888

Qwen-Plus-2025-12-01

qwen-plus-2025-12-01

MU1 x 4

¥216

¥104,472

Qwen-Plus-Character-2025-11-06

qwen-plus-character-2025-11-06

MU1 x 4

¥216

¥104,472

GLM

Model Name

Model Code

Model Unit Specification

Hourly Unit Price (CNY)

Minimum billing: minutes

Monthly Subscription Unit Price (CNY)

Minimum billing: days

GLM-5.1

glm-5.1

MU2 x 8

¥504

¥240,288

MU3 x 16(PD disaggregation mode)

PD disaggregation mode:¥2,192

PD disaggregation mode:¥1,055,504

MU6 x 16

¥400

¥193,424

GLM-5

glm-5

MU3 x 16(PD disaggregation mode)

PD disaggregation mode:¥2,192

PD disaggregation mode:¥1,055,504

GLM-4.7

glm-4.7

MU6 x 32(PD disaggregation mode)

PD disaggregation mode:¥800

PD disaggregation mode:¥386,848

DeepSeek

Model Name

Model Code

Model Unit Specification

Hourly Unit Price (CNY)

Minimum billing: minutes

Monthly Subscription Unit Price (CNY)

Minimum billing: days

DeepSeek-v4-Flash

deepseek-v4-flash

MU1 x 8

¥432

¥208,944

MU3 x 8

¥1,096

¥527,752

DeepSeek-v3.2

deepseek-v3.2

MU2 x 16(PD disaggregation mode)

PD disaggregation mode:¥1,008

PD disaggregation mode:¥480,576

More models

Model Name

Model Code

Model Unit Specification

Hourly Unit Price (CNY)

Minimum billing: minutes

Monthly Subscription Unit Price (CNY)

Minimum billing: days

Kimi-K2.5

kimi-k2.5

MU2 x 8

¥504

¥240,288

Model type:

  • Instruct - The model performs inference in non-thinking mode after deployment.
  • Thinking - The model performs inference in thinking mode after deployment.

Model deployment type:

  • PD disaggregation mode - Reduces first-token latency and increases throughput.

    For models deployed in this mode, during inference the first-token computation (Prefill) and the subsequent token computation (Decode) are split across different compute nodes for execution.

Multimodal

Qwen-VL

Model Name

Model Code

Model Unit Specification

Hourly Unit Price (CNY)

Minimum billing: minutes

Monthly Subscription Unit Price (CNY)

Minimum billing: days

Qwen3-VL-235B-A22B-Thinking

qwen3-vl-235b-a22b-thinking

MU1 x 8

¥432

¥208,944

MU2 x 8

¥504

¥240,288

MU3 x 8

¥1,096

¥527,752

Qwen3-VL-32B-Instruct

qwen3-vl-32b-instruct

MU2 x 8

¥504

¥240,288

MU3 x 8

¥1,096

¥527,752

Qwen3-VL-8B-Instruct

qwen3-vl-8b-instruct

MU1 x 2

¥108

¥52,236

MU5 x 1

¥21

¥10,139

Qwen3-VL-4B-Instruct

qwen3-vl-4b-instruct

MU1 x 2

¥108

¥52,236

Qwen3-VL-2B-Instruct

qwen3-vl-2b-instruct

MU5 x 1

¥21

¥10,139

Qwen3-VL-Embedding-2B

qwen3-vl-embedding-2b

MU5 x 1

¥21

¥10,139

Qwen3-VL-Flash-2025-10-15

qwen3-vl-flash-2025-10-15

MU1 x 4

¥216

¥104,472

Qwen3-VL-Plus-2025-09-23

qwen3-vl-plus-2025-09-23

MU1 x 4

¥216

¥104,472

Qwen-VL-Max-2025-08-13

qwen-vl-max-2025-08-13

MU6 x 4

¥100

¥48,356

Qwen Omni

Model Name

Model Code

Model Unit Specification

Hourly Unit Price (CNY)

Minimum billing: minutes

Monthly Subscription Unit Price (CNY)

Minimum billing: days

Qwen3.5-Omni-Flash

qwen3.5-omni-flash

MU8 x 1

¥47

¥22,400

MU9 x 1

¥51

¥24,600

Model Type:

  • Instruct - The model performs inference in non-thinking mode after deployment.
  • Thinking - The model performs inference in thinking mode after deployment.
  • Instruct/Thinking - You can choose whether to enable thinking mode when deploying the model.

Speech Synthesis

CosyVoice

Model Name

Model Code

Model Unit Specification

Hourly Unit Price (CNY)

Monthly Subscription Unit Price (CNY)

cosyvoice-v3-flash

cosyvoice-v3-flash

MU5

¥21

¥10,139

Deployment Creation

DTU Deployment Creation

Before using DTU, enable the DTU feature in the console and apply for a resource quota. Once enabled, select the target model on the Dedicated Deployment page of the Bailian console and choose DTU as the billing method to create a deployment.

WarningDTU deployment does not currently support creation and management via API. Please complete enablement, creation, scaling, and renewal in the Bailian console.

For the basic workflow of general model deployment, see Dedicated Deployment Overview.

The form fields for creating a deployment are shown in the table below.

Parameter

Description

Required

Value Description

Service Name

Name of the deployment service

Yes

Custom

Model

Target model to deploy

Yes

Dropdown selection

Deployment Template

Deployment architecture

No

Dropdown selection, default to the first

Payment Type

Billing method

Yes

Pre-paid (monthly)

Input Throughput Quota

Purchased input TPM capacity

Yes

Integer multiple of baseline input TPM (kTPM)

Output Throughput Quota

Purchased output TPM capacity

Yes

Integer multiple of baseline output TPM (kTPM)

Purchase Duration

Purchase period

Yes

1-12 (integer, months)

Auto Renewal

Auto-renew on expiry

No

On/Off

Single Renewal Duration

Required when auto-renewal is enabled

Yes

1-12 (integer, months)

MU Deployment Configuration

Model Unit

Configuration item

Configuration details

Service name

A custom name for the deployment service.

Model

Select the model to deploy, including platform preset models and fine-tuned models.

Model unit type

Select the deployment specification. Different specifications correspond to different computing power and performance.

Replica count

Set the initial number of deployment replicas, which affects the concurrent processing capability of the service.

Deployment template

Select a deployment template (for example, "single-node deployment"). Different templates correspond to different resource configuration schemes. Available only in the model unit billing mode.

Model inference mode

For some models, when deployed inModel Unit mode, you can configure the inference mode, maximum context, and more.

  • Instruct - The model performs inference in non-thinking mode after deployment.

  • Thinking - The model performs inference in thinking mode after deployment.

Maximum context

TheModel Unit deployment mode of some models supports this setting. The maximum context length depends on the model type.

Service throttling

TheModel Unit deployment mode of some models supports this setting, which can limit the RPM and TPM of model calls.

Capacity Planning

DTU is sold by input/output TPM; the purchased TPM defines the maximum tokens processable per minute. Capacity planning aims to derive the purchase multiple from your business's peak token demand, then verify through load testing that the actually sustainable concurrency and latency meet requirements. Baseline input/output TPM and standard-workload performance references for each model are in the tables above.

Use a load-testing tool such as evalscope to test against your real business scenario (input/output length, concurrency, latency requirements), then determine the purchase multiple against the pricing table.

Capacity Planning Method

  1. Profile peak business metrics: typical request input/output token length, target concurrency, and requirements for first-token latency (TTFT) and per-token generation latency (TPOT).
  2. Estimate peak throughput demand: peak input TPM ≈ concurrency × per-request input length ÷ per-request processing time (minutes); peak output TPM ≈ concurrency × per-request output length ÷ per-request generation time (minutes).
  3. Compute the purchase multiple: input multiple = ⌈peak input TPM ÷ baseline input TPM⌉, output multiple = ⌈peak output TPM ÷ baseline output TPM⌉. DTU requires input and output to scale together, so take the larger of the two as the final multiple.
  4. Verify with load testing: after purchasing the computed multiple, re-test with evalscope under the target workload to confirm actual throughput and latency meet business requirements. If latency is high, raise concurrency within the TPM headroom to improve effective throughput.

If latency requirements are not strict, you can raise concurrency — within the purchased TPM cap — to increase actual throughput.

Business estimation example:

For a model with baseline input 656K TPM and baseline output 82K TPM: business peak profiling shows peak input ~1,300K TPM and peak output ~160K TPM. Compute multiples: input ⌈1300÷656⌉=2, output ⌈160÷82⌉=2; take the larger (2×), purchasing input 1,312K TPM (656K×2) + output 164K TPM (82K×2). Then load-test with evalscope at the peak workload to confirm latency passes. If business volume doubles (peak input ~2,600K, output ~320K), multiples become input ⌈2600÷656⌉=4, output ⌈320÷82⌉=4; purchase 4×: input 2,624K TPM (656K×4) + output 328K TPM (82K×4), and re-test.

Estimation formula:

Purchase multiple = max(⌈peak input TPM ÷ baseline input TPM⌉, ⌈peak output TPM ÷ baseline output TPM⌉)

Scaling, Renewal & Unsubscription

DTU Scaling, Renewal & Unsubscription

The deployment list shows the purchased input/output TPM quotas; the details page shows the TPM capacity details.

Scaling: Click "Scale" in the deployment list to modify the input/output TPM capacity. Scaling up shows the additional amount due; scaling down shows the estimated refund. Input and output must be increased or decreased together; one cannot increase while the other decreases. Only running deployments can be operated.

Renewal: Pre-paid monthly deployments can be renewed upon expiry. The renewal button on the details page enters the renewal flow, supporting an auto-renewal toggle and a single renewal duration.

Unsubscription: Prepaid monthly orders support early termination; the used portion is settled at a 1.2x coefficient for refund. For details, see Refund rules for configuration downgrades. Unsubscription is irreversible; after unsubscription, dedicated resources are released and the service stops.

MU Scaling and Renewal

  • Model unit (billed by duration): Click the Scaling button to self-service, manually adjust the number of instances (replica count). You can also configure an auto-scaling policy (including scaling thresholds, minimum/maximum replica count, scheduled scaling, etc.) via the scaling configuration button in the operation column.
  • Renewal: Prepaid monthly services can be renewed to extend the service time; auto-renewal is supported.
  • Unsubscription: If you cancel early within the first month of a prepaid purchase, the daily unit price is billed at 1.2 times the rate. For details, see Refund rules for configuration downgrades.

FAQ

What is the difference between DTU deployment and Token pay-as-you-go billing?

Token pay-as-you-go bills by token usage on a shared resource pool; DTU uses dedicated GPU resources and bills by input/output TPM, suitable for scenarios requiring stable throughput and dedicated deployment.

What billing methods does DTU deployment support?

Only pre-paid (monthly) billing is supported.

How do I view the usage of purchased TPM?

You can view the purchased input/output TPM quotas in the deployment list and details page of Dedicated Deployment in the Bailian console. See Dedicated Deployment Overview.