DTU Dedicated compute deployment
Dedicated compute deployment offers two options: Dedicated Throughput Unit (DTU) and Model Unit (MU), providing dedicated GPU resources for specified models with performance and throughput guarantees. DTU targets newly released models (billed by input/output TPM × duration); MU is used for existing models (billed by model unit count × duration). This document introduces both options' features, billing, usage flow, and supported models.
Overview
Dedicated compute deployment provides dedicated GPU resources and performance guarantees for specified models, supporting both base models and custom model deployment. Bailian offers two dedicated compute options:
- DTU (Dedicated Throughput Unit): Billed by Input/Output TPM, providing more direct throughput guarantees. Fully managed inference service with dedicated underlying GPU resources maintained by the platform.
- Model Unit (MU): Configures compute power based on usage duration and the number of Model Units, with dedicated resources. Supports custom performance metrics and PD disaggregated computing mode. Billing granularity is model unit count × duration.
DTU and MU are differentiated by model applicability: newly released models use DTU (metered by input/output TPM), while existing models continue to use MU (metered by model unit count). Both provide dedicated compute, resource isolation, and fine-tuned model deployment. For new deployments, choose DTU if the model supports it.
Shared capabilities of both options:
- Dedicated GPU resources; the inference environment is physically isolated from other users.
- Supports base models and custom models (Bailian fine-tuned models or user-uploaded models).
- The platform maintains the underlying operations, no need to manage GPUs.
- No proactive RPM/TPM limits; traffic is bound by actual capacity.
For the general model deployment workflow, see Dedicated Deployment Overview and Provisioned Throughput.
Solution Selection
Bailian offers multiple capacity and billing plans for inference calls. DTU suits scenarios requiring dedicated deployment, fine-tuned model deployment, low latency with high concurrency, or data isolation. The plans are compared in the table below.
For the full comparison of all billing methods including Token pay-as-you-go and PTU, see the main table in Dedicated Deployment Overview.
Comparison Dimension | DTU | PTU | Token Pay-as-you-go | PAI/Lingjun |
|---|---|---|---|---|
Feature | Dedicated compute + fine-tuned models | Throughput guarantee + low latency | Elastic and zero-threshold | Custom runtime |
Resource isolation | Physical isolation | Logical isolation | Shared pool | Physical isolation |
Fine-tuned models | Full-parameter + LoRA | Not supported | LoRA | Supported |
Custom framework | Not supported | Not supported | Not supported | Supported |
Billing mode | TPM quota × duration | TPM quota × duration | Token usage | GPU × duration |
Billing Rules
DTU Billing
DTU is billed separately by input TPM and output TPM, using a pre-paid (monthly) model. Each model has fixed input/output baseline TPM (see table below), and you must purchase in integer multiples of the baseline TPM. Purchase at least 1× baseline TPM for both input and output; for production, 2× or above is recommended for each.
Fees are calculated by the backend based on model, input/output TPM, service region, and purchase duration. Prices for each model are shown in the table below; the starting and scaling prices are both the total for 1× baseline TPM of input and output. Fine-tuned model prices are the same as the corresponding base model, subject to the actual console display.
MU Billing
Cost = Usage duration (hours) × Number of Model Units × Model Unit unit price
The "Model Unit unit price" takes the "Hourly unit price" column in the table below for postpaid scenarios; for monthly subscription prepaid billing, the formula becomes Number of monthly subscriptions × Number of Model Units × Monthly unit price.
- For the first month of a prepaid purchase, if you cancel the subscription early within the first month, the daily unit price (≈ Monthly unit price / 30) will be billed at 1.2 times the rate (less than one day is billed as one day)
NoteThe compute resources for the Model Unit postpaid method are first-come, first-served. If the purchase fails, a full refund will be issued.
Both MU and DTU use pre-paid billing. Early termination settles the used portion at a 1.2x coefficient for refund. For details, see Refund rules for configuration downgrades.
Supported Models and Pricing
DTU Price Table and Performance Baseline
China (Beijing) 2
Model | Spec | Baseline Input (TPM) | Baseline Output (TPM) | Starting Price (CNY/month) | Scaling Price (CNY/month) |
|---|---|---|---|---|---|
qwen3.7-plus-2026-05-26 | Fast | 1,372,000 | 170,000 | 556,466 | 556,466 |
qwen3.6-27b | Standard | 273,000 | 34,000 | 49,118 | 49,118 |
qwen3.5-397b-a17b | Standard | 896,000 | 112,000 | 1,107,904 | 1,107,904 |
qwen3.5-122b-a10b | Fast | 7,288,000 | 904,000 | 2,226,984 | 2,226,984 |
qwen3.5-35b-a3b | Standard | 656,000 | 82,000 | 109,962 | 109,962 |
qwen3.5-27b | Standard | 2,432,000 | 304,000 | 1,108,688 | 1,108,688 |
Efficient | 291,000 | 36,000 | 48,996 | 48,996 | |
qwen3.5-4b | Standard | 1,072,000 | 134,000 | 109,478 | 109,478 |
glm-5.2 | Fast | 1,004,000 | 120,000 | 1,895,260 | 1,895,260 |
glm-5.1 | Standard | 256,000 | 32,000 | 504,704 | 504,704 |
deepseek-v4-flash | Standard | 2,240,000 | 280,000 | 1,107,400 | 1,107,400 |
deepseek-v4-flash-0731 | Fast | 1,676,000 | 213,000 | 277,984 | 277,984 |
The performance reference data for each model under standard workload is shown in the table below, subject to the actual console display.
The following performance reference data was measured at a 0% cache hit rate. In actual use, as the cache hit rate increases, model performance improves accordingly.
Model | Spec | Input Length | Output Length | Cache Hit Rate | First-token Latency (ms) | Per-token Latency (ms) |
|---|---|---|---|---|---|---|
qwen3.7-plus-2026-05-26 | Fast | 16,000 | 2,000 | 0 | 2,418 | 15 |
qwen3.6-27b | Standard | 4,000 | 500 | 0 | 1,292 | 19 |
qwen3.5-397b-a17b | Standard | 4,000 | 500 | 0 | 996 | 27 |
qwen3.5-122b-a10b | Fast | 16,000 | 2,000 | 0 | 568 | 8 |
qwen3.5-35b-a3b | Standard | 4,000 | 500 | 0 | 471 | 10 |
qwen3.5-27b | Standard | 4,000 | 500 | 0 | 703 | 14 |
Efficient | 4,000 | 500 | 0 | 1,448 | 23 | |
qwen3.5-4b | Standard | 4,000 | 500 | 0 | 552 | 6 |
glm-5.2 | Fast | 16,000 | 2,000 | 0 | 1,558 | 15 |
glm-5.1 | Standard | 4,000 | 500 | 0 | 769 | 27 |
deepseek-v4-flash | Standard | 4,000 | 500 | 0 | 651 | 19 |
deepseek-v4-flash-0731 | Fast | 16,000 | 2,000 | 0 | 1,079 | 13 |
MU Price Table
Text generation
Qwen
Model Name | Model Code | Model Unit Specification | Hourly Unit Price (CNY) Minimum billing: minutes | Monthly Subscription Unit Price (CNY) Minimum billing: days |
|---|---|---|---|---|
Qwen3.8-27B | qwen3.8-27b | MU9 x 4 | ¥204 | ¥98,400 |
Qwen3.7-Plus-2026-05-26 | qwen3.7-plus-2026-05-26 | MU2 x 8 | ¥504 | ¥240,288 |
MU3 x 8 | ¥1,096 | ¥527,752 | ||
Qwen3.6-35B-A3B | qwen3.6-35b-a3b | MU1 x 8 | ¥432 | ¥208,944 |
MU2 x 8 | ¥504 | ¥240,288 | ||
MU3 x 8 | ¥1,096 | ¥527,752 | ||
MU8 x 1 | ¥47 | ¥22,400 | ||
MU9 x 1 | ¥51 | ¥24,600 | ||
Qwen3.6-27B | qwen3.6-27b | MU9 x 1 | ¥51 | ¥24,600 |
Qwen3.6-Flash-2026-04-16 | qwen3.6-flash-2026-04-16 | MU1 x 2 | ¥108 | ¥52,236 |
MU3 x 8 | ¥1,096 | ¥527,752 | ||
Qwen3.6-Plus-2026-04-02 | qwen3.6-plus-2026-04-02 | MU1 x 8 MU1 x 16(PD disaggregation mode) | ¥432 PD disaggregation mode:¥864 | ¥208,944 PD disaggregation mode:¥417,888 |
MU2 x 8 | ¥504 | ¥240,288 | ||
Qwen3.5-397B-A17B | qwen3.5-397b-a17b | MU3 x 8 MU3 x 16(PD disaggregation mode) | ¥1,096 PD disaggregation mode:¥2,192 | ¥527,752 PD disaggregation mode:¥1,055,504 |
MU6 x 16 | ¥400 | ¥193,424 | ||
Qwen3.5-122B-A10B | qwen3.5-122b-a10b | MU1 x 4 | ¥216 | ¥104,472 |
MU6 x 16 | ¥400 | ¥193,424 | ||
Qwen3.5-35B-A3B | qwen3.5-35b-a3b | MU1 x 2 | ¥108 | ¥52,236 |
MU2 x 8 | ¥504 | ¥240,288 | ||
MU3 x 8 | ¥1,096 | ¥527,752 | ||
MU9 x 1 | ¥51 | ¥24,600 | ||
Qwen3.5-27B | qwen3.5-27b | MU2 x 8 | ¥504 | ¥240,288 |
MU3 x 8 | ¥1,096 | ¥527,752 | ||
MU8 x 1 | ¥47 | ¥22,400 | ||
MU9 x 1 | ¥51 | ¥24,600 | ||
Qwen3.5-9B | qwen3.5-9b | MU1 x 2 | ¥108 | ¥52,236 |
MU2 x 8 | ¥504 | ¥240,288 | ||
Qwen3.5-Flash-2026-02-23 | qwen3.5-flash-2026-02-23 | MU1 x 2 | ¥108 | ¥52,236 |
Qwen3.5-Plus-2026-02-15 | qwen3.5-plus-2026-02-15 | MU1 x 8 MU1 x 16(PD disaggregation mode) | ¥432 PD disaggregation mode:¥864 | ¥208,944 PD disaggregation mode:¥417,888 |
MU2 x 8 | ¥504 | ¥240,288 | ||
MU3 x 8 MU3 x 16(PD disaggregation mode) | ¥1,096 PD disaggregation mode:¥2,192 | ¥527,752 PD disaggregation mode:¥1,055,504 | ||
Qwen3-235B-A22B-Instruct-2507 | qwen3-235b-a22b-instruct-2507 | MU1 x 4 | ¥216 | ¥104,472 |
MU2 x 8 | ¥504 | ¥240,288 | ||
MU3 x 8 | ¥1,096 | ¥527,752 | ||
Qwen3-32B | qwen3-32b | MU6 x 16 | ¥400 | ¥193,424 |
Qwen3-30B-A3B-Thinking-2507 | qwen3-30b-a3b-thinking-2507 | MU1 x 2 | ¥108 | ¥52,236 |
Qwen3-4B | qwen3-4b | MU1 x 2 | ¥108 | ¥52,236 |
MU5 x 1 | ¥21 | ¥10,139 | ||
Qwen3-Embedding-0.6B | qwen3-embedding-0.6b | MU5 x 1 | ¥21 | ¥10,139 |
MU6 x 1 | ¥25 | ¥12,089 | ||
Qwen3-MoE-Rerank-0.6B | qwen3-moe-rerank-0.6b | MU5 x 1 | ¥21 | ¥10,139 |
Qwen3-Rerank-0.6B | qwen3-rerank-0.6b | MU5 x 1 | ¥21 | ¥10,139 |
MU6 x 1 | ¥25 | ¥12,089 | ||
Qwen3-Max-2025-09-23 | qwen3-max-2025-09-23 | MU2 x 8 | ¥504 | ¥240,288 |
MU3 x 8 | ¥1,096 | ¥527,752 | ||
Qwen3-Rerank | qwen3-rerank | MU5 x 1 | ¥21 | ¥10,139 |
Qwen2.5-Open-Source-72B | qwen2.5-72b-instruct | MU1 x 8 | ¥432 | ¥208,944 |
Qwen2.5-Open-Source-14B | qwen2.5-14b-instruct | MU1 x 2 | ¥108 | ¥52,236 |
Qwen2.5-Open-Source-7B | qwen2.5-7b-instruct | MU1 x 2 | ¥108 | ¥52,236 |
MU5 x 1 | ¥21 | ¥10,139 | ||
Qwen-Plus-2025-07-28 | qwen-plus-2025-07-28 | MU1 x 4 MU1 x 16(PD disaggregation mode) | ¥216 PD disaggregation mode:¥864 | ¥104,472 PD disaggregation mode:¥417,888 |
Qwen-Plus-2025-12-01 | qwen-plus-2025-12-01 | MU1 x 4 | ¥216 | ¥104,472 |
Qwen-Plus-Character-2025-11-06 | qwen-plus-character-2025-11-06 | MU1 x 4 | ¥216 | ¥104,472 |
GLM
Model Name | Model Code | Model Unit Specification | Hourly Unit Price (CNY) Minimum billing: minutes | Monthly Subscription Unit Price (CNY) Minimum billing: days |
|---|---|---|---|---|
GLM-5.1 | glm-5.1 | MU2 x 8 | ¥504 | ¥240,288 |
MU3 x 16(PD disaggregation mode) | PD disaggregation mode:¥2,192 | PD disaggregation mode:¥1,055,504 | ||
MU6 x 16 | ¥400 | ¥193,424 | ||
GLM-5 | glm-5 | MU3 x 16(PD disaggregation mode) | PD disaggregation mode:¥2,192 | PD disaggregation mode:¥1,055,504 |
GLM-4.7 | glm-4.7 | MU6 x 32(PD disaggregation mode) | PD disaggregation mode:¥800 | PD disaggregation mode:¥386,848 |
DeepSeek
Model Name | Model Code | Model Unit Specification | Hourly Unit Price (CNY) Minimum billing: minutes | Monthly Subscription Unit Price (CNY) Minimum billing: days |
|---|---|---|---|---|
DeepSeek-v4-Flash | deepseek-v4-flash | MU1 x 8 | ¥432 | ¥208,944 |
MU3 x 8 | ¥1,096 | ¥527,752 | ||
DeepSeek-v3.2 | deepseek-v3.2 | MU2 x 16(PD disaggregation mode) | PD disaggregation mode:¥1,008 | PD disaggregation mode:¥480,576 |
More models
Model Name | Model Code | Model Unit Specification | Hourly Unit Price (CNY) Minimum billing: minutes | Monthly Subscription Unit Price (CNY) Minimum billing: days |
|---|---|---|---|---|
Kimi-K2.5 | kimi-k2.5 | MU2 x 8 | ¥504 | ¥240,288 |
Model type:
- Instruct - The model performs inference in non-thinking mode after deployment.
- Thinking - The model performs inference in thinking mode after deployment.
Model deployment type:
-
PD disaggregation mode - Reduces first-token latency and increases throughput.
For models deployed in this mode, during inference the first-token computation (Prefill) and the subsequent token computation (Decode) are split across different compute nodes for execution.
Multimodal
Qwen-VL
Model Name | Model Code | Model Unit Specification | Hourly Unit Price (CNY) Minimum billing: minutes | Monthly Subscription Unit Price (CNY) Minimum billing: days |
|---|---|---|---|---|
Qwen3-VL-235B-A22B-Thinking | qwen3-vl-235b-a22b-thinking | MU1 x 8 | ¥432 | ¥208,944 |
MU2 x 8 | ¥504 | ¥240,288 | ||
MU3 x 8 | ¥1,096 | ¥527,752 | ||
Qwen3-VL-32B-Instruct | qwen3-vl-32b-instruct | MU2 x 8 | ¥504 | ¥240,288 |
MU3 x 8 | ¥1,096 | ¥527,752 | ||
Qwen3-VL-8B-Instruct | qwen3-vl-8b-instruct | MU1 x 2 | ¥108 | ¥52,236 |
MU5 x 1 | ¥21 | ¥10,139 | ||
Qwen3-VL-4B-Instruct | qwen3-vl-4b-instruct | MU1 x 2 | ¥108 | ¥52,236 |
Qwen3-VL-2B-Instruct | qwen3-vl-2b-instruct | MU5 x 1 | ¥21 | ¥10,139 |
Qwen3-VL-Embedding-2B | qwen3-vl-embedding-2b | MU5 x 1 | ¥21 | ¥10,139 |
Qwen3-VL-Flash-2025-10-15 | qwen3-vl-flash-2025-10-15 | MU1 x 4 | ¥216 | ¥104,472 |
Qwen3-VL-Plus-2025-09-23 | qwen3-vl-plus-2025-09-23 | MU1 x 4 | ¥216 | ¥104,472 |
Qwen-VL-Max-2025-08-13 | qwen-vl-max-2025-08-13 | MU6 x 4 | ¥100 | ¥48,356 |
Qwen Omni
Model Name | Model Code | Model Unit Specification | Hourly Unit Price (CNY) Minimum billing: minutes | Monthly Subscription Unit Price (CNY) Minimum billing: days |
|---|---|---|---|---|
Qwen3.5-Omni-Flash | qwen3.5-omni-flash | MU8 x 1 | ¥47 | ¥22,400 |
MU9 x 1 | ¥51 | ¥24,600 |
Model Type:
- Instruct - The model performs inference in non-thinking mode after deployment.
- Thinking - The model performs inference in thinking mode after deployment.
- Instruct/Thinking - You can choose whether to enable thinking mode when deploying the model.
Speech Synthesis
CosyVoiceModel Name | Model Code | Model Unit Specification | Hourly Unit Price (CNY) | Monthly Subscription Unit Price (CNY) |
|---|---|---|---|---|
cosyvoice-v3-flash | cosyvoice-v3-flash | MU5 | ¥21 | ¥10,139 |
Deployment Creation
DTU Deployment Creation
Before using DTU, enable the DTU feature in the console and apply for a resource quota. Once enabled, select the target model on the Dedicated Deployment page of the Bailian console and choose DTU as the billing method to create a deployment.
WarningDTU deployment does not currently support creation and management via API. Please complete enablement, creation, scaling, and renewal in the Bailian console.
For the basic workflow of general model deployment, see Dedicated Deployment Overview.
The form fields for creating a deployment are shown in the table below.
Parameter | Description | Required | Value Description |
|---|---|---|---|
Service Name | Name of the deployment service | Yes | Custom |
Model | Target model to deploy | Yes | Dropdown selection |
Deployment Template | Deployment architecture | No | Dropdown selection, default to the first |
Payment Type | Billing method | Yes | Pre-paid (monthly) |
Input Throughput Quota | Purchased input TPM capacity | Yes | Integer multiple of baseline input TPM (kTPM) |
Output Throughput Quota | Purchased output TPM capacity | Yes | Integer multiple of baseline output TPM (kTPM) |
Purchase Duration | Purchase period | Yes | 1-12 (integer, months) |
Auto Renewal | Auto-renew on expiry | No | On/Off |
Single Renewal Duration | Required when auto-renewal is enabled | Yes | 1-12 (integer, months) |
MU Deployment Configuration
Model Unit
Configuration item | Configuration details |
|---|---|
Service name | A custom name for the deployment service. |
Model | Select the model to deploy, including platform preset models and fine-tuned models. |
Model unit type | Select the deployment specification. Different specifications correspond to different computing power and performance. |
Replica count | Set the initial number of deployment replicas, which affects the concurrent processing capability of the service. |
Deployment template | Select a deployment template (for example, "single-node deployment"). Different templates correspond to different resource configuration schemes. Available only in the model unit billing mode. |
Model inference mode | For some models, when deployed inModel Unit mode, you can configure the inference mode, maximum context, and more.
|
Maximum context | TheModel Unit deployment mode of some models supports this setting. The maximum context length depends on the model type. |
Service throttling | TheModel Unit deployment mode of some models supports this setting, which can limit the RPM and TPM of model calls. |
Capacity Planning
DTU is sold by input/output TPM; the purchased TPM defines the maximum tokens processable per minute. Capacity planning aims to derive the purchase multiple from your business's peak token demand, then verify through load testing that the actually sustainable concurrency and latency meet requirements. Baseline input/output TPM and standard-workload performance references for each model are in the tables above.
Use a load-testing tool such as evalscope to test against your real business scenario (input/output length, concurrency, latency requirements), then determine the purchase multiple against the pricing table.
Capacity Planning Method
- Profile peak business metrics: typical request input/output token length, target concurrency, and requirements for first-token latency (TTFT) and per-token generation latency (TPOT).
- Estimate peak throughput demand: peak input TPM ≈ concurrency × per-request input length ÷ per-request processing time (minutes); peak output TPM ≈ concurrency × per-request output length ÷ per-request generation time (minutes).
- Compute the purchase multiple: input multiple = ⌈peak input TPM ÷ baseline input TPM⌉, output multiple = ⌈peak output TPM ÷ baseline output TPM⌉. DTU requires input and output to scale together, so take the larger of the two as the final multiple.
- Verify with load testing: after purchasing the computed multiple, re-test with evalscope under the target workload to confirm actual throughput and latency meet business requirements. If latency is high, raise concurrency within the TPM headroom to improve effective throughput.
If latency requirements are not strict, you can raise concurrency — within the purchased TPM cap — to increase actual throughput.
Business estimation example:
For a model with baseline input 656K TPM and baseline output 82K TPM: business peak profiling shows peak input ~1,300K TPM and peak output ~160K TPM. Compute multiples: input ⌈1300÷656⌉=2, output ⌈160÷82⌉=2; take the larger (2×), purchasing input 1,312K TPM (656K×2) + output 164K TPM (82K×2). Then load-test with evalscope at the peak workload to confirm latency passes. If business volume doubles (peak input ~2,600K, output ~320K), multiples become input ⌈2600÷656⌉=4, output ⌈320÷82⌉=4; purchase 4×: input 2,624K TPM (656K×4) + output 328K TPM (82K×4), and re-test.
Estimation formula:
Purchase multiple = max(⌈peak input TPM ÷ baseline input TPM⌉, ⌈peak output TPM ÷ baseline output TPM⌉)
Scaling, Renewal & Unsubscription
DTU Scaling, Renewal & Unsubscription
The deployment list shows the purchased input/output TPM quotas; the details page shows the TPM capacity details.
Scaling: Click "Scale" in the deployment list to modify the input/output TPM capacity. Scaling up shows the additional amount due; scaling down shows the estimated refund. Input and output must be increased or decreased together; one cannot increase while the other decreases. Only running deployments can be operated.
Renewal: Pre-paid monthly deployments can be renewed upon expiry. The renewal button on the details page enters the renewal flow, supporting an auto-renewal toggle and a single renewal duration.
Unsubscription: Prepaid monthly orders support early termination; the used portion is settled at a 1.2x coefficient for refund. For details, see Refund rules for configuration downgrades. Unsubscription is irreversible; after unsubscription, dedicated resources are released and the service stops.
MU Scaling and Renewal
- Model unit (billed by duration): Click the Scaling button to self-service, manually adjust the number of instances (replica count). You can also configure an auto-scaling policy (including scaling thresholds, minimum/maximum replica count, scheduled scaling, etc.) via the scaling configuration button in the operation column.
- Renewal: Prepaid monthly services can be renewed to extend the service time; auto-renewal is supported.
- Unsubscription: If you cancel early within the first month of a prepaid purchase, the daily unit price is billed at 1.2 times the rate. For details, see Refund rules for configuration downgrades.
FAQ
What is the difference between DTU deployment and Token pay-as-you-go billing?Token pay-as-you-go bills by token usage on a shared resource pool; DTU uses dedicated GPU resources and bills by input/output TPM, suitable for scenarios requiring stable throughput and dedicated deployment.
What billing methods does DTU deployment support?Only pre-paid (monthly) billing is supported.
How do I view the usage of purchased TPM?You can view the purchased input/output TPM quotas in the deployment list and details page of Dedicated Deployment in the Bailian console. See Dedicated Deployment Overview.