Throughput Reservation Billing
This article describes the billing rules, pricing, refunds, and capacity conversion rules of Alibaba Cloud Model Studio throughput reservation (formerly TPM reservation).
Billing Rules
Throughput reservation purchases dedicated inference throughput via subscription. Calls within the reserved capacity incur no additional charges; the excess is handled per the overflow strategy. The dedicated model Code (invocation identifier, used to replace the model parameter of the API) generated after you create a throughput reservation is used for billing and invocation statistics.
| Performance Mode | Standard mode: same TPS as the standard API. High-speed mode: 1.5~2× TPS improvement over the standard API (i.e., PTU model deployment). |
| Billing Method | Subscription (by day): supported for both standard and high-speed modes, billed by natural day. Input is priced per 10K-TPM and output per 1K-TPM. Subscription (by 8-hour time slot): only for standard mode, fixed 8 hours, effective immediately. For more information, see 8-Hour Time Slot Reservation Billing. Pay-as-you-go (by hour): only for high-speed mode, billed by actual duration, no duration purchase required. |
| Billing Formula | Subscription: Pay-as-you-go (hour granularity, minute precision): |
| Billed Upon Activation | Billing starts as soon as the reservation is successfully created; calls within the reserved capacity incur no additional charges; capacity fees are billed based on the purchased amount, not reduced by actual usage. |
| New Purchase Capacity Expiration Time | UTC+8: if purchased before 22:00 on the current day, the expiration time is 23:59:59 on the current day. If purchased after 22:00, the expiration time is 23:59:59 on the next day. |
| Input Capacity | Starting from 200 kTPM, step size 10 kTPM. |
| Output Capacity | Starting from 20 kTPM, step size 1 kTPM. |
| Duration | Supports 1~30, 60, 90, 120, 365 days. |
| Renew |
|
| Auto-Renew on Expiration | Auto-deducts and renews 1 day before expiration, with up to 4 attempts between 8:20–20:00 on that day. Enabled by default, can be disabled. |
Standard Mode
Taking Qwen3.8-Max (China (Beijing)) as an example (200 input + 100 output kTPM, 1 day):
- Assume standard mode input unit price ¥121/10K-TPM·day, output unit price ¥36.29/1K-TPM·day.
- Input fee =
121 × 200 ÷ 10= 2420 CNY. - Output fee =
36.29 × 100= 3629 CNY. - Reservation fee =
1 × (2420 + 3629)= 6049 CNY.
High-Speed Mode
Taking Qwen3.8-Max (China (Beijing)) as an example (200 input + 100 output kTPM, 1 day):
- Assume high-speed mode subscription input unit price ¥345.6/10K-TPM·day, output unit price ¥103.68/1K-TPM·day.
- Input fee =
345.6 × 200 ÷ 10= 6912 CNY. - Output fee =
103.68 × 100= 10368 CNY. - Reservation fee =
1 × (6912 + 10368)= 17280 CNY.
Unit Description
- kTPM = 1,000 Tokens/minute, the purchase quantity unit.
- 10K-TPM = 10,000 Tokens/minute = 10 kTPM. Input unit price is quoted per 10K-TPM; in the formula kTPM ÷ 10 converts to 10K-TPM then multiplied by the unit price; output unit price is quoted per 1K-TPM (1K-TPM = 1 kTPM), multiplied directly.
Supported Models and Pricing
The subscription prices for standard mode (same TPS as pay-as-you-go invocations) and high-speed mode (1.5~2× TPS) are as follows:
China (Beijing)
Standard Mode
Model | Subscription | |
|---|---|---|
Input (10K-TPM·day) | Output (1K-TPM·day) | |
Qwen3.8-Max | ¥121.00 | ¥36.29 |
Qwen3.7-Flash-2026-07-15 | ¥2.00 | ¥0.81 |
Qwen3.7-Max-2026-05-20 | ¥121.00 | ¥36.29 |
Qwen3.7-Plus-2026-05-26 | ¥20.20 | ¥8.06 |
Qwen3.6-Flash-2026-04-16 | ¥12.10 | ¥7.26 |
GLM-5.3 | ¥80.60 | ¥28.22 |
GLM-5.2 | ¥80.60 | ¥28.22 |
GLM-5.1 | ¥60.50 | ¥24.19 |
DeepSeek-v4-Flash | ¥10.10 | ¥2.02 |
DeepSeek-v4-Flash-0731 | ¥30.20 | ¥9.07 |
DeepSeek-v4-Pro | ¥121.00 | ¥24.19 |
DeepSeek-v4-Pro-0813 | ¥90.70 | ¥27.22 |
Kimi-K2.6 | ¥65.50 | ¥27.22 |
High-Speed Mode
Model | Pay-as-you-go | Subscription | ||
|---|---|---|---|---|
Input (10K-TPM·hour) | Output (1K-TPM·hour) | Input (10K-TPM·day) | Output (1K-TPM·day) | |
Qwen3.8-Max | ¥28.8 | ¥8.64 | ¥345.6 | ¥103.68 |
Qwen3.7-Flash-2026-07-15 | ¥0.48 | ¥0.19 | ¥5.76 | ¥2.3 |
Qwen3.7-Max-2026-05-20 | ¥28.8 | ¥8.64 | ¥345.6 | ¥103.68 |
Qwen3.7-Plus-2026-05-26 | ¥4.8 | ¥1.92 | ¥57.6 | ¥23.04 |
Qwen3.6-Flash-2026-04-16 | ¥2.88 | ¥1.73 | ¥34.56 | ¥20.74 |
GLM-5.3 | — | — | — | — |
GLM-5.2 | ¥28.8 | ¥10.08 | ¥345.6 | ¥120.96 |
GLM-5.1 | — | — | — | — |
DeepSeek-v4-Flash | ¥3.6 | ¥0.72 | ¥43.2 | ¥8.64 |
DeepSeek-v4-Flash-0731 | ¥7.2 | ¥1.44 | ¥86.4 | ¥17.28 |
DeepSeek-v4-Pro | ¥43.2 | ¥8.64 | ¥518.4 | ¥103.68 |
DeepSeek-v4-Pro-0813 | — | — | — | — |
Kimi-K2.6 | — | — | — | — |
Singapore
Standard Mode
Model | Subscription | |
|---|---|---|
Input (10K-TPM·day) | Output (1K-TPM·day) | |
Qwen3.8-Max | ¥151.10 | ¥45.32 |
Qwen3.7-Flash-2026-07-15 | ¥2.30 | ¥0.98 |
Qwen3.7-Max-2026-05-20 | ¥188.90 | ¥56.66 |
Qwen3.7-Plus-2026-05-26 | ¥30.20 | ¥12.09 |
Qwen3.6-Flash-2026-04-16 | ¥18.90 | ¥11.33 |
GLM-5.3 | ¥102.90 | ¥32.34 |
GLM-5.2 | ¥105.80 | ¥33.24 |
GLM-5.1 | ¥105.80 | ¥33.24 |
DeepSeek-v4-Flash | ¥15.10 | ¥3.02 |
DeepSeek-v4-Flash-0731 | ¥31.40 | ¥9.43 |
DeepSeek-v4-Pro | ¥181.30 | ¥36.26 |
DeepSeek-v4-Pro-0813 | ¥94.30 | ¥28.29 |
High-Speed Mode
Model | Pay-as-you-go | Subscription | ||
|---|---|---|---|---|
Input (10K-TPM·hour) | Output (1K-TPM·hour) | Input (10K-TPM·day) | Output (1K-TPM·day) | |
Qwen3.8-Max | ¥35.97 | ¥10.79 | ¥431.7 | ¥129.5 |
Qwen3.7-Flash-2026-07-15 | ¥0.54 | ¥0.23 | ¥6.47 | ¥2.81 |
Qwen3.7-Max-2026-05-20 | ¥44.97 | ¥13.49 | ¥539.6 | ¥161.87 |
Qwen3.7-Plus-2026-05-26 | ¥7.19 | ¥2.88 | ¥86.3 | ¥34.53 |
Qwen3.6-Flash-2026-04-16 | — | — | — | — |
GLM-5.3 | — | — | — | — |
GLM-5.2 | ¥37.8 | ¥11.87 | ¥453.3 | ¥142.45 |
GLM-5.1 | — | — | — | — |
DeepSeek-v4-Flash | ¥5.4 | ¥1.08 | ¥64.8 | ¥12.95 |
DeepSeek-v4-Flash-0731 | ¥10.79 | ¥2.16 | ¥129.5 | ¥25.9 |
DeepSeek-v4-Pro | ¥64.75 | ¥12.95 | ¥777 | ¥155.4 |
DeepSeek-v4-Pro-0813 | — | — | — | — |
NoteCache hit rate only affects input kTPM, not output.
Capacity Conversion Rules
Long input is converted by a tiered coefficient, and the cache-hit portion is converted by a cache conversion coefficient; the converted amount is deducted from the purchased quota. Parameters for each model are as follows.
Model | Max Input Token | Cache Conversion Coefficient | Long Input Tiered Coefficient |
|---|---|---|---|
qwen3.8-max | 1M | 0.125 | No tier (1.0) |
qwen3.7-flash-2026-07-15 | 1M | 0.2 | Same for input and output |
qwen3.7-max-2026-05-20 | 256K | 0.1 | No tier (1.0) |
qwen3.7-plus-2026-05-26 | 256K | 0.2 | No tier (1.0) |
qwen3.6-flash-2026-04-16 | 256K | 1 | No tier (1.0) |
glm-5.3 | 1M | 0.25 | No tier (1.0) |
glm-5.2 | 1M | 0.25 | No tier (1.0) |
glm-5.1 | 200K | 0.2 | (0,32K] Input 1x / Output 1x |
deepseek-v4-flash | 256K | 0.1 | No tier (1.0) |
deepseek-v4-flash-0731 | 1M | 0.1 | No tier (1.0) |
deepseek-v4-pro | 256K | 0.08 | No tier (1.0) |
deepseek-v4-pro-0813 | 1M | 0.1 | No tier (1.0) |
kimi-k2.6 | 256K | 0.2 | No tier (1.0) |
Capacity conversion examples are as follows.
Tiered · Qwen3.7-Flash
Long input 50K Tokens, no cache hit.
- First 32K portion (coefficient 1x):
32K × 1= 32 kTPM. - Exceeding 32K portion (50K − 32K = 18K, coefficient 3x):
18K × 3= 54 kTPM. - Total consumption:
32 + 54= 86 kTPM.
Same 50K Token input, with the first 30K cache-hit (cache conversion coefficient 0.2).
- Cache-hit portion (within the first 32K tier, coefficient 1x, then multiplied by conversion coefficient 0.2):
30K × 1 × 0.2= 6 kTPM. - Non-hit portion (remaining 2K within the first 32K tier = 32K − 30K, coefficient 1x):
2K × 1= 2 kTPM. - Exceeding 32K portion (50K − 32K = 18K, coefficient 3x):
18K × 3= 54 kTPM. - Total consumption:
6 + 2 + 54= 62 kTPM (about 28% less than without cache).
No Tier · GLM-5.2
Long input 50K Tokens, no cache hit.
- No tier, all converted at coefficient 1x:
50K × 1= 50 kTPM.
Same 50K Token input, with the first 30K cache-hit (cache conversion coefficient 0.25).
- Cache-hit portion (coefficient 1x, then multiplied by conversion coefficient 0.25):
30K × 0.25= 7.5 kTPM. - Non-hit portion (remaining 20K = 50K − 30K, coefficient 1x):
20K × 1= 20 kTPM. - Total consumption:
7.5 + 20= 27.5 kTPM (about 45% less than without cache).
8-Hour Time Slot Reservation Billing
8-hour time slot reservation (8h Block) provides a continuous 8-hour TPM capacity certainty guarantee for scenarios where the load is concentrated in specific time slots.
Purchase and Billing Rules
- Billing method: Subscription (by 8-hour time slot), only for standard mode; enjoy a 10% discount, subject to the price displayed in the Model Studio console.
- Orderable time slots: daily 20:00–next day 04:00 (the order moment must fall within this time slot).
- Effective: effective immediately after purchase, fixed duration of 8 hours, can cross natural days; the effective start point is rounded down to the current whole hour based on the purchase moment and cannot be customized; less than 1 hour is counted as 1 hour.
- Two purchase forms:
- New independent reservation: generates a dedicated model Code.
- Added as an add-on capacity package to an existing reservation: reuses the base reservation's dedicated model Code, increasing capacity within that 8-hour window. For more information, see Add-on Capacity Package Billing.
New 8-Hour Time Slot Reservation
Taking Qwen3.8-Max (China (Beijing)) standard mode as an example (200 input + 100 output kTPM, 8-hour time slot, fee multiplier 0.9; assume input unit price ¥121/10K-TPM·day, output unit price ¥36.29/1K-TPM·day):
- By-day-basis fee =
121 × 200 ÷ 10 + 36.29 × 100=2420 + 3629= 6049 CNY (1 day). - 8-hour time slot fee =
6049 × (8 ÷ 24) × 0.9≈ 1814.70 CNY (10% discount).
Add-on 8-Hour Add-on Capacity Package
Adding an 8-hour add-on capacity package on top of an existing by-day reservation, taking Qwen3.8-Max (China (Beijing)) standard mode as an example (adding 8-hour capacity of 200 input + 100 output kTPM, fee multiplier 0.9; assume input unit price ¥121/10K-TPM·day, output unit price ¥36.29/1K-TPM·day):
- By-day-basis fee of the added capacity =
121 × 200 ÷ 10 + 36.29 × 100= 6049 CNY (1-day basis). - 8-hour add-on package fee =
6049 × (8 ÷ 24) × 0.9≈ 1814.70 CNY. - After adding, the capacity increases for the same dedicated model Code within that 8-hour window; outside the window, the add-on capacity becomes invalid.
Limits
- Scale-out/scale-in and renew/unsubscribe are not supported.
- Capacity automatically becomes invalid after the time slot expires.
- Orders cannot be placed between 04:00–20:00 on the current day.
Expiration and Lifecycle
The service stops immediately upon expiration, and resources are released 2 hours after the stop.
Phase (after expiration) | Instance Status | Can Invoke/Renew |
|---|---|---|
0~2 hours | Stopped | Cannot invoke, can still renew |
After 2 hours | Released | Cannot be recovered |
It is recommended to enable Auto-Renew on Expiration in advance to avoid service interruption.
Scale Out
Subscription Instances
- Subscription purchase granularity is by day, billing precision is by hour.
- Newly added capacity from scale-out is billed by the remaining validity period, converted to days: remaining hours ÷ 24, rounded up.
- Less than 1 day (24 hours) remaining is still counted as 1 day (e.g., 5 hours remaining counts as 1 day).
Scale-Out Fee = Rounded Up(Remaining Hours ÷ 24) × (Input Capacity Difference (kTPM) ÷ 10 × Input Unit Price + Output Capacity Difference (kTPM) × Output Unit Price)
5 Hours Remaining · Scale Out 100 Input + 50 Output kTPM
Taking Qwen3.8-Max (China (Beijing)) standard mode as an example (assume daily unit price ¥121/10K-TPM·day, output unit price ¥36.29/1K-TPM·day), original 100 input + 50 output kTPM, scale out 100 input + 50 output kTPM, 5 hours remaining:
- Remaining days =
Rounded Up(5 ÷ 24)= 1 day (5 hours is less than 1 day, rounded up). - Input scale-out fee =
1 × 100 ÷ 10 × 121= 1210 CNY. - Output scale-out fee =
1 × 50 × 36.29= 1814.5 CNY. - Scale-out fee =
1210 + 1814.5= 3024.5 CNY.
500 Hours Remaining · Scale Out 100 Input + 50 Output kTPM
Taking Qwen3.8-Max (China (Beijing)) standard mode as an example (assume daily unit price ¥121/10K-TPM·day, output unit price ¥36.29/1K-TPM·day), original 100 input + 50 output kTPM, scale out 100 input + 50 output kTPM, 500 hours remaining:
- Remaining days =
Rounded Up(500 ÷ 24)= 21 days. - Input scale-out fee =
21 × 100 ÷ 10 × 121= 25410 CNY. - Output scale-out fee =
21 × 50 × 36.29= 37904.5 CNY. - Scale-out fee =
25410 + 37904.5= 63314.5 CNY.
Pay-as-you-go Instances
- Pay-as-you-go purchase granularity is by hour, billing precision is by minute.
- Scale-out is billed by actual usage duration; less than 1 hour (60 minutes) is billed by actual minutes.
- Fee = Usage Minutes ÷ 60 × Unit Price × Capacity (e.g., 30 minutes billed as 0.5 hour).
Scale-Out Fee = (Usage Minutes ÷ 60) × (Pay-as-you-go Input Unit Price × Input Capacity Difference (kTPM) ÷ 10 + Pay-as-you-go Output Unit Price × Output Capacity Difference (kTPM))
30 Minutes Used · Less Than 1 Hour
Taking Qwen3.8-Max (China (Beijing)) high-speed mode as an example (assume pay-as-you-go input unit price ¥28.8/10K-TPM·hour, output unit price ¥8.64/1K-TPM·hour), scale out 100 input + 50 output kTPM, used 30 minutes:
- Input scale-out fee =
(30 ÷ 60) × 28.8 × (100 ÷ 10)=0.5 × 28.8 × 10= 144 CNY. - Output scale-out fee =
(30 ÷ 60) × 8.64 × 50=0.5 × 8.64 × 50= 216 CNY. - Scale-out fee =
144 + 216= 360 CNY.
10 Hours Used
Taking Qwen3.8-Max (China (Beijing)) high-speed mode as an example (assume pay-as-you-go input unit price ¥28.8/10K-TPM·hour, output unit price ¥8.64/1K-TPM·hour), scale out 100 input + 50 output kTPM, used 10 hours:
- Input scale-out fee =
(600 ÷ 60) × 28.8 × (100 ÷ 10)=10 × 28.8 × 10= 2880 CNY. - Output scale-out fee =
(600 ÷ 60) × 8.64 × 50=10 × 8.64 × 50= 4320 CNY. - Scale-out fee =
2880 + 4320= 7200 CNY.
Scale In
Subscription Instances
- Subscription purchase granularity is by day, billing precision is by minute (settled amount rounded up by hour).
- Settled amount = Rounded Up(used hours) × Penalty coefficient.
- New spec fee is precise to the minute; used ≤30 days at 1.2×, >30 days at 1.0×.
Refund = max(Original Spec Remaining Fee - New Spec Purchase Fee, 0)
- Original Spec Remaining Fee = Paid Amount - Settled Amount
- Settled Amount =
(Daily Unit Price ÷ 24) × Rounded Up(Used Hours) × Penalty Coefficient × Original Capacity (kTPM) ÷ 10 - New Spec Purchase Fee =
Daily Unit Price × (Remaining Hours ÷ 24) × New Capacity (kTPM) ÷ 10
NoteZeroing (capacity adjusted to 0) no longer incurs capacity fees; the dedicated model Code is retained, and it is treated as a downgrade settled by the penalty coefficient; remaining hours are precise to the minute (e.g., 27 hours 37 minutes ≈ 27.62 hours).
Purchased 30 Days · Used 20 Days (≤30 Days, 1.2×)
Taking Qwen3.8-Max (China (Beijing)) standard mode as an example (assume daily unit price ¥121/10K-TPM·day, output unit price ¥36.29/1K-TPM·day), original 100 kTPM purchased 30 days, scaled in to 50 kTPM. Paid 36300 CNY = 100 ÷ 10 × 30 × 121, used 20 days (480 hours), remaining 240 hours:
- Settled amount =
(121 ÷ 24) × 480 × 1.2 × 100 ÷ 10= 29040 CNY. - Original spec remaining fee =
36300 - 29040= 7260 CNY. - New spec purchase fee =
121 × (240 ÷ 24) × 50 ÷ 10= 6050 CNY. - Refund =
max(7260 - 6050, 0)= 1210 CNY.
Purchased 60 Days · Used 40 Days (>30 Days, 1.0×)
Taking Qwen3.8-Max (China (Beijing)) standard mode as an example (assume daily unit price ¥121/10K-TPM·day, output unit price ¥36.29/1K-TPM·day), original 100 kTPM purchased 60 days, scaled in to 50 kTPM. Paid 72600 CNY = 100 ÷ 10 × 60 × 121, used 40 days (960 hours), remaining 480 hours:
- Settled amount =
(121 ÷ 24) × 960 × 1.0 × 100 ÷ 10= 48400 CNY. - Original spec remaining fee =
72600 - 48400= 24200 CNY. - New spec purchase fee =
121 × (480 ÷ 24) × 50 ÷ 10= 12100 CNY. - Refund =
max(24200 - 12100, 0)= 12100 CNY.
Zeroing (Capacity Adjusted to 0)
Taking Qwen3.8-Max (China (Beijing)) standard mode as an example (assume daily unit price ¥121/10K-TPM·day, original 100 kTPM, purchased 30 days, used 20 days):
- Settled amount =
(121 ÷ 24) × 480 × 1.2 × 100 ÷ 10= 29040 CNY (used 20 days ≤30 days, 1.2× penalty coefficient). - Original spec remaining fee =
36300 - 29040= 7260 CNY. - New spec purchase fee = 0 (capacity adjusted to 0).
- Refund =
max(7260 - 0, 0)= 7260 CNY.
After zeroing, no capacity fee is incurred and the dedicated model Code is retained. Zeroing is a downgrade; the used portion is settled by the penalty coefficient (same as the scale-in formula, new capacity = 0, refund = original spec remaining fee).
Pay-as-you-go Instances
- Pay-as-you-go purchase granularity is by hour, billing precision is by minute.
- Used amount is settled by usage minutes ÷ 60, no refund is generated.
- After scale-in, the new capacity is billed at the new capacity from the effective time; deletion or zeroing stops billing.
Used Amount = Pay-as-you-go Input Unit Price × Input Capacity (kTPM) ÷ 10 × Used Minutes ÷ 60 + Pay-as-you-go Output Unit Price × Output Capacity (kTPM) × Used Minutes ÷ 60
30 Minutes Used · Scaled In to Half
Taking Qwen3.8-Max (China (Beijing)) high-speed mode as an example (assume pay-as-you-go input unit price ¥28.8/10K-TPM·hour, output unit price ¥8.64/1K-TPM·hour), original 100 input + 50 output kTPM, scaled in to 50 input + 25 output kTPM, used 30 minutes:
- Used amount =
28.8 × (100 ÷ 10) × 30 ÷ 60 + 8.64 × 50 × 30 ÷ 60=28.8 × 10 × 0.5 + 8.64 × 50 × 0.5= 144 + 216 = 360 CNY (actual consumption, non-refundable). - After scale-in, the new capacity 50 input + 25 output kTPM is billed at the new capacity from the effective time.
2 Hours Used · Scaled In to Half
Taking Qwen3.8-Max (China (Beijing)) high-speed mode as an example (assume pay-as-you-go input unit price ¥28.8/10K-TPM·hour, output unit price ¥8.64/1K-TPM·hour), original 100 input + 50 output kTPM, scaled in to 50 input + 25 output kTPM, used 2 hours:
- Used amount =
28.8 × (100 ÷ 10) × 120 ÷ 60 + 8.64 × 50 × 120 ÷ 60=28.8 × 10 × 2 + 8.64 × 50 × 2= 576 + 864 = 1440 CNY (actual consumption, non-refundable). - After scale-in, the new capacity is billed at the new capacity from the effective time.
Unsubscribe
Subscription Instances
- Subscription purchase granularity is by day, billing precision is by minute (settled amount rounded up by hour).
- Unsubscribe is a special case of scale-in (new spec = 0); the settled amount is settled by Rounded Up(used hours) × Penalty coefficient.
- ≤30 days at 1.2×, >30 days at 1.0×; unsubscribe is irreversible, and the dedicated model Code becomes invalid immediately.
Refund = max(Original Spec Remaining Fee, 0) (special case of scale-in where New Spec Purchase Fee = 0)
The unsubscribe process is handled in the Expenses Center.
Purchased 30 Days · Used 20 Days (≤30 Days, 1.2×)
Taking Qwen3.8-Max (China (Beijing)) standard mode as an example (assume daily unit price ¥121/10K-TPM·day, output unit price ¥36.29/1K-TPM·day), original 100 kTPM purchased 30 days, unsubscribe:
- Settled amount =
(121 ÷ 24) × 480 × 1.2 × 100 ÷ 10= 29040 CNY. - Original spec remaining fee =
36300 - 29040= 7260 CNY. - Refund =
max(7260, 0)= 7260 CNY.
Purchased 60 Days · Used 40 Days (>30 Days, 1.0×)
Taking Qwen3.8-Max (China (Beijing)) standard mode as an example (assume daily unit price ¥121/10K-TPM·day, output unit price ¥36.29/1K-TPM·day), original 100 kTPM purchased 60 days, unsubscribe:
- Settled amount =
(121 ÷ 24) × 960 × 1.0 × 100 ÷ 10= 48400 CNY. - Original spec remaining fee =
72600 - 48400= 24200 CNY. - Refund =
max(24200, 0)= 24200 CNY.
Pay-as-you-go Instances
- Pay-as-you-go purchase granularity is by hour, billing precision is by minute.
- Used amount is settled by usage minutes ÷ 60, no refund is generated.
- Deleting the instance stops billing.
Used Amount = Pay-as-you-go Input Unit Price × Input Capacity (kTPM) ÷ 10 × Used Minutes ÷ 60 + Pay-as-you-go Output Unit Price × Output Capacity (kTPM) × Used Minutes ÷ 60
Deleted After 30 Minutes · Stops Billing
Taking Qwen3.8-Max (China (Beijing)) high-speed mode as an example (assume pay-as-you-go input unit price ¥28.8/10K-TPM·hour, output unit price ¥8.64/1K-TPM·hour), original 100 input + 50 output kTPM, deleted after 30 minutes:
- Used amount =
28.8 × (100 ÷ 10) × 30 ÷ 60 + 8.64 × 50 × 30 ÷ 60= 360 CNY (actual consumption). - After deletion, billing stops, no refund is generated.
Add-on Capacity Package Billing
An additionally purchased capacity instance (add-on capacity package) reuses the same dedicated model Code; the total capacity increases accordingly, with no API changes required.
Billing Item | Description |
|---|---|
Billing Cycle | The selectable billing cycle of the add-on package is constrained by the base reservation's billing method: when the base is subscription, only a subscription add-on package can be selected; when the base is pay-as-you-go, either a subscription or pay-as-you-go add-on package can be selected. Subscription (by day): supported for both standard and high-speed modes, billed by natural day; billing rules are the same as a new purchase. Subscription (by 8-hour time slot): only for standard mode, rules same as 8-Hour Time Slot Reservation Billing. After adding, the capacity increases for the same dedicated model Code within that 8-hour window; outside the window, the 8h package capacity becomes invalid. Pay-as-you-go (by hour): only for high-speed mode, billed by actual duration, no duration purchase required. |
Input Capacity | Starting from 200 kTPM, step size 10 kTPM. |
Output Capacity | Starting from 20 kTPM, step size 1 kTPM. |
Duration | Only required for Subscription/by day: 1~30, 60, 90, 120, 365 days. |
Auto-Renew on Expiration | Only applies to Subscription/by day; auto-deducts and renews 1 day before expiration, with up to 4 attempts between 8:20–20:00 on that day. |
NoteA Model Code can have at most one undeleted pay-as-you-go (by hour) capacity instance: if a pay-as-you-go instance already exists, the billing cycle of the add-on capacity package can only be Subscription/by day.
Overflow Strategy and Billing
The overflow strategy determines how the excess is handled and billed after the capacity is exhausted:
Overflow Strategy | Excess Behavior | Billing Method |
|---|---|---|
Auto Overflow (Default) | Excess requests are automatically downgraded to the model standard pay-as-you-go billing (Token billing), without service interruption. | The excess is billed at the pay-as-you-go model invocation standard rate, no longer enjoying the reserved pricing. You can view the number of degradations in Overage Degradation Statistics on the details page. |
Use Reserved Capacity Only | Excess requests return a 429 error; the business needs to retry or degrade on its own. | The excess incurs no additional fees. |
In the auto-overflow scenario, when a single input exceeds the model maximum input Token, that invocation is also converted to pay-as-you-go billing.
Bills and Usage Query
- Expense bills: Subscription orders can be viewed in the Expenses Center, supporting order details and consumption records.
- Usage monitoring: On the details page of the Model Studio console, the Overview tab shows usage trends, utilization, and overage degradation statistics; the Monitor tab shows the number of invocations inside and outside the quota and the cache hit volume. For more information, see Monitoring.
- Utilization description:
Utilization = Converted Consumption ÷ Purchased Quota. The long-input tiered coefficient has been included in the converted consumption.
FAQ
Q: Why is the refund 0?
When the original spec remaining fee ≤ the new spec purchase fee, the refund is 0. Settled amount = (Daily Unit Price/24) × Rounded Up(Used Hours) × Penalty Coefficient × Original Capacity (kTPM) ÷ 10; used duration ≤30 days is settled at the 1.2× penalty coefficient; >30 days at 1.0× (no penalty). The longer the usage, the higher the settled amount, the lower the remaining fee, and the refund may be 0.
Q: What is the difference between standard mode and high-speed mode?
Standard mode has the same TPS as the standard API; high-speed mode has a 1.5~2× TPS improvement over the standard API. The billing rules of the two modes are similar, but high-speed mode additionally supports pay-as-you-go/by hour, and the billing cycle options and prices are different.
Q: How to estimate the reservation fee?
Use the formula Reservation Fee = Duration (days) × (Input Unit Price × Input kTPM ÷ 10 + Output Unit Price × Output kTPM); the unit price is taken from the corresponding mode column of the price table. For more information, see the example in Billing Rules.
Q: Can I get a refund after purchase? How much?
You can scale in or unsubscribe. Refund = max(Original Spec Remaining Fee - New Spec Purchase Fee, 0); the used portion is settled by the penalty coefficient. For more information, see Scale In.
Q: What happens upon expiration?
Upon expiration, the service stops immediately and resources are released; the dedicated model Code becomes invalid and cannot be recovered. It is recommended to enable Auto-Renew on Expiration in advance to avoid service interruption. For more information, see Expiration and Lifecycle.
Q: Can the dedicated model Code still be used after zeroing?
After zeroing, no capacity fee is incurred and the dedicated model Code is retained, but the capacity is 0 and cannot be invoked. You need to scale out again to restore invocations.