Training and deployment pricing
This topic describes the billing rules and pricing for model training and model deployment on Alibaba Cloud Model Studio.
Training billing
Text generation models – Qwen
NoteFor the training workflow, see Introduction to model fine-tuning. After training completes, deploy the new model before evaluating or calling it.
| Method | Billed by training tokens |
| Formula | Model training fee = (Total tokens in training data + Total tokens in mixed training data) × Number of epochs × Training unit price (Minimum billing unit: 1 token)
|
Qwen
Model service | Model Code | Price |
|---|---|---|
Qwen3.8-27B | qwen3.8-27b | ¥0.05/1K tokens |
Qwen3.7-Plus-2026-05-26 | qwen3.7-plus-2026-05-26 | ¥0.35/1K tokens |
Qwen3.6-27B | qwen3.6-27b | ¥0.05/1K tokens |
Qwen3.6-Flash-2026-04-16 | qwen3.6-flash-2026-04-16 | ¥0.05/1K tokens |
Qwen3.6-Plus-2026-04-02 | qwen3.6-plus-2026-04-02 | ¥0.3/1K tokens |
Qwen3.5-27B | qwen3.5-27b | ¥0.05/1K tokens |
Qwen3.5-9B | qwen3.5-9b | ¥0.02/1K tokens |
Qwen3.5-4B | qwen3.5-4b | ¥0.015/1K tokens |
Qwen3.5-Flash-2026-02-23 | qwen3.5-flash-2026-02-23 | ¥0.05/1K tokens |
Qwen3.5-Plus-2026-02-15 | qwen3.5-plus-2026-02-15 | ¥0.3/1K tokens |
Qwen3-32B | qwen3-32b | ¥0.04/1K tokens |
Qwen3-30B-A3B-Instruct-2507 | qwen3-30b-a3b-instruct-2507 | ¥0.03/1K tokens |
Qwen3-14B | qwen3-14b | ¥0.03/1K tokens |
Qwen3-8B | qwen3-8b | ¥0.006/1K tokens |
Qwen3-4B-Instruct-2507 | qwen3-4b-instruct-2507 | ¥0.006/1K tokens |
Qwen3-1.7B | qwen3-1.7b | ¥0.0045/1K tokens |
Qwen3-0.6B | qwen3-0.6b | ¥0.003/1K tokens |
Qwen2.5-72B-Instruct | qwen2.5-72b-instruct | ¥0.15/1K tokens |
Qwen2.5-32B-Instruct | qwen2.5-32b-instruct | ¥0.03/1K tokens |
Qwen2.5-14B-Instruct | qwen2.5-14b-instruct | ¥0.03/1K tokens |
Qwen2.5-7B-Instruct | qwen2.5-7b-instruct | ¥0.006/1K tokens |
Qwen-Plus-Character-2025-11-06 | qwen-plus-character-2025-11-06 | ¥0.15/1K tokens |
Qwen-VL
Model service | Model Code | Price |
|---|---|---|
Qwen3-VL-8B-Instruct | qwen3-vl-8b-instruct | ¥0.012/1K tokens |
Qwen3-VL-8B-Thinking | qwen3-vl-8b-thinking | ¥0.012/1K tokens |
Qwen3-VL-4B-Instruct | qwen3-vl-4b-instruct | ¥0.006/1K tokens |
Qwen2.5-VL-72B-Instruct | qwen2.5-vl-72b-instruct | ¥0.05/1K tokens |
Qwen2.5-VL-32B-Instruct | qwen2.5-vl-32b-instruct | ¥0.02/1K tokens |
Qwen2.5-VL-7B-Instruct | qwen2.5-vl-7b-instruct | ¥0.01/1K tokens |
Calculate tokens for images and videos
Images
Formula: Image Tokens = h_bar * w_bar / token_pixels + 2
-
h_bar, w_bar: The height and width of the scaled image. Before processing an image, the model performs pre-processing to scale it down to a specific pixel limit. This limit depends on the values of themax_pixelsandvl_high_resolution_imagesparameters. For more information, see Process high-resolution images. -
token_pixels: The pixel value corresponding to each visualtoken. This varies by model:qwen3.8-series,qwen3.7-series,qwen3.6-series,qwen3.5-series,Qwen3-VL,qwen-vl-max, andqwen-vl-plus:Eachtokencorresponds to32x32pixels.QVQand otherQwen2.5-VLmodels:Each token corresponds to28x28pixels.
The following code demonstrates the approximate image scaling logic used by the model. Use it to estimate the tokens for an image. For actual billing, refer to the API response.
import math
from PIL import Image # pip install Pillow
def smart_resize(image_path, max_pixels, vl_high_resolution_images):
"""Calculates the scaled dimensions of an image based on model parameters to estimate image tokens."""
image = Image.open(image_path)
height, width = image.height, image.width
# The scaling factor is 32 for models such as Qwen3.6, Qwen3.5, and Qwen3-VL. For other models, it is 28.
factor = 32
h_bar = round(height / factor) * factor
w_bar = round(width / factor) * factor
# Token lower limit: 4 tokens
min_pixels = 4 * factor * factor
# If vl_high_resolution_images=True, the token upper limit is fixed at 16384, and max_pixels is ignored.
if vl_high_resolution_images:
max_pixels = 16384 * factor * factor
# Constrains the total number of pixels to the range [min_pixels, max_pixels].
if h_bar * w_bar > max_pixels:
beta = math.sqrt((height * width) / max_pixels)
h_bar = math.floor(height / beta / factor) * factor
w_bar = math.floor(width / beta / factor) * factor
elif h_bar * w_bar < min_pixels:
beta = math.sqrt(min_pixels / (height * width))
h_bar = math.ceil(height * beta / factor) * factor
w_bar = math.ceil(width * beta / factor) * factor
return h_bar, w_bar
if __name__ == "__main__":
# Note: The values of max_pixels and vl_high_resolution_images must match the parameters passed when calling the model.
h_bar, w_bar = smart_resize("xxx/test.jpg", max_pixels=2560 * 32 * 32, vl_high_resolution_images=False)
print(f"Scaled image dimensions: Height {h_bar}, Width {w_bar}")
# Each image includes one <vision_bos> and one <vision_eos> token.
token = int(h_bar * w_bar / (32 * 32)) + 2
print(f"Number of image tokens: {token}")
Videos
-
Video files:
When processing a video file, the model first extracts frames and then calculates the total number of tokens for all video frames. Because this calculation is complex, you can use the following code to estimate the total token consumption for a video by providing its path:
# Before use, install: pip install opencv-python
import math
import os
import logging
import cv2
logger = logging.getLogger(__name__)
FRAME_FACTOR = 2
# For models such as Qwen3.6, Qwen3.5, Qwen3-VL, qwen-vl-max-0813, qwen-vl-plus-0815, and qwen-vl-plus-0710, the image scaling factor is 32.
IMAGE_FACTOR = 32
# For other models, the image scaling factor is 28.
# IMAGE_FACTOR = 28
# Maximum aspect ratio for video frames
MAX_RATIO = 200
# Pixel lower limit for video frames
VIDEO_MIN_PIXELS = 4 * 32 * 32
# Pixel upper limit for video frames. For the Qwen3-VL-Plus model, VIDEO_MAX_PIXELS is 640 * 32 * 32. For other models, it is 768 * 32 * 32.
VIDEO_MAX_PIXELS = 640 * 32 * 32
# If the user does not pass the FPS parameter, the default value is used for fps.
FPS = 2.0
# Minimum number of extracted frames
FPS_MIN_FRAMES = 4
# Maximum number of extracted frames (set based on the selected model)
FPS_MAX_FRAMES = 2000
# Maximum pixel value for video input. For the Qwen3-VL-Plus model, set VIDEO_TOTAL_PIXELS to 131072 * 32 * 32. For other models, set it to 65536 * 32 * 32.
VIDEO_TOTAL_PIXELS = int(float(os.environ.get('VIDEO_TOTAL_PIXELS', 131072 * 32 * 32)))
def round_by_factor(number: int, factor: int) -> int:
"""Returns the integer closest to 'number' that is divisible by 'factor'."""
return round(number / factor) * factor
def ceil_by_factor(number: int, factor: int) -> int:
"""Returns the smallest integer that is greater than or equal to 'number' and divisible by 'factor'."""
return math.ceil(number / factor) * factor
def floor_by_factor(number: int, factor: int) -> int:
"""Returns the largest integer that is less than or equal to 'number' and divisible by 'factor'."""
return math.floor(number / factor) * factor
def extract_vision_info(conversations):
vision_infos = []
if isinstance(conversations[0], dict):
conversations = [conversations]
for conversation in conversations:
for message in conversation:
if isinstance(message["content"], list):
for ele in message["content"]:
if (
"image" in ele
or "image_url" in ele
or "video" in ele
or ele.get("type","") in ("image", "image_url", "video")
):
vision_infos.append(ele)
return vision_infos
def smart_nframes(ele,total_frames,video_fps):
"""Calculates the number of extracted video frames.
Args:
ele (dict): A dictionary containing the video configuration.
- fps: Controls the number of input frames extracted for the model.
total_frames (int): The original total number of frames in the video.
video_fps (int | float): The original frame rate of the video.
Raises:
An error is reported if nframes is not within the interval [FRAME_FACTOR, total_frames].
Returns:
The number of video frames for model input.
"""
assert not ("fps" in ele and "nframes" in ele), "Only accept either `fps` or `nframes`"
fps = ele.get("fps", FPS)
min_frames = ceil_by_factor(ele.get("min_frames", FPS_MIN_FRAMES), FRAME_FACTOR)
max_frames = floor_by_factor(ele.get("max_frames", min(FPS_MAX_FRAMES, total_frames)), FRAME_FACTOR)
duration = total_frames / video_fps if video_fps != 0 else 0
if duration-int(duration)>(1/fps):
total_frames = math.ceil(duration * video_fps)
else:
total_frames = math.ceil(int(duration)*video_fps)
nframes = total_frames / video_fps * fps
if nframes > total_frames:
logger.warning(f"smart_nframes: nframes[{nframes}] > total_frames[{total_frames}]")
nframes = int(min(min(max(nframes, min_frames), max_frames), total_frames))
if not (FRAME_FACTOR <= nframes and nframes <= total_frames):
raise ValueError(f"nframes should in interval [{FRAME_FACTOR}, {total_frames}], but got {nframes}.")
return nframes
def get_video(video_path):
# Get video information
cap = cv2.VideoCapture(video_path)
frame_width = int(cap.get(cv2.CAP_PROP_FRAME_WIDTH))
# Get video height
frame_height = int(cap.get(cv2.CAP_PROP_FRAME_HEIGHT))
total_frames = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))
video_fps = cap.get(cv2.CAP_PROP_FPS)
return frame_height, frame_width, total_frames, video_fps
def smart_resize(ele, path, factor=IMAGE_FACTOR):
# Get the original width and height of the video
height, width, total_frames, video_fps = get_video(path)
# Token lower limit for video frames
min_pixels = VIDEO_MIN_PIXELS
total_pixels = VIDEO_TOTAL_PIXELS
# Number of extracted video frames
nframes = smart_nframes(ele, total_frames, video_fps)
max_pixels = max(min(VIDEO_MAX_PIXELS, total_pixels / nframes * FRAME_FACTOR),int(min_pixels * 1.05))
# The aspect ratio of the video should not exceed 200:1 or 1:200.
if max(height, width) / min(height, width) > MAX_RATIO:
raise ValueError(
f"absolute aspect ratio must be smaller than {MAX_RATIO}, got {max(height, width) / min(height, width)}"
)
h_bar = max(factor, round_by_factor(height, factor))
w_bar = max(factor, round_by_factor(width, factor))
if h_bar * w_bar > max_pixels:
beta = math.sqrt((height * width) / max_pixels)
h_bar = floor_by_factor(height / beta, factor)
w_bar = floor_by_factor(width / beta, factor)
elif h_bar * w_bar < min_pixels:
beta = math.sqrt(min_pixels / (height * width))
h_bar = ceil_by_factor(height * beta, factor)
w_bar = ceil_by_factor(width * beta, factor)
return h_bar, w_bar
def token_calculate(video_path, fps):
# Pass the video path and the fps frame extraction parameter.
messages = [{"content": [{"video": video_path, "fps": fps}]}]
vision_infos = extract_vision_info(messages)[0]
resized_height, resized_width = smart_resize(vision_infos, video_path)
height, width, total_frames, video_fps = get_video(video_path)
num_frames = smart_nframes(vision_infos, total_frames, video_fps)
print(f"Original video dimensions: {height}*{width}, Model input dimensions: {resized_height}*{resized_width}, Total video frames: {total_frames}, Total frames extracted when fps is {fps}: {num_frames}", end=", ")
video_token = int(math.ceil(num_frames / 2) * resized_height / 32 * resized_width / 32)
video_token += 2 # The system automatically adds <|vision_bos|> and <|vision_eos|> visual markers (1 token each).
return video_token
video_token = token_calculate("xxx/test.mp4", 1)
print("Video tokens:", video_token)
-
Image list:
When a video is passed as a list of images, it means that frame extraction has already been performed. Use the following code to calculate the token consumption by providing the path and number of images:
# Before use, install: pip install Pillow
import math
import os
import logging
from typing import Tuple
from PIL import Image
logger = logging.getLogger(__name__)
# ==================== Constant Definitions ====================
FRAME_FACTOR = 2
# For models such as Qwen3-VL, qwen-vl-max-0813, qwen-vl-plus-0815, and qwen-vl-plus-0710, the scaling factor is 32.
IMAGE_FACTOR = 32
# For other models, the scaling factor is 28.
# IMAGE_FACTOR = 28
# Constants for token calculation
TOKEN_DIVISOR = 32 # Divisor for token calculation
VISION_SPECIAL_TOKENS = 2 # <|vision_bos|> and <|vision_eos|> markers
# Maximum aspect ratio for video frames
MAX_RATIO = 200
# Pixel lower limit for video frames
VIDEO_MIN_PIXELS = 4 * 32 * 32
# Pixel upper limit for video frames. For the Qwen3-VL-Plus model, VIDEO_MAX_PIXELS is 640 * 32 * 32. For other models, it is 768 * 32 * 32.
VIDEO_MAX_PIXELS = 640 * 32 * 32
# Maximum pixel value for video input. For the Qwen3-VL-Plus model, set VIDEO_TOTAL_PIXELS to 131072 * 32 * 32. For other models, set it to 65536 * 32 * 32.
VIDEO_TOTAL_PIXELS = int(float(os.environ.get('VIDEO_TOTAL_PIXELS', 131072 * 32 * 32)))
def round_by_factor(number: int, factor: int) -> int:
"""Returns the integer closest to 'number' that is divisible by 'factor'."""
return round(number / factor) * factor
def ceil_by_factor(number: int, factor: int) -> int:
"""Returns the smallest integer that is greater than or equal to 'number' and divisible by 'factor'."""
return math.ceil(number / factor) * factor
def floor_by_factor(number: int, factor: int) -> int:
"""Returns the largest integer that is less than or equal to 'number' and divisible by 'factor'."""
return math.floor(number / factor) * factor
def get_image_size(image_path: str) -> Tuple[int, int]:
if not os.path.exists(image_path):
raise FileNotFoundError(f"Image file not found: {image_path}")
try:
image = Image.open(image_path)
height = image.height
width = image.width
image.close() # Close the file promptly
return height, width
except Exception as e:
raise ValueError(f"Cannot read image file {image_path}: {str(e)}")
def smart_resize(height: int, width: int, nframes: int, factor: int = IMAGE_FACTOR) -> Tuple[int, int]:
"""
Calculates the scaled dimensions of an image
Args:
height: Original image height
width: Original image width
nframes: Number of video frames
factor: Scaling factor, defaults to IMAGE_FACTOR
Returns:
(resized_height, resized_width) The scaled height and width
Raises:
ValueError: Aspect ratio exceeds the limit
"""
# Token lower limit for video frames
min_pixels = VIDEO_MIN_PIXELS
total_pixels = VIDEO_TOTAL_PIXELS
# Number of extracted video frames
max_pixels = max(min(VIDEO_MAX_PIXELS, total_pixels / nframes * FRAME_FACTOR), int(min_pixels * 1.05))
# The aspect ratio of the video should not exceed 200:1 or 1:200.
aspect_ratio = max(height, width) / min(height, width)
if aspect_ratio > MAX_RATIO:
raise ValueError(
f"Image aspect ratio must be less than {MAX_RATIO}:1, but is currently {aspect_ratio:.2f}:1"
)
h_bar = max(factor, round_by_factor(height, factor))
w_bar = max(factor, round_by_factor(width, factor))
if h_bar * w_bar > max_pixels:
beta = math.sqrt((height * width) / max_pixels)
h_bar = floor_by_factor(height / beta, factor)
w_bar = floor_by_factor(width / beta, factor)
elif h_bar * w_bar < min_pixels:
beta = math.sqrt(min_pixels / (height * width))
h_bar = ceil_by_factor(height * beta, factor)
w_bar = ceil_by_factor(width * beta, factor)
return h_bar, w_bar
def calculate_video_tokens(image_path: str, nframes: int = 1, factor: int = IMAGE_FACTOR, verbose: bool = True) -> int:
"""
Args:
image_path: Path to the video frame file
nframes: Number of video frames,
factor: Scaling factor, defaults to IMAGE_FACTOR
verbose: Whether to print detailed information
Returns:
The number of tokens consumed
Raises:
FileNotFoundError: The file does not exist
ValueError: The file format is invalid or the aspect ratio exceeds the limit
"""
# Get the original image dimensions (read only once)
height, width = get_image_size(image_path)
# Calculate the scaled dimensions
resized_height, resized_width = smart_resize(height, width, nframes, factor)
# Calculate the number of tokens
# Formula: ceil(nframes/2) * (height/TOKEN_DIVISOR) * (width/TOKEN_DIVISOR) + VISION_SPECIAL_TOKENS
video_token = int(
math.ceil(nframes / 2) *
(resized_height / TOKEN_DIVISOR) *
(resized_width / TOKEN_DIVISOR)
)
# Add visual marker tokens (<|vision_bos|> and <|vision_eos|>)
video_token += VISION_SPECIAL_TOKENS
if verbose:
print(f"Original video frame dimensions: {height}x{width}, Model input dimensions: {resized_height}x{resized_width}, ", end="")
return video_token
if __name__ == "__main__":
try:
video_token = calculate_video_tokens("xxx/test.jpg", nframes=30)
print(f"Video tokens: {video_token}\n")
except Exception as e:
print(f"Error: {str(e)}\n")
Image generation models – Wan
NoteFor the training workflow, see Fine-tune image generation models. After training completes, deploy the new model before calling it.
Method | Billed by training tokens |
Formula | Model training fee = Total training tokens × Training unit price (Billing unit: per 1,000 tokens) |
Model | Code | Training price (per 1K tokens) |
|---|---|---|
Wan image generation | wan2.7-image-pro | CNY 0.08 |
Wan image generation | wan2.7-image | CNY 0.08 |
Billing example
Suppose you fine-tune the wan2.7-image-pro model for t2i. The parameters are: max_steps = 200, max_token_length = "1k", and the training price is CNY 0.08 per 1,000 tokens:
- From the table: Lmax = 12,800 (generation_type=t2i, max_token_length=1k), Lstep ≈ Lmax = 12,800
- Total training tokens ≈ 200 × 12800 = 2560000 = 2560 thousand tokens
- Model training fee ≈ 2560 × 0.08 = CNY 204.8
Image generation models – Qwen
NoteFor the training workflow, see Fine-tune image generation models. After training completes, deploy the new model before calling it.
Method | Billed by training tokens |
Formula | Model training fee = Total training tokens × Training unit price (Billing unit: per 1,000 tokens) |
Model | Code | Training price (per 1K tokens) |
|---|---|---|
Qwen image generation | qwen-image-2.0 | CNY 0.02 |
Qwen image generation | qwen-image-2.0-pro | CNY 0.02 |
Billing example
Suppose you fine-tune the qwen-image-2.0 model with a training set of 1 image. The parameters are: n_epochs = 1000, max_pixels = "1k", the training price is CNY 0.02 per 1,000 tokens, and the GPU coefficient is 8:
- Tokens per image = max_pixels / Compression ratio = 1024×1024 / 1024 = 1024
- Total training tokens = 1000 × 1 × 1024 × 8 = 8192000 = 8192 thousand tokens
- Model training fee = 8192 × 0.02 = CNY 163.84
The following table estimates token consumption and fees per training image, assuming a GPU coefficient of 8. For a training set of N images, multiply the values by N.
max_pixels | n_epochs | Estimated tokens | Estimated fee (CNY) |
|---|---|---|---|
1k | 800 | 6,553,600 | 131.07 |
1k | 1,000 | 8,192,000 | 163.84 |
1k | 2,000 | 16,384,000 | 327.68 |
2k | 800 | 26,214,400 | 524.29 |
2k | 1,000 | 32,768,000 | 655.36 |
2k | 2,000 | 65,536,000 | 1310.72 |
Note
- The GPU coefficient is dynamically adjusted based on job scheduling. The value 8 in the example and the table above is used only to illustrate the calculation. For actual token consumption, see the
usagefield returned by the Query a fine-tuning job operation. Your bill is the final authority on fees. batch_sizedoes not affect billing or the total training tokens.
Video generation models – Wan
NoteFor the training workflow, see Fine-tuning video generation models. After training completes, deploy the new model before calling it.
Method | Billed by training tokens |
Formula | Model training fee = Total training tokens × Training unit price (Billing unit: per 1,000 tokens) |
Model | Code | Training price (per 1K tokens) |
|---|---|---|
Wan image-to-video (first frame-based) | wan2.7-i2v | CNY 2 |
wan2.2-i2v-flash | CNY 0.06 | |
wan2.5-i2v-preview | CNY 0.32 | |
Image-to-video (first and last frame-based) | wan2.7-i2v | CNY 2 |
wan2.2-kf2v-flash | CNY 0.06 |
Speech synthesis models – CosyVoice
NoteCosyVoice model fine-tuning is available only in the China (Beijing) region. For the training workflow, see Fine-tune CosyVoice. After training completes, deploy the new model before calling it.
Method | Billed by tokens consumed during training |
Unit price | CNY 0.2 per 1,000 tokens |
Formula | Model training fee = Total consumed tokens × Training unit price |
Deployment billing
Text generation models: Qwen
Billing by usage duration (Provisioned Throughput)
Fee = Usage duration × (Input TPM unit price × Input TPM + Output TPM unit price × Output TPM)
Post-paid is calculated hourly: the usage duration unit is hours, and the unit price is taken from the "Continuous 1 hour" column in the table below; prepaid is calculated daily: the usage duration unit is days, and the unit price is taken from the "Continuous 1 day" column in the table below.
- Prepaid orders take effect in real time after payment, with a validity period of N days ending at 23:59 on day N. If the order is placed after 22:00, the expiration date will be automatically extended by 1 day.
- After a prepaid order expires, the service will be stopped with a 2-hour delay, and resources will be retained for 14 hours after the stop and then released.
- Prepaid orders support early termination; the used portion is settled at a 1.2x coefficient for refund. For details, see Refund rules for configuration downgrades.
- For post-paid billing, if the account is in arrears, the deployed resources will continue to be retained and billed for 24 hours, during which the service can still be used normally. After 24 hours, the system stops billing, the model deployment enters an arrears state, and the underlying resources will be deleted, but the model deployment task will be retained. After the arrears are paid, the system will reallocate resources and restore usage (fees will continue to accrue after restoration). If you do not want to continue incurring fees, you can delete the model deployment task, and billing will stop after successful deletion.
When the model input exceeds the maximum input Token, the relevant call will automatically switch to the pay-as-you-go mode of the current model; when the purchased TPM is exceeded, it is handled according to the overflow strategy selected at creation ("auto-overflow" switches to pay-as-you-go, "use-only-PTU-capacity" returns 429). At this time, inference performance may degrade and will be subject to the public traffic control of the current snapshot model in the business space, and fees will be charged according to the model invocation (pay-as-you-go) standard.
- In this case (only under the "auto-overflow" strategy), the API response Header will include:
x-dashscope-ptu-overflow:true. - For TPM statistics, go to: Model Monitoring (Beijing).
For the specific fee reduction and refund rules in scale-down (downgrade) scenarios, please refer to: Refund rules for configuration downgrades.
NotePTU deployment supports long-input tiered capacity coefficients and cache discounts; see Provisioned Throughput long input and caching for details.
North China 2 (Beijing)
Qwen
Model Name | Model Code | Max input tokens | Post-paid input Per 10K TPM/hour | Post-paid output Per 1K TPM/hour | Prepaid input Per 10K TPM/day | Prepaid output Per 1K TPM/day |
|---|---|---|---|---|---|---|
Qwen3.8-Max | qwen3.8-max | 1M | ¥28.8 | ¥8.64 | ¥345.6 | ¥103.68 |
Qwen3.7-Flash-2026-07-15 | qwen3.7-flash-2026-07-15 | 128K | ¥0.48 | ¥0.19 | ¥5.76 | ¥2.3 |
Qwen3.7-Max-2026-05-20 | qwen3.7-max-2026-05-20 | 256K | ¥28.8 | ¥8.64 | ¥345.6 | ¥103.68 |
Qwen3.7-Plus-2026-05-26 | qwen3.7-plus-2026-05-26 | 256K | ¥4.8 | ¥1.92 | ¥57.6 | ¥23.04 |
Qwen3.6-Plus-2026-04-02 | qwen3.6-plus-2026-04-02 | 128K | ¥4.8 | ¥2.88 | ¥57.6 | ¥34.56 |
Qwen3.5-Plus-2026-04-20 | qwen3.5-plus-2026-04-20 | 128K | ¥1.92 | ¥1.15 | ¥23.04 | ¥13.82 |
Qwen3-Max-2025-09-23 | qwen3-max-2025-09-23 | 128K | ¥7.68 | ¥3.08 | ¥92.16 | ¥36.96 |
Qwen-Flash-2025-07-28 | qwen-flash-2025-07-28 | 128K | ¥0.36 | ¥0.36 | ¥4.32 | ¥4.32 |
Qwen-Plus-2025-12-01 | qwen-plus-2025-12-01 | 128K | ¥1.92 | Non-thinking: ¥0.48 Thinking: ¥1.92 | ¥23.04 | Non-thinking: ¥5.76 Thinking: ¥23.04 |
DeepSeek
Model Name | Model Code | Max input tokens | Post-paid input Per 10K TPM/hour | Post-paid output Per 1K TPM/hour | Prepaid input Per 10K TPM/day | Prepaid output Per 1K TPM/day |
|---|---|---|---|---|---|---|
DeepSeek-v4-Flash | deepseek-v4-flash | 256K | ¥3.6 | ¥0.72 | ¥43.2 | ¥8.64 |
DeepSeek-v4-Flash-0731 | deepseek-v4-flash-0731 | 256K | ¥7.2 | ¥1.44 | ¥86.4 | ¥17.28 |
DeepSeek-v4-Pro | deepseek-v4-pro | 256K | ¥43.2 | ¥8.64 | ¥518.4 | ¥103.68 |
DeepSeek-v3 | deepseek-v3 | 64K | ¥7.2 | ¥2.88 | ¥86.4 | ¥34.56 |
Qwen-VL
Model Name | Model Code | Max input tokens | Post-paid input Per 10K TPM/hour | Post-paid output Per 1K TPM/hour | Prepaid input Per 10K TPM/day | Prepaid output Per 1K TPM/day |
|---|---|---|---|---|---|---|
Qwen3-VL-Plus-2025-09-23 | qwen3-vl-plus-2025-09-23 | 128K | ¥2.4 | ¥2.4 | ¥28.8 | ¥28.8 |
GLM
Model Name | Model Code | Max input tokens | Post-paid input Per 10K TPM/hour | Post-paid output Per 1K TPM/hour | Prepaid input Per 10K TPM/day | Prepaid output Per 1K TPM/day |
|---|---|---|---|---|---|---|
GLM-5.2 | glm-5.2 | 1M | ¥28.8 | ¥10.08 | ¥345.6 | ¥120.96 |
Singapore
Qwen
Model Name | Model Code | Max input tokens | Post-paid input Per 10K TPM/hour | Post-paid output Per 1K TPM/hour | Prepaid input Per 10K TPM/day | Prepaid output Per 1K TPM/day |
|---|---|---|---|---|---|---|
Qwen3.8-Max | qwen3.8-max | 1M | ¥35.97 | ¥10.79 | ¥431.7 | ¥129.5 |
Qwen3.7-Flash-2026-07-15 | qwen3.7-flash-2026-07-15 | 128K | ¥0.54 | ¥0.23 | ¥6.47 | ¥2.81 |
Qwen3.7-Max-2026-05-20 | qwen3.7-max-2026-05-20 | 256K | ¥44.97 | ¥13.49 | ¥539.6 | ¥161.87 |
Qwen3.7-Plus-2026-05-26 | qwen3.7-plus-2026-05-26 | 256K | ¥7.19 | ¥2.88 | ¥86.3 | ¥34.53 |
Qwen3.6-Plus-2026-04-02 | qwen3.6-plus-2026-04-02 | 128K | ¥9 | ¥5.4 | ¥107.9 | ¥64.75 |
Qwen3.5-Plus-2026-04-20 | qwen3.5-plus-2026-04-20 | 128K | ¥7.2 | ¥4.32 | ¥86.3 | ¥51.8 |
DeepSeek
Model Name | Model Code | Max input tokens | Post-paid input Per 10K TPM/hour | Post-paid output Per 1K TPM/hour | Prepaid input Per 10K TPM/day | Prepaid output Per 1K TPM/day |
|---|---|---|---|---|---|---|
DeepSeek-v4-Flash | deepseek-v4-flash | 256K | ¥5.4 | ¥1.08 | ¥64.8 | ¥12.95 |
DeepSeek-v4-Flash-0731 | deepseek-v4-flash-0731 | 256K | ¥10.79 | ¥2.16 | ¥129.5 | ¥25.9 |
DeepSeek-v4-Pro | deepseek-v4-pro | 256K | ¥64.75 | ¥12.95 | ¥777 | ¥155.4 |
Qwen-VL
Model Name | Model Code | Max input tokens | Post-paid input Per 10K TPM/hour | Post-paid output Per 1K TPM/hour | Prepaid input Per 10K TPM/day | Prepaid output Per 1K TPM/day |
|---|---|---|---|---|---|---|
Qwen3-VL-Plus-2025-09-23 | qwen3-vl-plus-2025-09-23 | 128K | ¥3.6 | ¥2.88 | ¥43.2 | ¥34.53 |
GLM
Model Name | Model Code | Max input tokens | Post-paid input Per 10K TPM/hour | Post-paid output Per 1K TPM/hour | Prepaid input Per 10K TPM/day | Prepaid output Per 1K TPM/day |
|---|---|---|---|---|---|---|
GLM-5.2 | glm-5.2 | 1M | ¥37.8 | ¥11.87 | ¥453.3 | ¥142.45 |
Billing by usage duration (Dedicated Throughput)
- Prepaid orders support early termination; the used portion is settled at a 1.2x coefficient for refund. For details, see Refund rules for configuration downgrades.
North China 2 (Beijing)
Model | Type | Base input (TPM) | Base output (TPM) | Starting price (CNY/month) | Scale-out price (CNY/month) |
|---|---|---|---|---|---|
qwen3.7-plus-2026-05-26 | Fast | 1,372,000 | 170,000 | 556,466 | 556,466 |
qwen3.6-27b | Standard | 273,000 | 34,000 | 49,118 | 49,118 |
qwen3.5-397b-a17b | Standard | 896,000 | 112,000 | 1,107,904 | 1,107,904 |
qwen3.5-122b-a10b | Fast | 7,288,000 | 904,000 | 2,226,984 | 2,226,984 |
qwen3.5-35b-a3b | Standard | 656,000 | 82,000 | 109,962 | 109,962 |
qwen3.5-27b | Standard | 2,432,000 | 304,000 | 1,108,688 | 1,108,688 |
Efficient | 291,000 | 36,000 | 48,996 | 48,996 | |
qwen3.5-4b | Standard | 1,072,000 | 134,000 | 109,478 | 109,478 |
glm-5.2 | Fast | 1,004,000 | 120,000 | 1,895,260 | 1,895,260 |
glm-5.1 | Standard | 256,000 | 32,000 | 504,704 | 504,704 |
deepseek-v4-flash | Standard | 2,240,000 | 280,000 | 1,107,400 | 1,107,400 |
deepseek-v4-flash-0731 | Fast | 1,676,000 | 213,000 | 277,984 | 277,984 |
Model | Type | Input length | Output length | Cache Hit Rate | First token latency (ms) | Per-token latency (ms) |
|---|---|---|---|---|---|---|
qwen3.7-plus-2026-05-26 | Fast | 16,000 | 2,000 | 0 | 2,418 | 15 |
qwen3.6-27b | Standard | 4,000 | 500 | 0 | 1,292 | 19 |
qwen3.5-397b-a17b | Standard | 4,000 | 500 | 0 | 996 | 27 |
qwen3.5-122b-a10b | Fast | 16,000 | 2,000 | 0 | 568 | 8 |
qwen3.5-35b-a3b | Standard | 4,000 | 500 | 0 | 471 | 10 |
qwen3.5-27b | Standard | 4,000 | 500 | 0 | 703 | 14 |
Efficient | 4,000 | 500 | 0 | 1,448 | 23 | |
qwen3.5-4b | Standard | 4,000 | 500 | 0 | 552 | 6 |
glm-5.2 | Fast | 16,000 | 2,000 | 0 | 1,558 | 15 |
glm-5.1 | Standard | 4,000 | 500 | 0 | 769 | 27 |
deepseek-v4-flash | Standard | 4,000 | 500 | 0 | 651 | 19 |
deepseek-v4-flash-0731 | Fast | 16,000 | 2,000 | 0 | 1,079 | 13 |
By model token usage
Fee = Model input token count × Model input unit price + Model output token count × Model output unit price (minimum billing unit: 1 token)
- Billing by model token usage is supported only after you complete efficient SFT training (that is, LoRA efficient fine-tuning; the plan parameter is set to lora for API deployment) on the following base models and obtain a custom model.
Beijing
Base model | Model Code | Input CNY/million tokens | Output CNY/million tokens |
|---|---|---|---|
Qwen3.6-27B | qwen3.6-27b | <256K ¥3 | <256K ¥18 |
Qwen3.5-27B | qwen3.5-27b | <128K ¥0.6 128K-256K ¥1.8 | <128K ¥4.8 128K-256K ¥14.4 |
Qwen3-32B | qwen3-32b | Non-thinking mode: ¥2 Thinking mode: ¥2 | Non-thinking mode: ¥8 Thinking mode: ¥20 |
Qwen3-14B | qwen3-14b | Non-thinking mode: ¥1 Thinking mode: ¥1 | Non-thinking mode: ¥4 Thinking mode: ¥10 |
Qwen3-8B | qwen3-8b | Non-thinking mode: ¥0.5 Thinking mode: ¥0.5 | Non-thinking mode: ¥2 Thinking mode: ¥5 |
Qwen3-VL-8B-Instruct | qwen3-vl-8b-instruct | ¥0.5 | ¥2 |
Qwen3-4B-Instruct-2507 | qwen3-4b-instruct-2507 | Non-thinking mode: ¥0.3 Thinking mode: ¥0.3 | Non-thinking mode: ¥1.2 Thinking mode: ¥3 |
Qwen2.5-Open-Source-72B | qwen2.5-72b-instruct | ¥4 | ¥12 |
Qwen2.5-VL-72B | qwen2.5-vl-72b-instruct | ¥16 | ¥48 |
Qwen2.5-Open-Source-32B | qwen2.5-32b-instruct | ¥2 | ¥6 |
Qwen2.5-VL-32B | qwen2.5-vl-32b-instruct | ¥8 | ¥24 |
Qwen2.5-Open-Source-14B | qwen2.5-14b-instruct | ¥1 | ¥3 |
Qwen2.5-Open-Source-7B | qwen2.5-7b-instruct | ¥0.5 | ¥1 |
Qwen2.5-VL-7B | qwen2.5-vl-7b-instruct | ¥2 | ¥5 |
Singapore
Base model | Model Code | Input CNY/million tokens | Output CNY/million tokens |
|---|---|---|---|
Qwen3.6-27B | qwen3.6-27b | <256K ¥4.497 | <256K ¥26.979 |
Qwen3.5-27B | qwen3.5-27b | ¥2.202 | ¥17.614 |
Qwen3-32B | qwen3-32b | Non-thinking mode: ¥1.174 Thinking mode: ¥1.174 | Non-thinking mode: ¥4.697 Thinking mode: ¥4.697 |
Qwen3-14B | qwen3-14b | Non-thinking mode: ¥2.569 Thinking mode: ¥2.569 | Non-thinking mode: ¥10.275 Thinking mode: ¥30.825 |
Qwen3-8B | qwen3-8b | Non-thinking mode: ¥1.321 Thinking mode: ¥1.321 | Non-thinking mode: ¥5.137 Thinking mode: ¥15.412 |
Qwen3-VL-8B-Instruct | qwen3-vl-8b-instruct | ¥1.321 | ¥5.137 |
Qwen3-4B-Instruct-2507 | qwen3-4b-instruct-2507 | Non-thinking mode: ¥0.807 Thinking mode: ¥0.807 | Non-thinking mode: ¥3.082 Thinking mode: ¥9.247 |
Qwen2.5-Open-Source-72B | qwen2.5-72b-instruct | ¥10.275 | ¥41.1 |
Qwen2.5-VL-72B | qwen2.5-vl-72b-instruct | ¥20.55 | ¥61.65 |
Qwen2.5-Open-Source-32B | qwen2.5-32b-instruct | ¥5.137 | ¥20.55 |
Qwen2.5-VL-32B | qwen2.5-vl-32b-instruct | ¥10.275 | ¥30.825 |
Qwen2.5-Open-Source-14B | qwen2.5-14b-instruct | ¥2.569 | ¥10.275 |
Qwen2.5-Open-Source-7B | qwen2.5-7b-instruct | ¥1.284 | ¥5.137 |
Qwen2.5-VL-7B | qwen2.5-vl-7b-instruct | ¥2.569 | ¥7.706 |
Billing by usage duration (Model Unit)
Cost = Usage duration (hours) × Number of Model Units × Model Unit unit price
The "Model Unit unit price" takes the "Hourly unit price" column in the table below for postpaid scenarios; for monthly subscription prepaid billing, the formula becomes Number of monthly subscriptions × Number of Model Units × Monthly unit price.
- For the first month of a prepaid purchase, if you cancel the subscription early within the first month, the daily unit price (≈ Monthly unit price / 30) will be billed at 1.2 times the rate (less than one day is billed as one day)
NoteThe compute resources for the Model Unit postpaid method are first-come, first-served. If the purchase fails, a full refund will be issued.
Text generation
Qwen
Model Name | Model Code | Model Unit Specification | Hourly unit price (CNY) Minimum billing: minutes | Monthly subscription unit price (CNY) Minimum billing: days |
|---|---|---|---|---|
Qwen3.6-35B-A3B | qwen3.6-35b-a3b | MU1 x 8 | ¥432 | ¥208,944 |
MU2 x 8 | ¥504 | ¥240,288 | ||
MU3 x 8 | ¥1,096 | ¥527,752 | ||
MU8 x 1 | ¥47 | ¥22,400 | ||
MU9 x 1 | ¥51 | ¥24,600 | ||
Qwen3.6-27B | qwen3.6-27b | MU9 x 1 | ¥51 | ¥24,600 |
Qwen3.6-Flash-2026-04-16 | qwen3.6-flash-2026-04-16 | MU1 x 2 | ¥108 | ¥52,236 |
MU3 x 8 | ¥1,096 | ¥527,752 | ||
Qwen3.6-Plus-2026-04-02 | qwen3.6-plus-2026-04-02 | MU1 x 8 MU1 x 16 (PD disaggregation mode) | ¥432 PD disaggregation mode: ¥864 | ¥208,944 PD disaggregation mode: ¥417,888 |
MU2 x 8 | ¥504 | ¥240,288 | ||
Qwen3.5-397B-A17B | qwen3.5-397b-a17b | MU3 x 8 MU3 x 16 (PD disaggregation mode) | ¥1,096 PD disaggregation mode: ¥2,192 | ¥527,752 PD disaggregation mode: ¥1,055,504 |
MU6 x 16 | ¥400 | ¥193,424 | ||
Qwen3.5-122B-A10B | qwen3.5-122b-a10b | MU1 x 4 | ¥216 | ¥104,472 |
MU6 x 16 | ¥400 | ¥193,424 | ||
Qwen3.5-35B-A3B | qwen3.5-35b-a3b | MU1 x 2 | ¥108 | ¥52,236 |
MU2 x 8 | ¥504 | ¥240,288 | ||
MU3 x 8 | ¥1,096 | ¥527,752 | ||
MU9 x 1 | ¥51 | ¥24,600 | ||
Qwen3.5-27B | qwen3.5-27b | MU2 x 8 | ¥504 | ¥240,288 |
MU3 x 8 | ¥1,096 | ¥527,752 | ||
MU8 x 1 | ¥47 | ¥22,400 | ||
MU9 x 1 | ¥51 | ¥24,600 | ||
Qwen3.5-9B | qwen3.5-9b | MU1 x 2 | ¥108 | ¥52,236 |
MU2 x 8 | ¥504 | ¥240,288 | ||
Qwen3.5-Flash-2026-02-23 | qwen3.5-flash-2026-02-23 | MU1 x 2 | ¥108 | ¥52,236 |
Qwen3.5-Plus-2026-02-15 | qwen3.5-plus-2026-02-15 | MU1 x 8 MU1 x 16 (PD disaggregation mode) | ¥432 PD disaggregation mode: ¥864 | ¥208,944 PD disaggregation mode: ¥417,888 |
MU2 x 8 | ¥504 | ¥240,288 | ||
MU3 x 8 MU3 x 16 (PD disaggregation mode) | ¥1,096 PD disaggregation mode: ¥2,192 | ¥527,752 PD disaggregation mode: ¥1,055,504 | ||
Qwen3-235B-A22B-Instruct-2507 | qwen3-235b-a22b-instruct-2507 | MU1 x 4 | ¥216 | ¥104,472 |
MU2 x 8 | ¥504 | ¥240,288 | ||
MU3 x 8 | ¥1,096 | ¥527,752 | ||
Qwen3-32B | qwen3-32b | MU6 x 16 | ¥400 | ¥193,424 |
Qwen3-30B-A3B-Thinking-2507 | qwen3-30b-a3b-thinking-2507 | MU1 x 2 | ¥108 | ¥52,236 |
Qwen3-4B | qwen3-4b | MU1 x 2 | ¥108 | ¥52,236 |
MU5 x 1 | ¥21 | ¥10,139 | ||
Qwen3-Embedding-0.6B | qwen3-embedding-0.6b | MU5 x 1 | ¥21 | ¥10,139 |
MU6 x 1 | ¥25 | ¥12,089 | ||
Qwen3-MoE-Rerank-0.6B | qwen3-moe-rerank-0.6b | MU5 x 1 | ¥21 | ¥10,139 |
Qwen3-Rerank-0.6B | qwen3-rerank-0.6b | MU5 x 1 | ¥21 | ¥10,139 |
MU6 x 1 | ¥25 | ¥12,089 | ||
Qwen3-Max-2025-09-23 | qwen3-max-2025-09-23 | MU2 x 8 | ¥504 | ¥240,288 |
MU3 x 8 | ¥1,096 | ¥527,752 | ||
Qwen3-Rerank | qwen3-rerank | MU5 x 1 | ¥21 | ¥10,139 |
Qwen2.5-Open-Source-72B | qwen2.5-72b-instruct | MU1 x 8 | ¥432 | ¥208,944 |
Qwen2.5-Open-Source-14B | qwen2.5-14b-instruct | MU1 x 2 | ¥108 | ¥52,236 |
Qwen2.5-Open-Source-7B | qwen2.5-7b-instruct | MU1 x 2 | ¥108 | ¥52,236 |
MU5 x 1 | ¥21 | ¥10,139 | ||
Qwen-Plus-2025-07-28 | qwen-plus-2025-07-28 | MU1 x 4 | ¥216 | ¥104,472 |
Qwen-Plus-2025-12-01 | qwen-plus-2025-12-01 | MU1 x 4 | ¥216 | ¥104,472 |
Qwen-Plus-Character-2025-11-06 | qwen-plus-character-2025-11-06 | MU1 x 4 | ¥216 | ¥104,472 |
GLM
Model Name | Model Code | Model Unit Specification | Hourly unit price (CNY) Minimum billing: minutes | Monthly subscription unit price (CNY) Minimum billing: days |
|---|---|---|---|---|
GLM-5.1 | glm-5.1 | MU2 x 8 | ¥504 | ¥240,288 |
MU3 x 16 (PD disaggregation mode) | PD disaggregation mode: ¥2,192 | PD disaggregation mode: ¥1,055,504 | ||
MU6 x 16 | ¥400 | ¥193,424 | ||
GLM-5 | glm-5 | MU3 x 16 (PD disaggregation mode) | PD disaggregation mode: ¥2,192 | PD disaggregation mode: ¥1,055,504 |
GLM-4.7 | glm-4.7 | MU6 x 32 (PD disaggregation mode) | PD disaggregation mode: ¥800 | PD disaggregation mode: ¥386,848 |
DeepSeek
Model Name | Model Code | Model Unit Specification | Hourly unit price (CNY) Minimum billing: minutes | Monthly subscription unit price (CNY) Minimum billing: days |
|---|---|---|---|---|
DeepSeek-v4-Flash | deepseek-v4-flash | MU1 x 8 | ¥432 | ¥208,944 |
MU3 x 8 | ¥1,096 | ¥527,752 | ||
DeepSeek-v3.2 | deepseek-v3.2 | MU2 x 16 (PD disaggregation mode) | PD disaggregation mode: ¥1,008 | PD disaggregation mode: ¥480,576 |
More models
Model Name | Model Code | Model Unit Specification | Hourly unit price (CNY) Minimum billing: minutes | Monthly subscription unit price (CNY) Minimum billing: days |
|---|---|---|---|---|
Kimi-K2.5 | kimi-k2.5 | MU2 x 8 | ¥504 | ¥240,288 |
Model type:
- Instruct - The model performs inference in non-thinking mode after deployment.
- Thinking - The model performs inference in thinking mode after deployment.
Model deployment type:
-
PD disaggregation mode - Reduces first-token latency and increases throughput.
For models deployed in this mode, during inference the first-token computation (Prefill) and the subsequent token computation (Decode) are split across different compute nodes for execution.
Multimodal
Qwen-VL
Model Name | Model Code | Model Unit Specification | Hourly unit price (CNY) Minimum billing: minutes | Monthly subscription unit price (CNY) Minimum billing: days |
|---|---|---|---|---|
Qwen3-VL-235B-A22B-Thinking | qwen3-vl-235b-a22b-thinking | MU1 x 8 | ¥432 | ¥208,944 |
MU2 x 8 | ¥504 | ¥240,288 | ||
MU3 x 8 | ¥1,096 | ¥527,752 | ||
Qwen3-VL-32B-Instruct | qwen3-vl-32b-instruct | MU2 x 8 | ¥504 | ¥240,288 |
MU3 x 8 | ¥1,096 | ¥527,752 | ||
Qwen3-VL-30B-A3B-Instruct | qwen3-vl-30b-a3b-instruct | MU1 x 4 | ¥216 | ¥104,472 |
Qwen3-VL-8B-Instruct | qwen3-vl-8b-instruct | MU1 x 2 | ¥108 | ¥52,236 |
MU5 x 1 | ¥21 | ¥10,139 | ||
Qwen3-VL-4B-Instruct | qwen3-vl-4b-instruct | MU1 x 2 | ¥108 | ¥52,236 |
Qwen3-VL-2B-Instruct | qwen3-vl-2b-instruct | MU5 x 1 | ¥21 | ¥10,139 |
Qwen3-VL-Embedding-2B | qwen3-vl-embedding-2b | MU5 x 1 | ¥21 | ¥10,139 |
Qwen3-VL-Flash-2025-10-15 | qwen3-vl-flash-2025-10-15 | MU1 x 4 | ¥216 | ¥104,472 |
Qwen3-VL-Plus-2025-09-23 | qwen3-vl-plus-2025-09-23 | MU1 x 4 | ¥216 | ¥104,472 |
Qwen-VL-Max-2025-08-13 | qwen-vl-max-2025-08-13 | MU6 x 4 | ¥100 | ¥48,356 |
Qwen Omni
Model Name | Model Code | Model Unit Specification | Hourly unit price (CNY) Minimum billing: minutes | Monthly subscription unit price (CNY) Minimum billing: days |
|---|---|---|---|---|
Qwen3.5-Omni-Flash | qwen3.5-omni-flash | MU9 x 1 | ¥51 | ¥24,600 |
Model Type:
- Instruct - The model performs inference in non-thinking mode after deployment.
- Thinking - The model performs inference in thinking mode after deployment.
- Instruct/Thinking - You can choose whether to enable thinking mode when deploying the model.
Speech Synthesis
CosyVoiceModel Name | Model Code | Model Unit Specification | Hourly Unit Price (CNY) | Monthly Subscription Unit Price (CNY) |
|---|---|---|---|---|
cosyvoice-v3-flash | cosyvoice-v3-flash | MU5 | ¥21 | ¥10,139 |
Image generation models – Wan
Deployment is free. Invocations are billed at the standard rate of the fine-tuned base model. For the training workflow, see Fine-tune image generation models.
Model ID | LoRA Deployment & Invocation Price |
|---|---|
wan2.7-image-pro | CNY 0.50/image |
wan2.7-image | CNY 0.20/image |
Image generation models – Qwen
Deployment is free. Invocations are billed at the standard rate of the fine-tuned base model. For the training and deployment workflow, see Fine-tune image generation models.
Model ID | LoRA Deployment & Invocation Price |
|---|---|
qwen-image-2.0 | CNY 0.20/image |
qwen-image-2.0-pro | CNY 0.50/image |
Speech synthesis models – CosyVoice
NoteCosyVoice model fine-tuning is available only in the China (Beijing) region.
After deployment, a fine-tuned CosyVoice model is billed by the usage duration of model units. The formula is: Fee = Usage duration (hours) × Number of model units × Model unit price.
Where Number of model units = Model units per replica × Number of replicas. For the available deployment templates and their model unit types, see Deployment templates in Fine-tune CosyVoice.
Model name | Model code | Model unit type | Hourly price (CNY) | Monthly price (CNY) |
|---|---|---|---|---|
cosyvoice-v3-flash | cosyvoice-v3-flash | MU5 | ¥21 | ¥10,139 |
FAQ
Q: When does billing for model deployment start?
A: Billing starts when the model status changes to Running. No charges apply during Deploying, Overdue Payment, or Deployment Failed.
For monthly subscriptions, the billing period starts when the status changes to Running.
Q: Am I charged if I cancel a training job?
A: Yes. If you cancel training manually, you are charged for all tokens processed before cancellation. Training jobs interrupted by system errors or other non-user causes are not charged.
Q: How do I view invocation statistics for a deployed model?
A: Visit the Model Monitoring (Beijing), Model Monitoring (Virginia), or Model Monitoring (Singapore) page.
