Training and deployment pricing

Updated at:

This topic describes the billing rules and pricing for model training and model deployment on Alibaba Cloud Model Studio.

Training billing

Text generation models – Qwen

NoteFor the training workflow, see Introduction to model fine-tuning. After training completes, deploy the new model before evaluating or calling it.

Method

Billed by training tokens

Formula

Model training fee = (Total tokens in training data + Total tokens in mixed training data) × Number of epochs × Training unit price (Minimum billing unit: 1 token)

View the estimated training fee at the bottom of the model training console, and click Computing Details to view the total number of training tokens, number of epochs, and training unit price.

Qwen

Model service

Model Code

Price

Qwen3.8-27B

qwen3.8-27b

¥0.05/1K tokens

Qwen3.7-Plus-2026-05-26

qwen3.7-plus-2026-05-26

¥0.35/1K tokens

Qwen3.6-27B

qwen3.6-27b

¥0.05/1K tokens

Qwen3.6-Flash-2026-04-16

qwen3.6-flash-2026-04-16

¥0.05/1K tokens

Qwen3.6-Plus-2026-04-02

qwen3.6-plus-2026-04-02

¥0.3/1K tokens

Qwen3.5-27B

qwen3.5-27b

¥0.05/1K tokens

Qwen3.5-9B

qwen3.5-9b

¥0.02/1K tokens

Qwen3.5-4B

qwen3.5-4b

¥0.015/1K tokens

Qwen3.5-Flash-2026-02-23

qwen3.5-flash-2026-02-23

¥0.05/1K tokens

Qwen3.5-Plus-2026-02-15

qwen3.5-plus-2026-02-15

¥0.3/1K tokens

Qwen3-32B

qwen3-32b

¥0.04/1K tokens

Qwen3-30B-A3B-Instruct-2507

qwen3-30b-a3b-instruct-2507

¥0.03/1K tokens

Qwen3-14B

qwen3-14b

¥0.03/1K tokens

Qwen3-8B

qwen3-8b

¥0.006/1K tokens

Qwen3-4B-Instruct-2507

qwen3-4b-instruct-2507

¥0.006/1K tokens

Qwen3-1.7B

qwen3-1.7b

¥0.0045/1K tokens

Qwen3-0.6B

qwen3-0.6b

¥0.003/1K tokens

Qwen2.5-72B-Instruct

qwen2.5-72b-instruct

¥0.15/1K tokens

Qwen2.5-32B-Instruct

qwen2.5-32b-instruct

¥0.03/1K tokens

Qwen2.5-14B-Instruct

qwen2.5-14b-instruct

¥0.03/1K tokens

Qwen2.5-7B-Instruct

qwen2.5-7b-instruct

¥0.006/1K tokens

Qwen-Plus-Character-2025-11-06

qwen-plus-character-2025-11-06

¥0.15/1K tokens

Qwen-VL

Model service

Model Code

Price

Qwen3-VL-8B-Instruct

qwen3-vl-8b-instruct

¥0.012/1K tokens

Qwen3-VL-8B-Thinking

qwen3-vl-8b-thinking

¥0.012/1K tokens

Qwen3-VL-4B-Instruct

qwen3-vl-4b-instruct

¥0.006/1K tokens

Qwen2.5-VL-72B-Instruct

qwen2.5-vl-72b-instruct

¥0.05/1K tokens

Qwen2.5-VL-32B-Instruct

qwen2.5-vl-32b-instruct

¥0.02/1K tokens

Qwen2.5-VL-7B-Instruct

qwen2.5-vl-7b-instruct

¥0.01/1K tokens

Calculate tokens for images and videos

Images

Formula: Image Tokens = h_bar * w_bar / token_pixels + 2

  • h_bar, w_bar: The height and width of the scaled image. Before processing an image, the model performs pre-processing to scale it down to a specific pixel limit. This limit depends on the values of the max_pixels and vl_high_resolution_images parameters. For more information, see Process high-resolution images.

  • token_pixels: The pixel value corresponding to each visual token. This varies by model:

    • qwen3.8-series,qwen3.7-series, qwen3.6-series, qwen3.5-series, Qwen3-VL, qwen-vl-max, and qwen-vl-plus:Each token corresponds to 32x32 pixels.
    • QVQ and other Qwen2.5-VL models:Each token corresponds to 28x28 pixels.

The following code demonstrates the approximate image scaling logic used by the model. Use it to estimate the tokens for an image. For actual billing, refer to the API response.

import math
from PIL import Image  # pip install Pillow

def smart_resize(image_path, max_pixels, vl_high_resolution_images):
    """Calculates the scaled dimensions of an image based on model parameters to estimate image tokens."""
    image = Image.open(image_path)
    height, width = image.height, image.width

    # The scaling factor is 32 for models such as Qwen3.6, Qwen3.5, and Qwen3-VL. For other models, it is 28.
    factor = 32
    h_bar = round(height / factor) * factor
    w_bar = round(width / factor) * factor

    # Token lower limit: 4 tokens
    min_pixels = 4 * factor * factor

    # If vl_high_resolution_images=True, the token upper limit is fixed at 16384, and max_pixels is ignored.
    if vl_high_resolution_images:
        max_pixels = 16384 * factor * factor

    # Constrains the total number of pixels to the range [min_pixels, max_pixels].
    if h_bar * w_bar > max_pixels:
        beta = math.sqrt((height * width) / max_pixels)
        h_bar = math.floor(height / beta / factor) * factor
        w_bar = math.floor(width / beta / factor) * factor
    elif h_bar * w_bar < min_pixels:
        beta = math.sqrt(min_pixels / (height * width))
        h_bar = math.ceil(height * beta / factor) * factor
        w_bar = math.ceil(width * beta / factor) * factor

    return h_bar, w_bar

if __name__ == "__main__":
    # Note: The values of max_pixels and vl_high_resolution_images must match the parameters passed when calling the model.
    h_bar, w_bar = smart_resize("xxx/test.jpg", max_pixels=2560 * 32 * 32, vl_high_resolution_images=False)
    print(f"Scaled image dimensions: Height {h_bar}, Width {w_bar}")

    # Each image includes one <vision_bos> and one <vision_eos> token.
    token = int(h_bar * w_bar / (32 * 32)) + 2
    print(f"Number of image tokens: {token}")

Videos

  • Video files:

    When processing a video file, the model first extracts frames and then calculates the total number of tokens for all video frames. Because this calculation is complex, you can use the following code to estimate the total token consumption for a video by providing its path:

# Before use, install: pip install opencv-python
import math
import os
import logging
import cv2

logger = logging.getLogger(__name__)

FRAME_FACTOR = 2

# For models such as Qwen3.6, Qwen3.5, Qwen3-VL, qwen-vl-max-0813, qwen-vl-plus-0815, and qwen-vl-plus-0710, the image scaling factor is 32.
IMAGE_FACTOR = 32

# For other models, the image scaling factor is 28.
# IMAGE_FACTOR = 28

# Maximum aspect ratio for video frames
MAX_RATIO = 200
# Pixel lower limit for video frames
VIDEO_MIN_PIXELS = 4 * 32 * 32
# Pixel upper limit for video frames. For the Qwen3-VL-Plus model, VIDEO_MAX_PIXELS is 640 * 32 * 32. For other models, it is 768 * 32 * 32.
VIDEO_MAX_PIXELS = 640 * 32 * 32

# If the user does not pass the FPS parameter, the default value is used for fps.
FPS = 2.0
# Minimum number of extracted frames
FPS_MIN_FRAMES = 4
# Maximum number of extracted frames (set based on the selected model)
FPS_MAX_FRAMES = 2000

# Maximum pixel value for video input. For the Qwen3-VL-Plus model, set VIDEO_TOTAL_PIXELS to 131072 * 32 * 32. For other models, set it to 65536 * 32 * 32.
VIDEO_TOTAL_PIXELS = int(float(os.environ.get('VIDEO_TOTAL_PIXELS', 131072 * 32 * 32)))

def round_by_factor(number: int, factor: int) -> int:
    """Returns the integer closest to 'number' that is divisible by 'factor'."""
    return round(number / factor) * factor

def ceil_by_factor(number: int, factor: int) -> int:
    """Returns the smallest integer that is greater than or equal to 'number' and divisible by 'factor'."""
    return math.ceil(number / factor) * factor

def floor_by_factor(number: int, factor: int) -> int:
    """Returns the largest integer that is less than or equal to 'number' and divisible by 'factor'."""
    return math.floor(number / factor) * factor

def extract_vision_info(conversations):
    vision_infos = []
    if isinstance(conversations[0], dict):
        conversations = [conversations]
    for conversation in conversations:
        for message in conversation:
            if isinstance(message["content"], list):
                for ele in message["content"]:
                    if (
                        "image" in ele
                        or "image_url" in ele
                        or "video" in ele
                        or ele.get("type","") in ("image", "image_url", "video")
                    ):
                        vision_infos.append(ele)
    return vision_infos

def smart_nframes(ele,total_frames,video_fps):
    """Calculates the number of extracted video frames.

    Args:
        ele (dict): A dictionary containing the video configuration.
            - fps: Controls the number of input frames extracted for the model.
        total_frames (int): The original total number of frames in the video.
        video_fps (int | float): The original frame rate of the video.

    Raises:
        An error is reported if nframes is not within the interval [FRAME_FACTOR, total_frames].

    Returns:
        The number of video frames for model input.
    """
    assert not ("fps" in ele and "nframes" in ele), "Only accept either `fps` or `nframes`"
    fps = ele.get("fps", FPS)
    min_frames = ceil_by_factor(ele.get("min_frames", FPS_MIN_FRAMES), FRAME_FACTOR)
    max_frames = floor_by_factor(ele.get("max_frames", min(FPS_MAX_FRAMES, total_frames)), FRAME_FACTOR)
    duration = total_frames / video_fps if video_fps != 0 else 0
    if duration-int(duration)>(1/fps):
        total_frames = math.ceil(duration * video_fps)
    else:
        total_frames = math.ceil(int(duration)*video_fps)
    nframes = total_frames / video_fps * fps
    if nframes > total_frames:
        logger.warning(f"smart_nframes: nframes[{nframes}] > total_frames[{total_frames}]")
    nframes = int(min(min(max(nframes, min_frames), max_frames), total_frames))
    if not (FRAME_FACTOR <= nframes and nframes <= total_frames):
        raise ValueError(f"nframes should in interval [{FRAME_FACTOR}, {total_frames}], but got {nframes}.")

    return nframes

def get_video(video_path):
    # Get video information
    cap = cv2.VideoCapture(video_path)

    frame_width = int(cap.get(cv2.CAP_PROP_FRAME_WIDTH))
    # Get video height
    frame_height = int(cap.get(cv2.CAP_PROP_FRAME_HEIGHT))
    total_frames = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))

    video_fps = cap.get(cv2.CAP_PROP_FPS)
    return frame_height, frame_width, total_frames, video_fps

def smart_resize(ele, path, factor=IMAGE_FACTOR):
    # Get the original width and height of the video
    height, width, total_frames, video_fps = get_video(path)
    # Token lower limit for video frames
    min_pixels = VIDEO_MIN_PIXELS
    total_pixels = VIDEO_TOTAL_PIXELS
    # Number of extracted video frames
    nframes = smart_nframes(ele, total_frames, video_fps)
    max_pixels = max(min(VIDEO_MAX_PIXELS, total_pixels / nframes * FRAME_FACTOR),int(min_pixels * 1.05))

    # The aspect ratio of the video should not exceed 200:1 or 1:200.
    if max(height, width) / min(height, width) > MAX_RATIO:
        raise ValueError(
            f"absolute aspect ratio must be smaller than {MAX_RATIO}, got {max(height, width) / min(height, width)}"
        )

    h_bar = max(factor, round_by_factor(height, factor))
    w_bar = max(factor, round_by_factor(width, factor))
    if h_bar * w_bar > max_pixels:
        beta = math.sqrt((height * width) / max_pixels)
        h_bar = floor_by_factor(height / beta, factor)
        w_bar = floor_by_factor(width / beta, factor)
    elif h_bar * w_bar < min_pixels:
        beta = math.sqrt(min_pixels / (height * width))
        h_bar = ceil_by_factor(height * beta, factor)
        w_bar = ceil_by_factor(width * beta, factor)
    return h_bar, w_bar

def token_calculate(video_path, fps):
    # Pass the video path and the fps frame extraction parameter.
    messages = [{"content": [{"video": video_path, "fps": fps}]}]
    vision_infos = extract_vision_info(messages)[0]

    resized_height, resized_width = smart_resize(vision_infos, video_path)

    height, width, total_frames, video_fps = get_video(video_path)
    num_frames = smart_nframes(vision_infos, total_frames, video_fps)
    print(f"Original video dimensions: {height}*{width}, Model input dimensions: {resized_height}*{resized_width}, Total video frames: {total_frames}, Total frames extracted when fps is {fps}: {num_frames}", end=", ")
    video_token = int(math.ceil(num_frames / 2) * resized_height / 32 * resized_width / 32)
    video_token += 2   # The system automatically adds <|vision_bos|> and <|vision_eos|> visual markers (1 token each).
    return video_token

video_token = token_calculate("xxx/test.mp4", 1)
print("Video tokens:", video_token)
  • Image list:

    When a video is passed as a list of images, it means that frame extraction has already been performed. Use the following code to calculate the token consumption by providing the path and number of images:

# Before use, install: pip install Pillow
import math
import os
import logging
from typing import Tuple
from PIL import Image

logger = logging.getLogger(__name__)

# ==================== Constant Definitions ====================
FRAME_FACTOR = 2
# For models such as Qwen3-VL, qwen-vl-max-0813, qwen-vl-plus-0815, and qwen-vl-plus-0710, the scaling factor is 32.
IMAGE_FACTOR = 32

# For other models, the scaling factor is 28.
# IMAGE_FACTOR = 28

# Constants for token calculation
TOKEN_DIVISOR = 32  # Divisor for token calculation
VISION_SPECIAL_TOKENS = 2  # <|vision_bos|> and <|vision_eos|> markers

# Maximum aspect ratio for video frames
MAX_RATIO = 200
# Pixel lower limit for video frames
VIDEO_MIN_PIXELS = 4 * 32 * 32
# Pixel upper limit for video frames. For the Qwen3-VL-Plus model, VIDEO_MAX_PIXELS is 640 * 32 * 32. For other models, it is 768 * 32 * 32.
VIDEO_MAX_PIXELS = 640 * 32 * 32

# Maximum pixel value for video input. For the Qwen3-VL-Plus model, set VIDEO_TOTAL_PIXELS to 131072 * 32 * 32. For other models, set it to 65536 * 32 * 32.
VIDEO_TOTAL_PIXELS = int(float(os.environ.get('VIDEO_TOTAL_PIXELS', 131072 * 32 * 32)))

def round_by_factor(number: int, factor: int) -> int:
    """Returns the integer closest to 'number' that is divisible by 'factor'."""
    return round(number / factor) * factor

def ceil_by_factor(number: int, factor: int) -> int:
    """Returns the smallest integer that is greater than or equal to 'number' and divisible by 'factor'."""
    return math.ceil(number / factor) * factor

def floor_by_factor(number: int, factor: int) -> int:
    """Returns the largest integer that is less than or equal to 'number' and divisible by 'factor'."""
    return math.floor(number / factor) * factor

def get_image_size(image_path: str) -> Tuple[int, int]:
    if not os.path.exists(image_path):
        raise FileNotFoundError(f"Image file not found: {image_path}")

    try:
        image = Image.open(image_path)
        height = image.height
        width = image.width
        image.close()  # Close the file promptly
        return height, width
    except Exception as e:
        raise ValueError(f"Cannot read image file {image_path}: {str(e)}")

def smart_resize(height: int, width: int, nframes: int, factor: int = IMAGE_FACTOR) -> Tuple[int, int]:
    """
    Calculates the scaled dimensions of an image

    Args:
        height: Original image height
        width: Original image width
        nframes: Number of video frames
        factor: Scaling factor, defaults to IMAGE_FACTOR

    Returns:
        (resized_height, resized_width) The scaled height and width

    Raises:
        ValueError: Aspect ratio exceeds the limit
    """
    # Token lower limit for video frames
    min_pixels = VIDEO_MIN_PIXELS
    total_pixels = VIDEO_TOTAL_PIXELS
    # Number of extracted video frames
    max_pixels = max(min(VIDEO_MAX_PIXELS, total_pixels / nframes * FRAME_FACTOR), int(min_pixels * 1.05))

    # The aspect ratio of the video should not exceed 200:1 or 1:200.
    aspect_ratio = max(height, width) / min(height, width)
    if aspect_ratio > MAX_RATIO:
        raise ValueError(
            f"Image aspect ratio must be less than {MAX_RATIO}:1, but is currently {aspect_ratio:.2f}:1"
        )

    h_bar = max(factor, round_by_factor(height, factor))
    w_bar = max(factor, round_by_factor(width, factor))
    if h_bar * w_bar > max_pixels:
        beta = math.sqrt((height * width) / max_pixels)
        h_bar = floor_by_factor(height / beta, factor)
        w_bar = floor_by_factor(width / beta, factor)
    elif h_bar * w_bar < min_pixels:
        beta = math.sqrt(min_pixels / (height * width))
        h_bar = ceil_by_factor(height * beta, factor)
        w_bar = ceil_by_factor(width * beta, factor)
    return h_bar, w_bar

def calculate_video_tokens(image_path: str, nframes: int = 1, factor: int = IMAGE_FACTOR, verbose: bool = True) -> int:
    """

    Args:
        image_path: Path to the video frame file
        nframes: Number of video frames,
        factor: Scaling factor, defaults to IMAGE_FACTOR
        verbose: Whether to print detailed information

    Returns:
        The number of tokens consumed

    Raises:
        FileNotFoundError: The file does not exist
        ValueError: The file format is invalid or the aspect ratio exceeds the limit
    """
    # Get the original image dimensions (read only once)
    height, width = get_image_size(image_path)

    # Calculate the scaled dimensions
    resized_height, resized_width = smart_resize(height, width, nframes, factor)

    # Calculate the number of tokens
    # Formula: ceil(nframes/2) * (height/TOKEN_DIVISOR) * (width/TOKEN_DIVISOR) + VISION_SPECIAL_TOKENS
    video_token = int(
        math.ceil(nframes / 2) *
        (resized_height / TOKEN_DIVISOR) *
        (resized_width / TOKEN_DIVISOR)
    )
    # Add visual marker tokens (<|vision_bos|> and <|vision_eos|>)
    video_token += VISION_SPECIAL_TOKENS

    if verbose:
        print(f"Original video frame dimensions: {height}x{width}, Model input dimensions: {resized_height}x{resized_width}, ", end="")

    return video_token

if __name__ == "__main__":
    try:
        video_token = calculate_video_tokens("xxx/test.jpg", nframes=30)
        print(f"Video tokens: {video_token}\n")
    except Exception as e:
        print(f"Error: {str(e)}\n")

Image generation models – Wan

NoteFor the training workflow, see Fine-tune image generation models. After training completes, deploy the new model before calling it.

Method

Billed by training tokens

Formula

Model training fee = Total training tokens × Training unit price (Billing unit: per 1,000 tokens)

Formula for total training tokens

Training Tokens Total ≈ max_steps × Lstep

Where:

  • max_steps: A hyperparameter specified during training, representing the maximum number of training steps (configured when creating a fine-tuning job).

  • Lstep: The token consumption per step. The formula is:

    Lstep = ∑i∈batch Litem(i) ≤ Lmax

    Lstep is approximately equal to Lmax. Lmax is determined by the max_token_length and generation_type, as shown below:

generation_type

max_token_length

Lmax

t2i (text-to-image)

1k

12,800

2k

23,220

i2i (image-to-image)

1k

23,220

2k

32,000

NoteThe above formula provides an approximation. Actual billing is based on the usage field returned by the system.

Model

Code

Training price (per 1K tokens)

Wan image generation

wan2.7-image-pro

CNY 0.08

Wan image generation

wan2.7-image

CNY 0.08

Billing example

Suppose you fine-tune the wan2.7-image-pro model for t2i. The parameters are: max_steps = 200, max_token_length = "1k", and the training price is CNY 0.08 per 1,000 tokens:

  • From the table: Lmax = 12,800 (generation_type=t2i, max_token_length=1k), Lstep ≈ Lmax = 12,800
  • Total training tokens ≈ 200 × 12800 = 2560000 = 2560 thousand tokens
  • Model training fee ≈ 2560 × 0.08 = CNY 204.8

Image generation models – Qwen

NoteFor the training workflow, see Fine-tune image generation models. After training completes, deploy the new model before calling it.

Method

Billed by training tokens

Formula

Model training fee = Total training tokens × Training unit price (Billing unit: per 1,000 tokens)

Formula for total training tokens

Total training tokens = n_epochs × N × (max_pixels / Compression ratio) × GPU coefficient

Where:

  • n_epochs: A hyperparameter specified during training, representing the number of training epochs (configured when creating a fine-tuning job).
  • N: The total number of images in the training set.
  • max_pixels: A hyperparameter specified during training, representing the maximum number of pixels per image (configured when creating a fine-tuning job). Valid values are 1k (1024×1024) and 2k (2048×2048).
  • Compression ratio: The fixed VAE compression ratio, which is 1024 (32×32).
  • GPU coefficient: A resource scheduling coefficient that is dynamically adjusted based on job scheduling.

Model

Code

Training price (per 1K tokens)

Qwen image generation

qwen-image-2.0

CNY 0.02

Qwen image generation

qwen-image-2.0-pro

CNY 0.02

Billing example

Suppose you fine-tune the qwen-image-2.0 model with a training set of 1 image. The parameters are: n_epochs = 1000, max_pixels = "1k", the training price is CNY 0.02 per 1,000 tokens, and the GPU coefficient is 8:

  • Tokens per image = max_pixels / Compression ratio = 1024×1024 / 1024 = 1024
  • Total training tokens = 1000 × 1 × 1024 × 8 = 8192000 = 8192 thousand tokens
  • Model training fee = 8192 × 0.02 = CNY 163.84

The following table estimates token consumption and fees per training image, assuming a GPU coefficient of 8. For a training set of N images, multiply the values by N.

max_pixels

n_epochs

Estimated tokens

Estimated fee (CNY)

1k

800

6,553,600

131.07

1k

1,000

8,192,000

163.84

1k

2,000

16,384,000

327.68

2k

800

26,214,400

524.29

2k

1,000

32,768,000

655.36

2k

2,000

65,536,000

1310.72

Note

  • The GPU coefficient is dynamically adjusted based on job scheduling. The value 8 in the example and the table above is used only to illustrate the calculation. For actual token consumption, see the usage field returned by the Query a fine-tuning job operation. Your bill is the final authority on fees.
  • batch_size does not affect billing or the total training tokens.

Video generation models – Wan

NoteFor the training workflow, see Fine-tuning video generation models. After training completes, deploy the new model before calling it.

Method

Billed by training tokens

Formula

Model training fee = Total training tokens × Training unit price (Billing unit: per 1,000 tokens)

Formula for total training tokens

Training Tokens Total = (∑i=1N billing duration of videoi) × (max_pixels / 1024) × n_epochs

Where:

  • N: Total number of videos in the training set.

  • max_pixels: A hyperparameter specified during training, representing the maximum number of pixels for a video (configured when creating a fine-tuning job).

  • n_epochs: A hyperparameter specified during training, representing the number of loops (configured when creating a fine-tuning job).

    • The conversion between n_epochs and steps is: steps = n_epochs × ⌈dataset_size / batch_size⌉, i.e., n_epochs = steps / ⌈dataset_size / batch_size⌉.
    • When the dataset contains only 1 sample and batch_size = 1, n_epochs = steps. We recommend a total of at least 800 steps.
  • Billing duration calculation rule for a single video: First, round the original video duration (in seconds) to the nearest integer, then determine the final value based on model limits.

    • wan2.7 model: Billing duration=min(10, rounded duration), meaning a single video is billed for a maximum of 10 seconds.
    • wan2.5 model: Billing duration=min(10, rounded duration), meaning a single video is billed for a maximum of 10 seconds.
    • wan2.2 model: Billing duration=min(5, rounded duration), meaning a single video is billed for a maximum of 5 seconds.

Model

Code

Training price (per 1K tokens)

Wan image-to-video (first frame-based)

wan2.7-i2v

CNY 2

wan2.2-i2v-flash

CNY 0.06

wan2.5-i2v-preview

CNY 0.32

Image-to-video (first and last frame-based)

wan2.7-i2v

CNY 2

wan2.2-kf2v-flash

CNY 0.06

Billing examples

  1. wan2.7-i2v cost estimation (single data)

Assume a training set contains 1 video with a duration of 10 seconds. With batch_size = 1 (recommended), n_epochs = steps / ⌈1(dataset_size) / 1(batch_size)⌉ = steps.

Training unit price = CNY 2/thousand tokens. Taking max_pixels = 36864 and n_epochs = 800 as an example:

  • Total training tokens = 10 × (36864 / 1024) × 800 = 288,000 = 288 thousand tokens
  • Model training fee = 288 × 2 = CNY 576

max_pixels

Common steps

n_epochs

Estimated tokens

Estimated cost (CNY)

36864

800

800

288,000

576

1,000

1,000

360,000

720

2,000

2,000

720,000

1,440

65536

800

800

512,000

1,024

1,000

1,000

640,000

1,280

2,000

2,000

1,280,000

2,560

102400

800

800

800,000

1,600

1,000

1,000

1,000,000

2,000

2,000

2,000

2,000,000

4,000

  1. wan2.7-i2v cost estimation (multiple data)

Assume a training set contains 2 videos with durations of 3.4 seconds and 11.5 seconds. Parameters: max_pixels = 36864, n_epochs = 800. Training unit price = CNY 2/thousand tokens:

  • Duration calculation:

    • Video 1: 3.4 seconds is rounded to 3. Billable duration = min(10, 3) = 3.
    • Video 2: 11.5 seconds is rounded to 11. Billable duration = min(10, 11) = 10.
    • Total billable duration = 3 + 10 = 13 seconds.
  • Total training tokens = 13 × (36864/1024) × 800 = 374,400 = 374.4 thousand tokens.

  • Model training fee = 374.4 × 2 = CNY 748.8.

Speech synthesis models – CosyVoice

NoteCosyVoice model fine-tuning is available only in the China (Beijing) region. For the training workflow, see Fine-tune CosyVoice. After training completes, deploy the new model before calling it.

Method

Billed by tokens consumed during training

Unit price

CNY 0.2 per 1,000 tokens

Formula

Model training fee = Total consumed tokens × Training unit price

Formula for total consumed tokens

The token consumption of a single job is estimated as follows:

Consumed tokens = (lm_max_epoch + fm_max_epoch) × 25 × Training set duration (s)

Where lm_max_epoch and fm_max_epoch are hyperparameters set when creating the fine-tuning job, representing the number of LM and FM training epochs. The training set duration is the total duration in seconds of all audio files in the fine-tuning dataset. Increasing either epoch count or enlarging the training set increases token consumption linearly.

Deployment billing

Text generation models: Qwen

Billing by usage duration (Provisioned Throughput)

Fee = Usage duration × (Input TPM unit price × Input TPM + Output TPM unit price × Output TPM)

Post-paid is calculated hourly: the usage duration unit is hours, and the unit price is taken from the "Continuous 1 hour" column in the table below; prepaid is calculated daily: the usage duration unit is days, and the unit price is taken from the "Continuous 1 day" column in the table below.

  • Prepaid orders take effect in real time after payment, with a validity period of N days ending at 23:59 on day N. If the order is placed after 22:00, the expiration date will be automatically extended by 1 day.
  • After a prepaid order expires, the service will be stopped with a 2-hour delay, and resources will be retained for 14 hours after the stop and then released.
  • Prepaid orders support early termination; the used portion is settled at a 1.2x coefficient for refund. For details, see Refund rules for configuration downgrades.
  • For post-paid billing, if the account is in arrears, the deployed resources will continue to be retained and billed for 24 hours, during which the service can still be used normally. After 24 hours, the system stops billing, the model deployment enters an arrears state, and the underlying resources will be deleted, but the model deployment task will be retained. After the arrears are paid, the system will reallocate resources and restore usage (fees will continue to accrue after restoration). If you do not want to continue incurring fees, you can delete the model deployment task, and billing will stop after successful deletion.

When the model input exceeds the maximum input Token, the relevant call will automatically switch to the pay-as-you-go mode of the current model; when the purchased TPM is exceeded, it is handled according to the overflow strategy selected at creation ("auto-overflow" switches to pay-as-you-go, "use-only-PTU-capacity" returns 429). At this time, inference performance may degrade and will be subject to the public traffic control of the current snapshot model in the business space, and fees will be charged according to the model invocation (pay-as-you-go) standard.

  • In this case (only under the "auto-overflow" strategy), the API response Header will include: x-dashscope-ptu-overflow:true.
  • For TPM statistics, go to: Model Monitoring (Beijing).

For the specific fee reduction and refund rules in scale-down (downgrade) scenarios, please refer to: Refund rules for configuration downgrades.

NotePTU deployment supports long-input tiered capacity coefficients and cache discounts; see Provisioned Throughput long input and caching for details.

North China 2 (Beijing)

Qwen

Model Name

Model Code

Max input tokens

Post-paid input

Per 10K TPM/hour

Post-paid output

Per 1K TPM/hour

Prepaid input

Per 10K TPM/day

Prepaid output

Per 1K TPM/day

Qwen3.8-Max

qwen3.8-max

1M

¥28.8

¥8.64

¥345.6

¥103.68

Qwen3.7-Flash-2026-07-15

qwen3.7-flash-2026-07-15

128K

¥0.48

¥0.19

¥5.76

¥2.3

Qwen3.7-Max-2026-05-20

qwen3.7-max-2026-05-20

256K

¥28.8

¥8.64

¥345.6

¥103.68

Qwen3.7-Plus-2026-05-26

qwen3.7-plus-2026-05-26

256K

¥4.8

¥1.92

¥57.6

¥23.04

Qwen3.6-Plus-2026-04-02

qwen3.6-plus-2026-04-02

128K

¥4.8

¥2.88

¥57.6

¥34.56

Qwen3.5-Plus-2026-04-20

qwen3.5-plus-2026-04-20

128K

¥1.92

¥1.15

¥23.04

¥13.82

Qwen3-Max-2025-09-23

qwen3-max-2025-09-23

128K

¥7.68

¥3.08

¥92.16

¥36.96

Qwen-Flash-2025-07-28

qwen-flash-2025-07-28

128K

¥0.36

¥0.36

¥4.32

¥4.32

Qwen-Plus-2025-12-01

qwen-plus-2025-12-01

128K

¥1.92

Non-thinking: ¥0.48

Thinking: ¥1.92

¥23.04

Non-thinking: ¥5.76

Thinking: ¥23.04

DeepSeek

Model Name

Model Code

Max input tokens

Post-paid input

Per 10K TPM/hour

Post-paid output

Per 1K TPM/hour

Prepaid input

Per 10K TPM/day

Prepaid output

Per 1K TPM/day

DeepSeek-v4-Flash

deepseek-v4-flash

256K

¥3.6

¥0.72

¥43.2

¥8.64

DeepSeek-v4-Flash-0731

deepseek-v4-flash-0731

256K

¥7.2

¥1.44

¥86.4

¥17.28

DeepSeek-v4-Pro

deepseek-v4-pro

256K

¥43.2

¥8.64

¥518.4

¥103.68

DeepSeek-v3

deepseek-v3

64K

¥7.2

¥2.88

¥86.4

¥34.56

Qwen-VL

Model Name

Model Code

Max input tokens

Post-paid input

Per 10K TPM/hour

Post-paid output

Per 1K TPM/hour

Prepaid input

Per 10K TPM/day

Prepaid output

Per 1K TPM/day

Qwen3-VL-Plus-2025-09-23

qwen3-vl-plus-2025-09-23

128K

¥2.4

¥2.4

¥28.8

¥28.8

GLM

Model Name

Model Code

Max input tokens

Post-paid input

Per 10K TPM/hour

Post-paid output

Per 1K TPM/hour

Prepaid input

Per 10K TPM/day

Prepaid output

Per 1K TPM/day

GLM-5.2

glm-5.2

1M

¥28.8

¥10.08

¥345.6

¥120.96

Singapore

Qwen

Model Name

Model Code

Max input tokens

Post-paid input

Per 10K TPM/hour

Post-paid output

Per 1K TPM/hour

Prepaid input

Per 10K TPM/day

Prepaid output

Per 1K TPM/day

Qwen3.8-Max

qwen3.8-max

1M

¥35.97

¥10.79

¥431.7

¥129.5

Qwen3.7-Flash-2026-07-15

qwen3.7-flash-2026-07-15

128K

¥0.54

¥0.23

¥6.47

¥2.81

Qwen3.7-Max-2026-05-20

qwen3.7-max-2026-05-20

256K

¥44.97

¥13.49

¥539.6

¥161.87

Qwen3.7-Plus-2026-05-26

qwen3.7-plus-2026-05-26

256K

¥7.19

¥2.88

¥86.3

¥34.53

Qwen3.6-Plus-2026-04-02

qwen3.6-plus-2026-04-02

128K

¥9

¥5.4

¥107.9

¥64.75

Qwen3.5-Plus-2026-04-20

qwen3.5-plus-2026-04-20

128K

¥7.2

¥4.32

¥86.3

¥51.8

DeepSeek

Model Name

Model Code

Max input tokens

Post-paid input

Per 10K TPM/hour

Post-paid output

Per 1K TPM/hour

Prepaid input

Per 10K TPM/day

Prepaid output

Per 1K TPM/day

DeepSeek-v4-Flash

deepseek-v4-flash

256K

¥5.4

¥1.08

¥64.8

¥12.95

DeepSeek-v4-Flash-0731

deepseek-v4-flash-0731

256K

¥10.79

¥2.16

¥129.5

¥25.9

DeepSeek-v4-Pro

deepseek-v4-pro

256K

¥64.75

¥12.95

¥777

¥155.4

Qwen-VL

Model Name

Model Code

Max input tokens

Post-paid input

Per 10K TPM/hour

Post-paid output

Per 1K TPM/hour

Prepaid input

Per 10K TPM/day

Prepaid output

Per 1K TPM/day

Qwen3-VL-Plus-2025-09-23

qwen3-vl-plus-2025-09-23

128K

¥3.6

¥2.88

¥43.2

¥34.53

GLM

Model Name

Model Code

Max input tokens

Post-paid input

Per 10K TPM/hour

Post-paid output

Per 1K TPM/hour

Prepaid input

Per 10K TPM/day

Prepaid output

Per 1K TPM/day

GLM-5.2

glm-5.2

1M

¥37.8

¥11.87

¥453.3

¥142.45

Billing by usage duration (Dedicated Throughput)

North China 2 (Beijing)

Model

Type

Base input (TPM)

Base output (TPM)

Starting price (CNY/month)

Scale-out price (CNY/month)

qwen3.7-plus-2026-05-26

Fast

1,372,000

170,000

556,466

556,466

qwen3.6-27b

Standard

273,000

34,000

49,118

49,118

qwen3.5-397b-a17b

Standard

896,000

112,000

1,107,904

1,107,904

qwen3.5-122b-a10b

Fast

7,288,000

904,000

2,226,984

2,226,984

qwen3.5-35b-a3b

Standard

656,000

82,000

109,962

109,962

qwen3.5-27b

Standard

2,432,000

304,000

1,108,688

1,108,688

Efficient

291,000

36,000

48,996

48,996

qwen3.5-4b

Standard

1,072,000

134,000

109,478

109,478

glm-5.2

Fast

1,004,000

120,000

1,895,260

1,895,260

glm-5.1

Standard

256,000

32,000

504,704

504,704

deepseek-v4-flash

Standard

2,240,000

280,000

1,107,400

1,107,400

deepseek-v4-flash-0731

Fast

1,676,000

213,000

277,984

277,984

Model

Type

Input length

Output length

Cache Hit Rate

First token latency (ms)

Per-token latency (ms)

qwen3.7-plus-2026-05-26

Fast

16,000

2,000

0

2,418

15

qwen3.6-27b

Standard

4,000

500

0

1,292

19

qwen3.5-397b-a17b

Standard

4,000

500

0

996

27

qwen3.5-122b-a10b

Fast

16,000

2,000

0

568

8

qwen3.5-35b-a3b

Standard

4,000

500

0

471

10

qwen3.5-27b

Standard

4,000

500

0

703

14

Efficient

4,000

500

0

1,448

23

qwen3.5-4b

Standard

4,000

500

0

552

6

glm-5.2

Fast

16,000

2,000

0

1,558

15

glm-5.1

Standard

4,000

500

0

769

27

deepseek-v4-flash

Standard

4,000

500

0

651

19

deepseek-v4-flash-0731

Fast

16,000

2,000

0

1,079

13

By model token usage

Fee = Model input token count × Model input unit price + Model output token count × Model output unit price (minimum billing unit: 1 token)

  • Billing by model token usage is supported only after you complete efficient SFT training (that is, LoRA efficient fine-tuning; the plan parameter is set to lora for API deployment) on the following base models and obtain a custom model.

Beijing

Base model

Model Code

Input

CNY/million tokens

Output

CNY/million tokens

Qwen3.6-27B

qwen3.6-27b

<256K ¥3

<256K ¥18

Qwen3.5-27B

qwen3.5-27b

<128K ¥0.6

128K-256K ¥1.8

<128K ¥4.8

128K-256K ¥14.4

Qwen3-32B

qwen3-32b

Non-thinking mode: ¥2

Thinking mode: ¥2

Non-thinking mode: ¥8

Thinking mode: ¥20

Qwen3-14B

qwen3-14b

Non-thinking mode: ¥1

Thinking mode: ¥1

Non-thinking mode: ¥4

Thinking mode: ¥10

Qwen3-8B

qwen3-8b

Non-thinking mode: ¥0.5

Thinking mode: ¥0.5

Non-thinking mode: ¥2

Thinking mode: ¥5

Qwen3-VL-8B-Instruct

qwen3-vl-8b-instruct

¥0.5

¥2

Qwen3-4B-Instruct-2507

qwen3-4b-instruct-2507

Non-thinking mode: ¥0.3

Thinking mode: ¥0.3

Non-thinking mode: ¥1.2

Thinking mode: ¥3

Qwen2.5-Open-Source-72B

qwen2.5-72b-instruct

¥4

¥12

Qwen2.5-VL-72B

qwen2.5-vl-72b-instruct

¥16

¥48

Qwen2.5-Open-Source-32B

qwen2.5-32b-instruct

¥2

¥6

Qwen2.5-VL-32B

qwen2.5-vl-32b-instruct

¥8

¥24

Qwen2.5-Open-Source-14B

qwen2.5-14b-instruct

¥1

¥3

Qwen2.5-Open-Source-7B

qwen2.5-7b-instruct

¥0.5

¥1

Qwen2.5-VL-7B

qwen2.5-vl-7b-instruct

¥2

¥5

Singapore

Base model

Model Code

Input

CNY/million tokens

Output

CNY/million tokens

Qwen3.6-27B

qwen3.6-27b

<256K ¥4.497

<256K ¥26.979

Qwen3.5-27B

qwen3.5-27b

¥2.202

¥17.614

Qwen3-32B

qwen3-32b

Non-thinking mode: ¥1.174

Thinking mode: ¥1.174

Non-thinking mode: ¥4.697

Thinking mode: ¥4.697

Qwen3-14B

qwen3-14b

Non-thinking mode: ¥2.569

Thinking mode: ¥2.569

Non-thinking mode: ¥10.275

Thinking mode: ¥30.825

Qwen3-8B

qwen3-8b

Non-thinking mode: ¥1.321

Thinking mode: ¥1.321

Non-thinking mode: ¥5.137

Thinking mode: ¥15.412

Qwen3-VL-8B-Instruct

qwen3-vl-8b-instruct

¥1.321

¥5.137

Qwen3-4B-Instruct-2507

qwen3-4b-instruct-2507

Non-thinking mode: ¥0.807

Thinking mode: ¥0.807

Non-thinking mode: ¥3.082

Thinking mode: ¥9.247

Qwen2.5-Open-Source-72B

qwen2.5-72b-instruct

¥10.275

¥41.1

Qwen2.5-VL-72B

qwen2.5-vl-72b-instruct

¥20.55

¥61.65

Qwen2.5-Open-Source-32B

qwen2.5-32b-instruct

¥5.137

¥20.55

Qwen2.5-VL-32B

qwen2.5-vl-32b-instruct

¥10.275

¥30.825

Qwen2.5-Open-Source-14B

qwen2.5-14b-instruct

¥2.569

¥10.275

Qwen2.5-Open-Source-7B

qwen2.5-7b-instruct

¥1.284

¥5.137

Qwen2.5-VL-7B

qwen2.5-vl-7b-instruct

¥2.569

¥7.706

Billing by usage duration (Model Unit)

Cost = Usage duration (hours) × Number of Model Units × Model Unit unit price

The "Model Unit unit price" takes the "Hourly unit price" column in the table below for postpaid scenarios; for monthly subscription prepaid billing, the formula becomes Number of monthly subscriptions × Number of Model Units × Monthly unit price.

  • For the first month of a prepaid purchase, if you cancel the subscription early within the first month, the daily unit price (≈ Monthly unit price / 30) will be billed at 1.2 times the rate (less than one day is billed as one day)

NoteThe compute resources for the Model Unit postpaid method are first-come, first-served. If the purchase fails, a full refund will be issued.

Text generation

Qwen

Model Name

Model Code

Model Unit Specification

Hourly unit price (CNY)

Minimum billing: minutes

Monthly subscription unit price (CNY)

Minimum billing: days

Qwen3.6-35B-A3B

qwen3.6-35b-a3b

MU1 x 8

¥432

¥208,944

MU2 x 8

¥504

¥240,288

MU3 x 8

¥1,096

¥527,752

MU8 x 1

¥47

¥22,400

MU9 x 1

¥51

¥24,600

Qwen3.6-27B

qwen3.6-27b

MU9 x 1

¥51

¥24,600

Qwen3.6-Flash-2026-04-16

qwen3.6-flash-2026-04-16

MU1 x 2

¥108

¥52,236

MU3 x 8

¥1,096

¥527,752

Qwen3.6-Plus-2026-04-02

qwen3.6-plus-2026-04-02

MU1 x 8

MU1 x 16 (PD disaggregation mode)

¥432

PD disaggregation mode: ¥864

¥208,944

PD disaggregation mode: ¥417,888

MU2 x 8

¥504

¥240,288

Qwen3.5-397B-A17B

qwen3.5-397b-a17b

MU3 x 8

MU3 x 16 (PD disaggregation mode)

¥1,096

PD disaggregation mode: ¥2,192

¥527,752

PD disaggregation mode: ¥1,055,504

MU6 x 16

¥400

¥193,424

Qwen3.5-122B-A10B

qwen3.5-122b-a10b

MU1 x 4

¥216

¥104,472

MU6 x 16

¥400

¥193,424

Qwen3.5-35B-A3B

qwen3.5-35b-a3b

MU1 x 2

¥108

¥52,236

MU2 x 8

¥504

¥240,288

MU3 x 8

¥1,096

¥527,752

MU9 x 1

¥51

¥24,600

Qwen3.5-27B

qwen3.5-27b

MU2 x 8

¥504

¥240,288

MU3 x 8

¥1,096

¥527,752

MU8 x 1

¥47

¥22,400

MU9 x 1

¥51

¥24,600

Qwen3.5-9B

qwen3.5-9b

MU1 x 2

¥108

¥52,236

MU2 x 8

¥504

¥240,288

Qwen3.5-Flash-2026-02-23

qwen3.5-flash-2026-02-23

MU1 x 2

¥108

¥52,236

Qwen3.5-Plus-2026-02-15

qwen3.5-plus-2026-02-15

MU1 x 8

MU1 x 16 (PD disaggregation mode)

¥432

PD disaggregation mode: ¥864

¥208,944

PD disaggregation mode: ¥417,888

MU2 x 8

¥504

¥240,288

MU3 x 8

MU3 x 16 (PD disaggregation mode)

¥1,096

PD disaggregation mode: ¥2,192

¥527,752

PD disaggregation mode: ¥1,055,504

Qwen3-235B-A22B-Instruct-2507

qwen3-235b-a22b-instruct-2507

MU1 x 4

¥216

¥104,472

MU2 x 8

¥504

¥240,288

MU3 x 8

¥1,096

¥527,752

Qwen3-32B

qwen3-32b

MU6 x 16

¥400

¥193,424

Qwen3-30B-A3B-Thinking-2507

qwen3-30b-a3b-thinking-2507

MU1 x 2

¥108

¥52,236

Qwen3-4B

qwen3-4b

MU1 x 2

¥108

¥52,236

MU5 x 1

¥21

¥10,139

Qwen3-Embedding-0.6B

qwen3-embedding-0.6b

MU5 x 1

¥21

¥10,139

MU6 x 1

¥25

¥12,089

Qwen3-MoE-Rerank-0.6B

qwen3-moe-rerank-0.6b

MU5 x 1

¥21

¥10,139

Qwen3-Rerank-0.6B

qwen3-rerank-0.6b

MU5 x 1

¥21

¥10,139

MU6 x 1

¥25

¥12,089

Qwen3-Max-2025-09-23

qwen3-max-2025-09-23

MU2 x 8

¥504

¥240,288

MU3 x 8

¥1,096

¥527,752

Qwen3-Rerank

qwen3-rerank

MU5 x 1

¥21

¥10,139

Qwen2.5-Open-Source-72B

qwen2.5-72b-instruct

MU1 x 8

¥432

¥208,944

Qwen2.5-Open-Source-14B

qwen2.5-14b-instruct

MU1 x 2

¥108

¥52,236

Qwen2.5-Open-Source-7B

qwen2.5-7b-instruct

MU1 x 2

¥108

¥52,236

MU5 x 1

¥21

¥10,139

Qwen-Plus-2025-07-28

qwen-plus-2025-07-28

MU1 x 4

¥216

¥104,472

Qwen-Plus-2025-12-01

qwen-plus-2025-12-01

MU1 x 4

¥216

¥104,472

Qwen-Plus-Character-2025-11-06

qwen-plus-character-2025-11-06

MU1 x 4

¥216

¥104,472

GLM

Model Name

Model Code

Model Unit Specification

Hourly unit price (CNY)

Minimum billing: minutes

Monthly subscription unit price (CNY)

Minimum billing: days

GLM-5.1

glm-5.1

MU2 x 8

¥504

¥240,288

MU3 x 16 (PD disaggregation mode)

PD disaggregation mode: ¥2,192

PD disaggregation mode: ¥1,055,504

MU6 x 16

¥400

¥193,424

GLM-5

glm-5

MU3 x 16 (PD disaggregation mode)

PD disaggregation mode: ¥2,192

PD disaggregation mode: ¥1,055,504

GLM-4.7

glm-4.7

MU6 x 32 (PD disaggregation mode)

PD disaggregation mode: ¥800

PD disaggregation mode: ¥386,848

DeepSeek

Model Name

Model Code

Model Unit Specification

Hourly unit price (CNY)

Minimum billing: minutes

Monthly subscription unit price (CNY)

Minimum billing: days

DeepSeek-v4-Flash

deepseek-v4-flash

MU1 x 8

¥432

¥208,944

MU3 x 8

¥1,096

¥527,752

DeepSeek-v3.2

deepseek-v3.2

MU2 x 16 (PD disaggregation mode)

PD disaggregation mode: ¥1,008

PD disaggregation mode: ¥480,576

More models

Model Name

Model Code

Model Unit Specification

Hourly unit price (CNY)

Minimum billing: minutes

Monthly subscription unit price (CNY)

Minimum billing: days

Kimi-K2.5

kimi-k2.5

MU2 x 8

¥504

¥240,288

Model type:

  • Instruct - The model performs inference in non-thinking mode after deployment.
  • Thinking - The model performs inference in thinking mode after deployment.

Model deployment type:

  • PD disaggregation mode - Reduces first-token latency and increases throughput.

    For models deployed in this mode, during inference the first-token computation (Prefill) and the subsequent token computation (Decode) are split across different compute nodes for execution.

Multimodal

Qwen-VL

Model Name

Model Code

Model Unit Specification

Hourly unit price (CNY)

Minimum billing: minutes

Monthly subscription unit price (CNY)

Minimum billing: days

Qwen3-VL-235B-A22B-Thinking

qwen3-vl-235b-a22b-thinking

MU1 x 8

¥432

¥208,944

MU2 x 8

¥504

¥240,288

MU3 x 8

¥1,096

¥527,752

Qwen3-VL-32B-Instruct

qwen3-vl-32b-instruct

MU2 x 8

¥504

¥240,288

MU3 x 8

¥1,096

¥527,752

Qwen3-VL-30B-A3B-Instruct

qwen3-vl-30b-a3b-instruct

MU1 x 4

¥216

¥104,472

Qwen3-VL-8B-Instruct

qwen3-vl-8b-instruct

MU1 x 2

¥108

¥52,236

MU5 x 1

¥21

¥10,139

Qwen3-VL-4B-Instruct

qwen3-vl-4b-instruct

MU1 x 2

¥108

¥52,236

Qwen3-VL-2B-Instruct

qwen3-vl-2b-instruct

MU5 x 1

¥21

¥10,139

Qwen3-VL-Embedding-2B

qwen3-vl-embedding-2b

MU5 x 1

¥21

¥10,139

Qwen3-VL-Flash-2025-10-15

qwen3-vl-flash-2025-10-15

MU1 x 4

¥216

¥104,472

Qwen3-VL-Plus-2025-09-23

qwen3-vl-plus-2025-09-23

MU1 x 4

¥216

¥104,472

Qwen-VL-Max-2025-08-13

qwen-vl-max-2025-08-13

MU6 x 4

¥100

¥48,356

Qwen Omni

Model Name

Model Code

Model Unit Specification

Hourly unit price (CNY)

Minimum billing: minutes

Monthly subscription unit price (CNY)

Minimum billing: days

Qwen3.5-Omni-Flash

qwen3.5-omni-flash

MU9 x 1

¥51

¥24,600

Model Type:

  • Instruct - The model performs inference in non-thinking mode after deployment.
  • Thinking - The model performs inference in thinking mode after deployment.
  • Instruct/Thinking - You can choose whether to enable thinking mode when deploying the model.

Speech Synthesis

CosyVoice

Model Name

Model Code

Model Unit Specification

Hourly Unit Price (CNY)

Monthly Subscription Unit Price (CNY)

cosyvoice-v3-flash

cosyvoice-v3-flash

MU5

¥21

¥10,139

Image generation models – Wan

Deployment is free. Invocations are billed at the standard rate of the fine-tuned base model. For the training workflow, see Fine-tune image generation models.

Model ID

LoRA Deployment &amp; Invocation Price

wan2.7-image-pro

CNY 0.50/image

wan2.7-image

CNY 0.20/image

Image generation models – Qwen

Deployment is free. Invocations are billed at the standard rate of the fine-tuned base model. For the training and deployment workflow, see Fine-tune image generation models.

Model ID

LoRA Deployment &amp; Invocation Price

qwen-image-2.0

CNY 0.20/image

qwen-image-2.0-pro

CNY 0.50/image

Speech synthesis models – CosyVoice

NoteCosyVoice model fine-tuning is available only in the China (Beijing) region.

After deployment, a fine-tuned CosyVoice model is billed by the usage duration of model units. The formula is: Fee = Usage duration (hours) × Number of model units × Model unit price.

Where Number of model units = Model units per replica × Number of replicas. For the available deployment templates and their model unit types, see Deployment templates in Fine-tune CosyVoice.

Model name

Model code

Model unit type

Hourly price (CNY)

Monthly price (CNY)

cosyvoice-v3-flash

cosyvoice-v3-flash

MU5

¥21

¥10,139

FAQ

Q: When does billing for model deployment start?

A: Billing starts when the model status changes to Running. No charges apply during Deploying, Overdue Payment, or Deployment Failed.

For monthly subscriptions, the billing period starts when the status changes to Running.

Q: Am I charged if I cancel a training job?

A: Yes. If you cancel training manually, you are charged for all tokens processed before cancellation. Training jobs interrupted by system errors or other non-user causes are not charged.

Q: How do I view invocation statistics for a deployed model?

A: Visit the Model Monitoring (Beijing), Model Monitoring (Virginia), or Model Monitoring (Singapore) page.

image