inference-nv-pytorch 25.12

更新时间:
复制 MD 格式

This topic describes the release notes for inference-nv-pytorch 25.12.

Main features and bug fixes

Main features

  • This release includes images for two CUDA versions:

    • The CUDA 12.8 image supports only the amd64 architecture.

    • The CUDA 13.0 image supports both the amd64 and aarch64 architectures.

  • The vLLM images now use PyTorch 2.9.0, and the SGLang images use PyTorch 2.9.1.

  • In the CUDA 12.8 image, deepgpu-comfyui is now 1.3.2, and the deepgpu-torch optimization component is 0.1.12+torch2.9.0cu128.

  • In both the CUDA 12.8 and CUDA 13.0 images, vLLM is now v0.12.0, and SGLang is v0.5.6.post2.

Bug fixes

None.

Contents

Image name

inference-nv-pytorch

Tag

25.12-vllm0.12.0-pytorch2.9-cu128-20251215-serverless

25.12-sglang0.5.6.post2-pytorch2.9-cu128-20251215-serverless

25.12-vllm0.12.0-pytorch2.9-cu130-20251215-serverless

25.12-sglang0.5.6.post2-pytorch2.9-cu130-20251215-serverless

Supported architecture

amd64

amd64

amd64

aarch64

amd64

aarch64

Use case

large model inference

large model inference

large model inference

large model inference

large model inference

large model inference

Framework

pytorch

pytorch

pytorch

pytorch

pytorch

pytorch

Requirements

NVIDIA Driver release >= 570

NVIDIA Driver release >= 570

NVIDIA Driver release >= 580

NVIDIA Driver release >= 580

NVIDIA Driver release >= 580

NVIDIA Driver release >= 580

System components

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.9.0+cu128

  • CUDA 12.8

  • diffusers 0.36.0

  • deepgpu-comfyui 1.3.2

  • deepgpu-torch 0.1.12+torch2.9.0cu128

  • flash_attn 2.8.3

  • flashinfer-python 0.5.3

  • imageio 2.37.2

  • imageio-ffmpeg 0.6.0

  • ray 2.52.1

  • transformers 4.57.3

  • triton 3.5.0

  • torchaudio 2.9.0+cu128

  • torchvision 0.24.0+cu128

  • vllm 0.12.0

  • xfuser 0.4.5

  • xgrammar 0.1.27

  • ljperf 0.1.0+477686c5

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.9.1+cu128

  • CUDA 12.8

  • diffusers 0.36.0

  • decord 0.6.0

  • decord2 2.0.0

  • deepgpu-comfyui 1.3.2

  • deepgpu-torch 0.1.12+torch2.9.0cu128

  • flash_attn 2.8.3

  • flash_mla 1.0.0+1408756

  • flashinfer-python 0.5.3

  • imageio 2.37.2

  • imageio-ffmpeg 0.6.0

  • ray 2.52.1

  • transformers 4.57.1

  • sgl-kernel 0.3.19

  • sglang 0.5.6.post2

  • xgrammar 0.1.27

  • triton 3.5.1

  • torchao 0.9.0

  • torchaudio 2.9.1

  • torchvision 0.24.1

  • xfuser 0.4.5

  • ljperf 0.1.0+477686c5

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.9.0+cu130

  • CUDA 13.0.2

  • diffusers 0.36.0

  • flash_attn 2.8.3

  • flashinfer-python 0.5.3

  • imageio 2.37.2

  • imageio-ffmpeg 0.6.0

  • ray 2.52.1

  • transformers 4.57.3

  • triton 3.5.0

  • torchaudio 2.9.0+cu130

  • torchvision 0.24.0+cu130

  • vllm 0.12.0

  • xfuser 0.4.5

  • xgrammar 0.1.27

  • ljperf 0.1.0+d0e4a408

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.9.0+cu130

  • CUDA 13.0.2

  • diffusers 0.36.0

  • flash_attn 2.8.3

  • flashinfer-python 0.5.3

  • transformers 4.57.1

  • ray 2.53.0

  • vllm 0.12.0

  • triton 3.5.0

  • torchaudio 2.9.0

  • torchvision 0.24.0

  • xfuser 0.4.5

  • xgrammar 0.1.27

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.9.1+cu130

  • CUDA 13.0.2

  • diffusers 0.36.0

  • decord 0.6.0

  • decord2 2.0.0

  • flash_attn 2.8.3

  • flashinfer-python 0.5.3

  • imageio 2.37.2

  • imageio-ffmpeg 0.6.0

  • ray 2.52.1

  • transformers 4.57.3

  • sgl-kernel 0.3.19

  • sglang 0.5.6.post2

  • xgrammar 0.1.27

  • triton 3.5.1

  • torchao 0.9.0

  • torchaudio 2.9.1

  • torchvision 0.24.1+cu130

  • xfuser 0.4.5

  • ljperf 0.1.0+d0e4a408

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.9.1+cu130

  • CUDA 13.0.2

  • diffusers 0.36.0

  • decord2 2.0.0

  • flash_attn 2.8.3

  • flashinfer-python 0.5.3

  • imageio 2.37.2

  • imageio-ffmpeg 0.6.0

  • transformers 4.57.1

  • sgl-kernel 0.3.19

  • sglang 0.5.6.post2

  • xgrammar 0.1.27

  • triton 3.5.1

  • torchao 0.9.0

  • torchaudio 2.9.1

  • torchvision 0.24.1

  • xfuser 0.4.5

Assets

Public images

CUDA 12.8 assets

  • egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:25.12-vllm0.12.0-pytorch2.9-cu128-20251215-serverless

  • egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:25.12-sglang0.5.6.post2-pytorch2.9-cu128-20251215-serverless

CUDA 13.0 assets

  • egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:25.12-vllm0.12.0-pytorch2.9-cu130-20251215-serverless

  • egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:25.12-sglang0.5.6.post2-pytorch2.9-cu130-20251215-serverless

VPC images

To pull ACS AI container images quickly from within a Virtual Private Cloud (VPC), replace the specified public asset URI egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/{image:tag} with acs-registry-vpc.{region-id}.cr.aliyuncs.com/egslingjun/{image:tag}.

  • {region-id}: The region ID of the region where your ACS product is available. Examples: cn-beijing and cn-wulanchabu.

  • {image:tag}: The name and tag of the AI container image. Examples: inference-nv-pytorch:25.10-vllm0.11.0-pytorch2.8-cu128-20251028-serverless and training-nv-pytorch:25.10-serverless.

Note

These images are suitable for ACS and multi-tenant Lingjun environments. They are not compatible with single-tenant Lingjun environments.

Driver requirements

  • CUDA 12.8: NVIDIA Driver release >= 570

  • CUDA 13.0: NVIDIA Driver release >= 580

Quick start

The following example shows how to pull the inference-nv-pytorch image by using Docker and test the inference service with the Qwen2.5-7B-Instruct model.

Note

To use the inference-nv-pytorch image in ACS, select it from the Artifact Center page when you create a workload in the console, or specify the image reference in a YAML file. For more information about deploying model inference services with ACS GPU compute, see the following topics:

  1. Pull the inference container image.

    docker pull egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:[tag]
  2. Download the open source model from ModelScope.

    pip install modelscope
    cd /mnt
    modelscope download --model Qwen/Qwen2.5-7B-Instruct --local_dir ./Qwen2.5-7B-Instruct
  3. Run the following command to start and enter the container.

    docker run -it --rm --gpus all --network=host --privileged --init --ipc=host \
    --ulimit memlock=-1 --ulimit stack=67108864  \
    -v /mnt/:/mnt/ \
    egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:[tag]
  4. Run an inference test to verify the vLLM conversational inference feature.

    1. Start the server.

      python3 -m vllm.entrypoints.openai.api_server \
      --model /mnt/Qwen2.5-7B-Instruct \
      --trust-remote-code --disable-custom-all-reduce \
      --tensor-parallel-size 1
    2. Run the test on the client.

      curl http://localhost:8000/v1/chat/completions \
          -H "Content-Type: application/json" \
          -d '{
          "model": "/mnt/Qwen2.5-7B-Instruct",  
          "messages": [
          {"role": "system", "content": "You are a friendly AI assistant."},
          {"role": "user", "content": "Tell me about deep learning."}
          ]}'

      Output:

      {
        "id": "chat-d3c28759793d4376a65bfc4e40b59a71",
        "object": "chat.completion",
        "created": 1735278194,
        "model": "/mnt/Qwen2.5-7B-Instruct",
        "choices": [
          {
            "index": 0,
            "message": {
              "role": "assistant",
              "content": "Deep learning is a branch of machine learning inspired by the structure and function of the human brain's neural networks. It uses deep neural networks to process and analyze large amounts of data to identify effective predictive models. This technology has achieved remarkable success in various fields, including image recognition, speech recognition, and natural language processing.\n\nIn deep learning, a neural network consists of multiple layers: an input layer, one or more hidden layers, and an output layer. Each layer contains numerous nodes (or neurons) connected to nodes in other layers through weighted links. During training, the network adjusts the weights of these connections based on input data to minimize the error between the predicted and actual outputs. This process is typically accomplished using optimization algorithms like gradient descent.\n\nTraining deep learning models requires substantial computational resources and vast datasets. In recent years, advancements in computing hardware (such as GPUs and TPUs) and the exponential growth of available data have fueled the widespread application and development of deep learning. In addition to the applications mentioned, it is also extensively used in fields such as medical diagnosis, autonomous driving, gaming, and finance."
            },
            "tool_calls": [],
            "logprobs": null,
            "finish_reason": "stop",
            "stop_reason": null
          }
        ],
        "usage": {
          "prompt_tokens": 237,
          "completion_tokens": 213
        }
      }

      For more information about using vLLM, see the vLLM documentation.

Known issues

  • The deepgpu-comfyui plugin, which accelerates Wanx model video generation, currently supports only the GN8IS, G49E, and G59 instance types.