inference-nv-pytorch 26.02

更新时间:
复制 MD 格式

This document provides the release notes for inference-nv-pytorch 26.02.

Main features and bug fixes

Main features

  • This release provides images for CUDA 12.8 and CUDA 13.0:

    • The CUDA 12.8 images support only the amd64 architecture.

    • The CUDA 13.0 images support both the amd64 and aarch64 architectures.

  • In the vLLM images, Torch is upgraded to 2.10.0 and vLLM is upgraded to v0.15.1.

  • In the SGLang images, Torch is upgraded to 2.10.0 and SGLang is upgraded to v0.5.9.

Bug fixes

None.

Contents

Image

inference-nv-pytorch

Tag

26.02-vllm0.15.1-pytorch2.10-cu128-20260211-serverless

26.02-sglang0.5.9-pytorch2.10-cu128-20260227-serverless

26.02-vllm0.15.1-pytorch2.10-cu130-20260211-serverless

26.02-sglang0.5.9-pytorch2.10-cu130-20260227-serverless

Architecture

amd64

amd64

amd64

aarch64

amd64

aarch64

Use case

large model inference

large model inference

large model inference

large model inference

large model inference

large model inference

Framework

pytorch

pytorch

pytorch

pytorch

pytorch

pytorch

Requirements

NVIDIA Driver release >= 570

NVIDIA Driver release >= 570

NVIDIA Driver release >= 580

NVIDIA Driver release >= 580

NVIDIA Driver release >= 580

NVIDIA Driver release >= 580

System components

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.10.0

  • CUDA 12.8

  • NCCL 2.28.9

  • diffusers 0.36.0

  • flash_attn 2.8.3

  • flash_attn_3 3.0.0

  • flashinfer-python 0.6.1

  • imageio 2.37.2

  • imageio-ffmpeg 0.6.0

  • ray 2.53.0

  • transformers 4.57.6

  • triton 3.6.0

  • torchaudio 2.10.0

  • torchvision 0.25.0

  • vllm 0.15.1

  • xfuser 0.4.5

  • xgrammar 0.1.27

  • ljperf 0.1.0+477686c5

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.10.0

  • CUDA 12.8

  • NCCL 2.28.9

  • torchaudio 2.10.0

  • torchvision 0.25.0

  • diffusers 0.36.0

  • decord 0.6.0

  • decord2 3.0.0

  • flash_attn 2.8.3

  • flash_attn_3 3.0.0

  • flashinfer-python 0.6.3

  • imageio 2.37.2

  • imageio-ffmpeg 0.6.0

  • ray 2.54.0

  • transformers 4.57.1

  • sgl-kernel 0.3.21

  • sglang 0.5.9

  • xgrammar 0.1.27

  • triton 3.6.0

  • torchao 0.9.0

  • xfuser 0.4.5

  • ljperf 0.1.0+477686c5

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.10.0+cu130

  • CUDA 13.0.2

  • NCCL 2.28.9

  • diffusers 0.36.0

  • flash_attn 2.8.3

  • flash_attn_3 3.0.0

  • flashinfer-python 0.6.1

  • imageio 2.37.2

  • imageio-ffmpeg 0.6.0

  • ray 2.53.0

  • transformers 4.57.6

  • triton 3.6.0

  • torchaudio 2.10.0+cu130

  • torchvision 0.25.0+cu130

  • vllm 0.15.1

  • xfuser 0.4.5

  • xgrammar 0.1.27

  • ljperf 0.1.0+d0e4a408

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.10.0+cu130

  • CUDA 13.0.2

  • NCCL 2.28.9

  • diffusers 0.36.0

  • flash_attn 2.8.3

  • flashinfer-python 0.6.1

  • transformers 4.57.6

  • ray 2.53.0

  • vllm 0.15.1

  • triton 3.6.0

  • torchaudio 2.10.0+cu130

  • torchvision 2.10.0+cu130

  • xfuser 0.4.5

  • xgrammar 0.1.27

  • ljperf 0.1.0+477686c5

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.10.0+cu130

  • CUDA 13.0.2

  • NCCL 2.28.9

  • diffusers 0.36.0

  • decord 0.6.0

  • decord2 3.0.0

  • flash_attn 2.8.3

  • flashinfer-python 0.6.3

  • imageio 2.37.2

  • imageio-ffmpeg 0.6.0

  • ray 2.54.0

  • transformers 4.57.1

  • sgl-kernel 0.3.21

  • sglang 0.5.9

  • xgrammar 0.1.27

  • triton 3.6.0

  • torchao 0.9.0

  • torchaudio 2.10.0+cu130

  • torchvision 0.25.0+cu130

  • xfuser 0.4.5

  • ljperf 0.1.0+d0e4a408

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.10.0+cu130

  • CUDA 13.0.2

  • NCCL 2.28.9

  • diffusers 0.36.0

  • decord2 3.0.0

  • flash_attn 2.8.3

  • flashinfer-python 0.6.3

  • imageio 2.37.2

  • imageio-ffmpeg 0.6.0

  • ray 2.54.0

  • transformers 4.57.1

  • sgl-kernel 0.3.21

  • sglang 0.5.9

  • xgrammar 0.1.27

  • triton 3.6.0

  • torchao 0.9.0

  • torchaudio 2.10.0+cu130

  • torchvision 0.25.0+cu130

  • xfuser 0.4.5

  • ljperf 0.1.0+477686c5

Assets

Public images

CUDA 12.8 assets

  • egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:26.02-vllm0.15.1-pytorch2.10-cu128-20260211-serverless

  • egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:26.02-sglang0.5.9-pytorch2.10-cu128-20260227-serverless

CUDA 13.0 assets

  • egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:26.02-vllm0.15.1-pytorch2.10-cu130-20260211-serverless

  • egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:26.02-sglang0.5.9-pytorch2.10-cu130-20260227-serverless

VPC images

To quickly pull ACS AI container images from within a VPC, replace the public asset URI egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/{image:tag} with acs-registry-vpc.{region-id}.cr.aliyuncs.com/egslingjun/{image:tag}.

  • {region-id}: The ID of a supported ACS region. Examples: cn-beijing and cn-wulanchabu.

  • {image:tag}: The name and tag of the AI container image. Examples: inference-nv-pytorch:25.10-vllm0.11.0-pytorch2.8-cu128-20251028-serverless and training-nv-pytorch:25.10-serverless.

Note

These images are for ACS and the Lingjun multi-tenant offering. Do not use them in Lingjun single-tenant scenarios.

Driver requirements

  • CUDA 12.8: NVIDIA Driver release >= 570

  • CUDA 13.0: NVIDIA Driver release >= 580

Quick start

The following example shows how to pull the inference-nv-pytorch image by using Docker and test the inference service with the Qwen2.5-7B-Instruct model.

Note

To use the inference-nv-pytorch image in ACS, select it from the Artifacts Center on the Create Workload page in the console, or specify the image reference in a YAML file. For more information about building a model inference service by using ACS GPU compute, see the following topics:

  1. Pull the inference container image.

    docker pull egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:[tag]
  2. Use ModelScope to download the open-source model.

    pip install modelscope
    cd /mnt
    modelscope download --model Qwen/Qwen2.5-7B-Instruct --local_dir ./Qwen2.5-7B-Instruct
  3. Run the following command to enter the container.

    docker run -it --rm --gpus all --network=host --privileged --init --ipc=host \
    --ulimit memlock=-1 --ulimit stack=67108864  \
    -v /mnt/:/mnt/ \
    egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:[tag]
  4. Test the vLLM chat feature.

    1. Start the server.

      python3 -m vllm.entrypoints.openai.api_server \
      --model /mnt/Qwen2.5-7B-Instruct \
      --trust-remote-code --disable-custom-all-reduce \
      --tensor-parallel-size 1
    2. Run a client-side test.

      curl http://localhost:8000/v1/chat/completions \
          -H "Content-Type: application/json" \
          -d '{
          "model": "/mnt/Qwen2.5-7B-Instruct",  
          "messages": [
          {"role": "system", "content": "You are a friendly AI assistant."},
          {"role": "user", "content": "Introduce deep learning."}
          ]}'

      Output:

      image.png

      For more information, see the vLLM documentation.

Known issues

  • In this release, these images do not support the deepgpu-comfyui plugin.