ACK 2025 release notes

Updated at:

Monthly feature releases and updates for Container Service for Kubernetes (ACK) and related products throughout 2025.

For the latest release notes, see Release notes.

December 2025

Product

Feature

Description

Region

Related documentation

Container Service for Kubernetes

Node pools support pay-as-you-go vulnerability remediation from Security Center

Enable OS CVE vulnerability remediation to scan nodes for vulnerabilities and apply fixes in the console. Before you use this feature, you must activate Security Center Ultimate or purchase pay-as-you-go vulnerability remediation.

All regions

Remediate OS CVE vulnerabilities in node pools

APIG Ingress support

APIG Ingress is a cloud-native API gateway built on open-source Higress, compatible with Nginx Ingress. Ideal for API management and microservice scenarios.

All regions

Manage APIG Ingress

Kagent support

Kagent is a framework for building and running AI applications on Kubernetes. It provides declarative APIs to create agents and MCP servers and integrate with multiple large language models.

All regions

Kagent

Manage tags for NAS, OSS, and CPFS with CNFS

CNFS supports tagging NAS, OSS, and CPFS resources for fine-grained classification and permission management.

All regions

Manage tags for NAS, OSS, and CPFS

Use Kyverno as a policy engine

Kyverno is a Kubernetes-native policy engine that enforces security, compliance, and automation policies with a Policy-as-Code approach. Unlike the default OPA Gatekeeper, Kyverno uses YAML instead of Rego and supports mutating and generating resources at admission. Ideal for customized policies, automated O&M, or multi-cluster policy governance.

All regions

Use Kyverno as a policy engine

Secure confidential container environments with remote attestation

PeerPod remote attestation verifies that confidential containers run in a genuine confidential computing environment, such as Intel TDX. It automatically verifies nodes before deployment and provides on-demand runtime attestations for end-to-end workload security.

All regions

Use remote attestation to secure confidential container environments

Deploy A2A protocol servers in Knative

Agent2Agent (A2A) is an open standard for communication between AI agents. Deploy an A2A server in Knative to leverage auto scaling (including scale to zero) for on-demand resource usage and rapid iteration.

All regions

Deploy A2A in Knative

ack-agent-gateway best practices

  • A2A traffic governance and authentication: Install the ack-agent-gateway extension based on the Gateway API to manage A2A protocol traffic for external-facing AI agent applications.

  • Build an MCP service gateway: Install the ack-agent-gateway extension based on the Gateway API to securely expose MCP services to external LLMs.

All regions

ACK One

FederatedHPA best practice based on vLLM custom metrics

Large language model (LLM) services with fluctuating traffic often use multi-cluster architectures. Deploy a vLLM inference service on an ACK One fleet and use FederatedHPA for cross-cluster elastic scaling.

All regions

FederatedHPA best practice based on vLLM custom metrics

November 2025

Product

Feature name

Description

Region

Related documentation

Container Service for Kubernetes

Support for deploying and running GPU workloads in smart managed mode

Enable smart managed mode to dynamically scale GPU resources with smart managed node pools, reducing costs for variable-demand scenarios such as online inference.

All

Deploy and run GPU workloads

New servicemesh-operator component

The servicemesh-operator simplifies deploying, upgrading, and configuring Service Mesh (ASM) in ACK clusters, enabling features such as traffic management, security, and observability.

All

servicemesh-operator

New built-in FinOps rule library

ACK cluster policy management now includes a FinOps rule library alongside Compliance, Infra, K8s-general, and PSP for validating Pod deployment and update requests.

All

Container security policy rule libraries

Support for deploying MCP Server on Knative

Host MCP Server on Knative to leverage Serverless benefits such as on-demand autoscaling and event-driven capabilities.

All

Deploy MCP Server on Knative

Best practices for configuring rolling updates and graceful shutdown

Configure readiness probes, readinessGates, preStop hooks, and SLB graceful shutdown for zero-downtime application updates and smooth traffic migration.

All

Implement zero-downtime rolling deployments

Distributed Cloud Container Platform ACK One

Best practices for multi-cluster priority-based elastic scheduling

Set cluster priorities in multi-cluster fleets to prioritize IDC or primary-region resources, with cloud or secondary-region resources as backup. Combined with inventory-aware scheduling, this ensures business continuity.

All

Multi-cluster priority-based elastic scheduling

October 2025

Product

Feature name

Description

Region

Related documentation

Container Service for Kubernetes

Best practices for OSS volume performance tuning

Troubleshoot and optimize OSS volume performance issues such as high latency or low throughput with systematic diagnostic steps.

All

Best practices for OSS volume performance tuning

Schedule GPUs with DRA

Deploy the NVIDIA DRA driver to overcome traditional device plugin limitations. The Kubernetes DRA API enables dynamic GPU allocation and fine-grained resource control between pods, improving utilization and reducing costs.

All

Schedule GPUs with DRA

Route external MaaS services with Gateway with Inference Extension

Use Gateway with Inference Extension to centralize API key and request path management when connecting to external MaaS services such as Bailian. Configure HTTPRoute rules to inject credentials and rewrite URLs automatically.

All

Route external MaaS services with Gateway with Inference Extension

ACK One

ACS GPU-HPN capacity reservation for registered clusters

Register on-premises clusters with ACK One to reserve GPU-HPN capacity for unified GPU scheduling across hybrid cloud environments, supporting AI training and inference workloads.

All

Example: Use ACS GPU HPN computing power in an ACK One registered cluster

Collect control plane metrics with a self-managed Prometheus

Install Metrics Aggregator and configure a ServiceMonitor to integrate ACK One registered cluster control plane metrics into your self-managed Prometheus for unified alerting.

All

Collect control plane metrics with a self-managed Prometheus

Cloud-native AI Suite

Submit eRDMA-accelerated PyTorch distributed training jobs with Arena

Submit PyTorch distributed training jobs with Arena and configure eRDMA network acceleration for low-latency, high-throughput inter-node communication, improving training efficiency.

All

Submit eRDMA-accelerated PyTorch distributed training jobs with Arena

September 2025

Product

Feature

Description

Region

Related documentation

Container Service for Kubernetes

Support for Kubernetes 1.34

ACK supports Kubernetes 1.34. Create new clusters or upgrade existing compatible clusters to this version.

All

Kubernetes 1.34

Support for hybrid cloud node pools

Create a hybrid cloud node pool in an ACK managed Pro cluster to manage on-premises and cloud resources from a single plane. Add existing hybrid cloud nodes for elastic scaling and cost optimization.

All

Create and manage a hybrid cloud node pool

Support for configuring DNS resolution for hybrid cloud node pools

Frequent CoreDNS queries through a leased line can overload the connection and cause DNS failures. Configure NodeLocal DNSCache to cache queries locally on each node.

All

Configure NodeLocal DNSCache for a hybrid cloud node pool

Support for the Terway-Hybrid network plugin

Standard CNI plugins cannot handle the complex topologies of hybrid cloud node pools. Terway-Hybrid ensures seamless Pod connectivity across your data center and the cloud.

All

Use the Terway-Hybrid network plugin

RRSA authentication for ossfs 2.0

Mount an OSS bucket as an ossfs 2.0 volume with a dynamic PV. RRSA authentication is recommended for auto-rotating credentials and Pod-level permission isolation in production and multi-tenant environments.

All

Use an ossfs 2.0 dynamic volume

Support for AI-powered content moderation

Enforce content compliance for generative AI services with Gateway API and Inference Extension. Configure the ACKTrafficFilter plugin to integrate with Content Moderation service for gateway-level content blocking.

All

Use Gateway with Inference Extension to implement AI-powered content moderation

ACK One (Alibaba Cloud Distributed Cloud Container Platform)

Support for integrating cloud-based GPU compute power

By providing unified scheduling and O&M for heterogeneous computing resources, ACK One registered clusters significantly improve resource utilization.

All

Integrate cloud-based GPU compute power

Migrate single-cluster applications to a fleet for multi-cluster distribution

Use AMC to deploy applications to multiple clusters, preventing configuration drift and enabling centralized management with automatic synchronization.

All

Migrate single-cluster applications to a fleet and distribute them to multiple clusters

August 2025

Product

Feature

Description

Region

Documentation

Container Service for Kubernetes

KV Cache-aware load balancing with intelligent inference routing

KV Cache-aware load balancing dynamically routes AI inference requests to the optimal compute node, improving large language model (LLM) service efficiency.

All regions

Use prefix cache-aware routing in precise mode

Support for custom CNI plugins

ACK supports Bring Your Own CNI (BYOCNI) mode, letting you install a custom CNI plugin when the default Terway and Flannel plugins don't meet your needs.

All regions

Use a custom CNI plugin in an ACK cluster

Managed policy governance for smart managed mode clusters

Enable the security policy management feature to meet compliance requirements and enhance cluster security. Security policy rules include Infra, Compliance, PSP, and K8s-general.

All regions

Enable security policy management

Knative support for ACS compute resources

Configure Knative Services to use ACS compute resources. Leverage diverse compute types and service levels to optimize costs for various workloads.

All regions

Use ACS resources

More flexible configurations for Gateway with Inference Extension

  • Customize inference extension configurations: Adjust routing policies through annotations or override deployment configurations with a ConfigMap.

  • Customize Gateway configurations: Modify the EnvoyProxy resource to adjust Service type, replica count, and resource allocation.

All regions

Configure inference routing for SGLang PD-separated services

PD separation decouples the prefill and decode phases onto different GPUs, eliminating resource contention, reducing time to first token (TPOT), and increasing throughput.

All regions

Configure inference routing for SGLang PD-separated services by using Gateway with Inference Extension

Securely deploy vLLM inference services in ACK heterogeneous confidential computing clusters

LLM inference involves sensitive data and model assets at risk of exposure in untrusted environments. ACK-CAI integrates hardware-based confidential computing (Intel TDX, GPU TEE) to provide end-to-end security for model inference.

All regions

Securely deploy vLLM inference services in ACK heterogeneous confidential computing clusters

Cloud-native AI Suite

Introducing the AI Serving Stack

The AI Serving Stack is an end-to-end ACK-based solution for cloud-native AI inference, covering deployment, intelligent routing, elastic scaling, and observability across the full LLM inference lifecycle.

All regions

AI Serving Stack

July 2025

Product

Feature

Description

Region

Related documentation

Container Service for Kubernetes

Access ECS instance metadata in enforced mode only

ACK cluster nodes now support enforced-mode-only (IMDSv2) access to ECS instance metadata, enhancing metadata service security.

All regions

Access ECS instance metadata in enforced mode only

Subscribe to images from international registries

Use artifact subscription in ACR Enterprise Edition to automatically synchronize images from international registries such as Docker Hub, GCR, and Quay.

All regions

Obtain images from international registries through artifact subscription

Mount NAS by using the EFC client through CNFS

EFC improves NAS performance with distributed caching, supporting high concurrency and parallel access to large datasets. Ideal for data-intensive workloads such as big data analytics and AI training. Compared to standard NFS, EFC accelerates file access and improves read/write performance.

All regions

Mount NAS by using the EFC client through CNFS

ACK One (Distributed Cloud Container Platform)

Console-based management for GitOps

Manage GitOps capabilities from the console, including feature toggles, ACLs, ApplicationSet UI, Argo CD ConfigMap, component restarts, and monitoring.

All regions

GitOps quick start

Argo CD ConfigMap configuration for multi-cluster GitOps

ACK One lets you manage GitOps-related features and permissions by configuring the Argo CD ConfigMap.

All regions

Configure the Argo CD ConfigMap

Inventory-aware elastic scheduling for multi-cluster fleets

ACK One provides an inventory-aware scheduler for multi-cluster fleets. When a cluster lacks resources, the scheduler deploys applications to clusters with available inventory, which then scale up nodes automatically.

All regions

Cross-region multi-cluster elastic scheduling based on inventory awareness

Container Service for Edge (ACK@Edge)

Configure a private connection for a leased line connection

ACK@Edge clusters can connect to the cloud over a leased line, enabling secure access to ACK and ACR while resolving network conflicts and fixed IP address issues.

All regions

Configure a private connection for a leased line connection

June 2025

Product

Feature name

Description

Region

Related documentation

Container Service for Kubernetes

AI profiling

AI Profiling uses eBPF and dynamic process injection for non-invasive GPU performance analysis in Kubernetes. Attach and detach data collection on running workloads without code changes for real-time production analysis.

All regions

AI Profiling

GPU node auto-healing

Node auto-healing now repairs instances affected by GPU hardware and software failures.

ACK automates the full lifecycle for EGS and Lingjun node failures: detection, alerting, isolation, drain, and repair. Optional user authorization before repairs adds oversight while reducing O&M costs.

All regions

Enable Node Auto-healing

CPFS for AI static volumes

CPFS for AI delivers ultra-high throughput and IOPS with end-to-end RDMA acceleration, ideal for AIGC and autonomous driving. Create static volumes in your cluster for these workloads.

All regions

Use CPFS for AI Static Volumes

ACK VPD CNI component

ACK VPD CNI is the container network plugin for Lingjun nodes in ACK Pro clusters, allocating and managing network resources for Lingjun connections.

All regions

ACK VPD CNI

ack-kms-agent-webhook-injector component

The ack-kms-agent-webhook-injector injects a KMS Agent sidecar into Pods. Applications fetch and cache credentials from KMS through a local HTTP interface, eliminating hard-coded secrets.

All regions

Import Alibaba Cloud KMS Service Credentials for Applications

Expanded capabilities for Gateway with Inference Extension

Gateway with Inference Extension now supports vLLM and SGLang with canary releases, inference load balancing, model name-based routing, rate limiting, and circuit breaking.

All regions

Gateway with Inference Extension Traffic Management and Inference Service Management

CAA solution for confidential containers on confidential virtual machines

Deploy confidential computing workloads in ACK clusters with the CAA solution. Using Intel® TDX, it protects sensitive data from external and cloud-provider threats for compliance scenarios such as financial risk control and healthcare.

All regions

Implement CAA Confidential Container Solution Based on Confidential VMs

Cloud-native AI Suite

Schedule Dify workflows with XXL-JOB

Dify lacks a built-in scheduler for automating tasks such as risk monitoring and data analysis. Integrate XXL-JOB to schedule and monitor Dify workflow applications.

All regions

Schedule Dify Workflow Applications via XXL-JOB

May 2025

Product

Feature

Description

Region

Related documentation

Container Service for Kubernetes

Support for Kubernetes 1.33

ACK supports Kubernetes 1.33. Create new clusters or upgrade existing ones to this version.

All regions

Kubernetes 1.33

Default installation of ack-ram-authenticator component

Starting with Kubernetes 1.33, new ACK managed clusters automatically install the latest version of the managed ack-ram-authenticator component, without consuming additional cluster node resources.

All regions

Product announcement: ack-ram-authenticator installed by default on ACK managed clusters starting with Kubernetes 1.33

containerd 2.1.1 is available

containerd 2.1.1 introduces new features, such as the Node Resource Interface (NRI), Container Device Interface (CDI), and Sandbox API.

All regions

containerd runtime release notes

Support for ossfs 2.0

ossfs 2.0 is a FUSE-based client that mounts OSS buckets as local file systems with POSIX operations. Compared to ossfs 1.0, it improves sequential read/write performance and concurrent small-file throughput, ideal for AI training, big data, and autonomous driving.

All regions

ossfs 2.0

ACK One

Use ApplicationSet to coordinate multi-environment deployments and application dependencies

Combine Argo CD Progressive Syncs with ApplicationSet to build an automated deployment system that manages application dependencies between development and pre-production environments.

All regions

Use ApplicationSet to coordinate multi-environment deployments and application dependencies

April 2025

Product

Feature

Description

Release region

Related documentation

Container Service for Kubernetes

Create and manage Lingjun node pools

You can create and manage a Lingjun node pool in an ACK managed cluster Pro.

All regions

Lingjun node pools

Configure a node pool by specifying instance attributes

Configure a node pool by specifying instance attributes such as vCPUs and memory. The node pool automatically selects suitable instance types during scale-out, improving scaling success rates.

All regions

Configure a node pool by specifying instance attributes

Real-time AI Profiling

AI Profiling uses eBPF and dynamic process injection for non-intrusive GPU diagnostics in Kubernetes. Attach and detach the profiler on live services without code changes.

All regions

Use AI Profiling from the command line

Enable preemption

When resources are tight, enable preemption so ACK Scheduler evicts low-priority Pods to free resources for high-priority tasks.

All regions

Enable preemption

Access services through Gateway with Inference Extension

The Gateway with Inference Extension component is built on the Envoy Gateway project. It supports all the basic capabilities of the Gateway API and the extended resources of the open-source Envoy Gateway.

All regions

Access services through Gateway with Inference Extension

Generative AI service enhancements

Use Gateway with Inference Extension for intelligent routing, traffic management, canary releases, circuit breaking, and traffic mirroring for AI inference services.

All regions

Generative AI service enhancements

Back up and restore volumes from PVC to PVC

Back up and restore cloud disk data within or across ACK clusters and regions. Restore to new PVCs that can be mounted directly without modifying workload configurations.

All regions

Backup Center

alibabacloud-privateca-issuer released

AlibabaCloud Private CA Issuer is now available. It lets you create and manage Alibaba Cloud PCA certificates in your cluster by using cert-manager. The issuer is available in the ACK App Market.

All regions

None

Deploy a workload and implement load balancing in an ACK managed cluster (smart managed mode)

Deploy a workload in an ACK managed cluster (smart managed mode) and expose it to the internet with an ALB Ingress for domain-based access and load balancing.

All regions

Deploy a workload and implement load balancing

Datapath V2 best practices

Optimize network configuration after enabling Datapath V2 with the Terway plugin, including Conntrack parameters and Identity resource management.

All regions

Datapath V2 best practices

Dify component upgrade guide

Upgrade ack-dify from an earlier version to v1.0.0 or later. The process includes backing up data, installing the plugin migration tool, and enabling the new plugin ecosystem.

All regions

Upgrade the Dify component in an ACK cluster

ACK One

Use PrivateLink to resolve IP conflicts in a data center network

When Serverless computing resources conflict with data center CIDR blocks in a leased-line-connected ACK One registered cluster, use PrivateLink to resolve IP conflicts.

All regions

Use PrivateLink to resolve IP conflicts in a data center network

Schedule ACS Pods across regions

ACK One registered cluster integrates Serverless computing resources from multiple regions for dynamic cross-region GPU scheduling and unified management.

All regions

Schedule ACS Pods across regions

Log collection

You can configure log collection by using SLS CRDs or environment variables to automatically collect container logs with Alibaba Cloud Log Service (SLS).

All regions

Build a multi-cluster CD system

Combine Cloud Efficiency CD with ACK One Application Distribution to build a multi-cluster continuous delivery system.

All regions

Build a multi-cluster CD system by using ACK One and Cloud Efficiency

ACK Edge

Version 1.32 released

Version 1.32 is now supported. Features include optimizing requests from CoreDNS, kube-proxy, and kubelet to the kube-apiserver, reducing cloud-to-edge communication traffic, and more.

All regions

ACK Edge Kubernetes 1.32 release notes

Network element configuration in a leased line environment

Connect on-premises servers to ACK over the internet or a leased line. For leased line connections, configure network elements first.

All regions

Configure network elements in a leased line environment

AI Engineering Suite

HistoryServer component support

The native Ray Dashboard is only available while a cluster runs. RayCluster HistoryServer collects logs in real time and persists them to OSS for access after cluster termination.

All regions

Install the HistoryServer component in ACK

KubeRay component support

Deploy the KubeRay Operator and integrate with SLS and Prometheus for enhanced log management, observability, and high availability.

All regions

Install the KubeRay component in ACK

March 2025

Product

Feature

Description

Region

Related documentation

Container Service for Kubernetes

ACK Pro managed clusters support intelligent hosting mode

When you create an ACK managed cluster, you can enable intelligent hosting mode to quickly provision a Kubernetes cluster that follows best practices.

After the cluster is created, ACK automatically provisions an intelligent managed node pool. This node pool dynamically scales based on workload demand. ACK also handles all operational tasks for this node pool, including OS version upgrades, software updates, and security patching.

All regions

Enable tracing for control plane and data plane components

After you enable tracing for the cluster API Server or kubelet, trace data automatically flows to Managed Service for OpenTelemetry. This integration provides detailed trace visualizations, real-time topology maps, and other monitoring data.

All regions

High-risk KubeConfig SMS and email notifications

Receive SMS and email alerts for high-risk KubeConfig files that still pose security risks after deletion.

All regions

None

Intelligent routing and traffic management with ACK Gateway with Inference Extension

Use ACK Gateway with Inference Extension for intelligent routing and traffic management for inference services.

All regions

Implement intelligent routing and traffic management with ACK Gateway with Inference Extension

Deploy vLLM inference applications on Knative

Traditional GPU utilization-based autoscaling doesn't reflect LLM inference load accurately. KPA adjusts resources based on QPS or RPS for more precise scaling.

All regions

Deploy vLLM inference applications on Knative

ACK One (Distributed Cloud Container Platform)

Unified component management for multi-cluster fleets

Define component baselines with specific versions and deploy them to multiple clusters. Supports component configuration, batch deployment, and rollbacks.

All regions

Multi-cluster component management

Dynamic distribution and rescheduling

Use PropagationPolicy to distribute replicas across clusters based on available resources. The rescheduler checks every two minutes and triggers rescheduling for Pods unschedulable for over 30 seconds.

All regions

Dynamic distribution and rescheduling

Cloud-native AI Suite

Set Slurm queue priorities

Configure Slurm queue policies for optimal scheduling and performance when jobs are submitted or change state.

All regions

Set Slurm queue priorities in an ACK cluster

February 2025

Product

Feature

Description

Region

Related documentation

Container Service for Kubernetes

Support for modifying control plane security groups and time zones

Modify the control plane security group and cluster time zone on the Basic Information page when initial settings no longer meet your requirements.

All

View cluster information

Node pools support custom containerd configurations

Customize containerd parameters for node pool nodes, such as configuring mirror repositories or bypassing certificate verification for specific registries.

All

Customize containerd parameter configurations for a node pool

Elasticity strength indicator for node pools

The elasticity strength indicator assesses node pool configuration availability and instance supply health, with recommendations to improve scale-out success rates.

All

View the elasticity strength of a node pool

Support for batch task orchestration

Argo Workflows is a Kubernetes-native engine that orchestrates parallel tasks with YAML or Python for CI/CD, data processing, and machine learning. Install the Argo Workflows component and manage tasks via the Alibaba Cloud Argo CLI or console.

All

Enable batch task orchestration

GPU fault detection

The ack-node-problem-detector component enhances the open-source node-problem-detector with GPU-specific fault detection checks. When a fault is detected, it generates a corresponding Kubernetes Event or Kubernetes Node Condition based on the fault type.

All

GPU fault detection and automatic isolation

Accelerate pod startup in Knative services using Fluid

Fluid is a Kubernetes-native dataset orchestration engine for data-intensive applications. Use Fluid in Knative to accelerate model inference pod startup.

All

Accelerate pod startup by using Fluid

ACK One

Schedule and distribute multi-cluster Spark jobs based on actual remaining resources

Use an ACK One fleet and ACK Koordinator to schedule multi-cluster Spark jobs based on actual remaining resources rather than requested resources, maximizing idle utilization while prioritizing online workload stability.

All

Schedule and distribute multi-cluster Spark jobs based on actual remaining resources

Build a DeepSeek distilled model inference service in an ACK One registered cluster using ACS GPU compute power

Connect a Kubernetes cluster from an on-premises data center to an ACK One registered cluster to seamlessly scale out compute power. You can then use ACS GPU resources to deploy the DeepSeek distilled inference service efficiently.

All

Build a DeepSeek distilled model inference service in an ACK One registered cluster by using ACS GPU compute power

ACK Edge

Support for adding pod vSwitches

Add a pod vSwitch to an ACK Edge cluster using Terway Edge to expand available IP addresses when the existing vSwitch is exhausted or the pod CIDR block needs expansion.

All

Add a pod vSwitch

Deploy the DeepSeek-R1 model

Use an ACK Edge cluster to manage on-premises GPUs and add cloud ACS Serverless GPUs via virtual nodes. The cluster runs tasks on on-premises GPUs first and automatically scales to cloud GPUs when needed.

All

Deploy a DeepSeek distilled model inference service

GPU resource monitoring

Integrate Prometheus monitoring with ACK Edge cluster to provide on-premises and edge GPU nodes the same observability as cloud resources.

All

Best practices for monitoring GPU resources of an ACK Edge cluster

Cloud Native AI Suite

Deploy a DeepSeek distilled model inference service on ACK

Deploy a production-ready DeepSeek distilled model inference service on ACK with KServe, using DeepSeek-R1-Distill-Qwen-7B as an example.

All

Deploy a DeepSeek distilled model inference service on ACK

Tutorial: Deploy a full-parameter DeepSeek inference service on ACK using multi-machine distributed deployment

Deploy DeepSeek-R1-671B across two nodes on ACK with a hybrid parallel strategy and Arena. Integrate the DeepSeek-R1 service with Dify to build an enterprise Q&A system with long-context understanding.

All

Tutorial: Deploy a DeepSeek full-parameter inference service by using multi-machine distributed deployment on ACK

January 2025

Product

Feature

Description

Region

Related documentation

Container Service for Kubernetes

On-demand image acceleration for node pools

ACK supports on-demand container image loading with DADI, eliminating full downloads and decompressing data on the fly to reduce startup time.

All Regions

Accelerate container startup by using on-demand image loading

Support for Alibaba Cloud Linux 3 Container Optimized Edition

Alibaba Cloud Linux 3 Container Optimized Edition is optimized for containerized environments, delivering higher deployment density, faster startup, and enhanced security isolation compared to the standard image.

All Regions

Support for Kubernetes 1.32

ACK supports Kubernetes 1.32. Create new clusters or upgrade existing ones to this version.

All Regions

(End of support) Kubernetes 1.32

Improve resource utilization with ElasticQuotaTree and task queues

ack-kube-queue, ElasticQuotaTree, and ack-scheduler enable fair and isolated resource allocation, allowing different teams and tasks to share compute resources within a cluster.

All Regions

Improve resource utilization with ElasticQuotaTree and task queues

Best practice: Fine-grained resource control with resource groups

Organize ACK resources into resource groups by department, project, or environment. Combine with RAM for resource isolation and fine-grained permission management.

All Regions

Use resource groups for fine-grained resource control

ACK One

Connect ACK One registered clusters to ACS compute power

An ACK One registered cluster can use container compute power from ACS.

All Regions

Schedule pods to run on ACS by using virtual nodes

Cross-cluster service access using native service domain names

ACK One uses MultiClusterService for cross-cluster access through native service domain names, without modifying application code, Pod DNS, or CoreDNS settings.

All Regions

Access services across clusters by using native service domain names

Access multi-cluster resources using the Go SDK

Use the Go SDK to integrate an ACK One fleet into your platform and access member cluster resources.

All Regions

Access multi-cluster resources by using the Go SDK

ACK Edge

Cloud node scaling

When on-premises node resources are insufficient, the auto-scaling feature scales out cloud-based nodes for your ACK Edge cluster to increase scheduling capacity.

All Regions

Cloud ECS node elasticity

Deploy elastic inference services for LLMs in a hybrid cloud

Install ack-kserve and use ACK Edge clusters cloud elasticity to deploy elastic LLM inference services in hybrid cloud, flexibly scheduling on-premises and cloud resources.

All Regions

GPU sharing and scheduling

GPU sharing allows multiple pods to share the compute resources of a single GPU card. This improves GPU utilization and reduces costs.

  • Cloud nodes in an ACK Edge cluster support full GPU sharing, GPU memory isolation, and computing power isolation.

  • Edge node pools in an ACK Edge cluster support only GPU sharing. GPU memory isolation and computing power isolation are not supported.

All Regions

Use GPU sharing and scheduling

Centrally manage ECS resources across regions

Use an ACK Edge cluster to centrally manage compute resources across regions for full lifecycle management and efficient scheduling.

All Regions

Centrally manage ECS resources across regions