ACK 2025 release notes
Monthly feature releases and updates for Container Service for Kubernetes (ACK) and related products throughout 2025.
For the latest release notes, see Release notes.
December 2025
Product | Feature | Description | Region | Related documentation |
Container Service for Kubernetes | Node pools support pay-as-you-go vulnerability remediation from Security Center | Enable OS CVE vulnerability remediation to scan nodes for vulnerabilities and apply fixes in the console. Before you use this feature, you must activate Security Center Ultimate or purchase pay-as-you-go vulnerability remediation. | All regions | |
APIG Ingress support | APIG Ingress is a cloud-native API gateway built on open-source Higress, compatible with Nginx Ingress. Ideal for API management and microservice scenarios. | All regions | ||
Kagent support | Kagent is a framework for building and running AI applications on Kubernetes. It provides declarative APIs to create agents and MCP servers and integrate with multiple large language models. | All regions | ||
Manage tags for NAS, OSS, and CPFS with CNFS | CNFS supports tagging NAS, OSS, and CPFS resources for fine-grained classification and permission management. | All regions | ||
Use Kyverno as a policy engine | Kyverno is a Kubernetes-native policy engine that enforces security, compliance, and automation policies with a Policy-as-Code approach. Unlike the default OPA Gatekeeper, Kyverno uses YAML instead of Rego and supports mutating and generating resources at admission. Ideal for customized policies, automated O&M, or multi-cluster policy governance. | All regions | ||
Secure confidential container environments with remote attestation | PeerPod remote attestation verifies that confidential containers run in a genuine confidential computing environment, such as Intel TDX. It automatically verifies nodes before deployment and provides on-demand runtime attestations for end-to-end workload security. | All regions | Use remote attestation to secure confidential container environments | |
Deploy A2A protocol servers in Knative | Agent2Agent (A2A) is an open standard for communication between AI agents. Deploy an A2A server in Knative to leverage auto scaling (including scale to zero) for on-demand resource usage and rapid iteration. | All regions | ||
ack-agent-gateway best practices |
| All regions | ||
ACK One | FederatedHPA best practice based on vLLM custom metrics | Large language model (LLM) services with fluctuating traffic often use multi-cluster architectures. Deploy a vLLM inference service on an ACK One fleet and use FederatedHPA for cross-cluster elastic scaling. | All regions |
November 2025
Product | Feature name | Description | Region | Related documentation |
Container Service for Kubernetes | Support for deploying and running GPU workloads in smart managed mode | Enable smart managed mode to dynamically scale GPU resources with smart managed node pools, reducing costs for variable-demand scenarios such as online inference. | All | |
New servicemesh-operator component | The servicemesh-operator simplifies deploying, upgrading, and configuring Service Mesh (ASM) in ACK clusters, enabling features such as traffic management, security, and observability. | All | ||
New built-in FinOps rule library | ACK cluster policy management now includes a FinOps rule library alongside Compliance, Infra, K8s-general, and PSP for validating Pod deployment and update requests. | All | ||
Support for deploying MCP Server on Knative | Host MCP Server on Knative to leverage Serverless benefits such as on-demand autoscaling and event-driven capabilities. | All | ||
Best practices for configuring rolling updates and graceful shutdown | Configure readiness probes, readinessGates, preStop hooks, and SLB graceful shutdown for zero-downtime application updates and smooth traffic migration. | All | ||
Distributed Cloud Container Platform ACK One | Best practices for multi-cluster priority-based elastic scheduling | Set cluster priorities in multi-cluster fleets to prioritize IDC or primary-region resources, with cloud or secondary-region resources as backup. Combined with inventory-aware scheduling, this ensures business continuity. | All |
October 2025
Product | Feature name | Description | Region | Related documentation |
Container Service for Kubernetes | Best practices for OSS volume performance tuning | Troubleshoot and optimize OSS volume performance issues such as high latency or low throughput with systematic diagnostic steps. | All | |
Schedule GPUs with DRA | Deploy the NVIDIA DRA driver to overcome traditional device plugin limitations. The Kubernetes DRA API enables dynamic GPU allocation and fine-grained resource control between pods, improving utilization and reducing costs. | All | ||
Route external MaaS services with Gateway with Inference Extension | Use Gateway with Inference Extension to centralize API key and request path management when connecting to external MaaS services such as Bailian. Configure HTTPRoute rules to inject credentials and rewrite URLs automatically. | All | Route external MaaS services with Gateway with Inference Extension | |
ACK One | ACS GPU-HPN capacity reservation for registered clusters | Register on-premises clusters with ACK One to reserve GPU-HPN capacity for unified GPU scheduling across hybrid cloud environments, supporting AI training and inference workloads. | All | Example: Use ACS GPU HPN computing power in an ACK One registered cluster |
Collect control plane metrics with a self-managed Prometheus | Install Metrics Aggregator and configure a ServiceMonitor to integrate ACK One registered cluster control plane metrics into your self-managed Prometheus for unified alerting. | All | Collect control plane metrics with a self-managed Prometheus | |
Cloud-native AI Suite | Submit eRDMA-accelerated PyTorch distributed training jobs with Arena | Submit PyTorch distributed training jobs with Arena and configure eRDMA network acceleration for low-latency, high-throughput inter-node communication, improving training efficiency. | All | Submit eRDMA-accelerated PyTorch distributed training jobs with Arena |
September 2025
Product | Feature | Description | Region | Related documentation |
Container Service for Kubernetes | Support for Kubernetes 1.34 | ACK supports Kubernetes 1.34. Create new clusters or upgrade existing compatible clusters to this version. | All | |
Support for hybrid cloud node pools | Create a hybrid cloud node pool in an ACK managed Pro cluster to manage on-premises and cloud resources from a single plane. Add existing hybrid cloud nodes for elastic scaling and cost optimization. | All | ||
Support for configuring DNS resolution for hybrid cloud node pools | Frequent CoreDNS queries through a leased line can overload the connection and cause DNS failures. Configure NodeLocal DNSCache to cache queries locally on each node. | All | ||
Support for the Terway-Hybrid network plugin | Standard CNI plugins cannot handle the complex topologies of hybrid cloud node pools. Terway-Hybrid ensures seamless Pod connectivity across your data center and the cloud. | All | ||
RRSA authentication for ossfs 2.0 | Mount an OSS bucket as an ossfs 2.0 volume with a dynamic PV. RRSA authentication is recommended for auto-rotating credentials and Pod-level permission isolation in production and multi-tenant environments. | All | ||
Support for AI-powered content moderation | Enforce content compliance for generative AI services with Gateway API and Inference Extension. Configure the ACKTrafficFilter plugin to integrate with Content Moderation service for gateway-level content blocking. | All | Use Gateway with Inference Extension to implement AI-powered content moderation | |
ACK One (Alibaba Cloud Distributed Cloud Container Platform) | Support for integrating cloud-based GPU compute power | By providing unified scheduling and O&M for heterogeneous computing resources, ACK One registered clusters significantly improve resource utilization. | All | |
Migrate single-cluster applications to a fleet for multi-cluster distribution | Use AMC to deploy applications to multiple clusters, preventing configuration drift and enabling centralized management with automatic synchronization. | All | Migrate single-cluster applications to a fleet and distribute them to multiple clusters |
August 2025
Product | Feature | Description | Region | Documentation |
Container Service for Kubernetes | KV Cache-aware load balancing with intelligent inference routing | KV Cache-aware load balancing dynamically routes AI inference requests to the optimal compute node, improving large language model (LLM) service efficiency. | All regions | |
Support for custom CNI plugins | ACK supports Bring Your Own CNI (BYOCNI) mode, letting you install a custom CNI plugin when the default Terway and Flannel plugins don't meet your needs. | All regions | ||
Managed policy governance for smart managed mode clusters | Enable the security policy management feature to meet compliance requirements and enhance cluster security. Security policy rules include Infra, Compliance, PSP, and K8s-general. | All regions | ||
Knative support for ACS compute resources | Configure Knative Services to use ACS compute resources. Leverage diverse compute types and service levels to optimize costs for various workloads. | All regions | ||
More flexible configurations for Gateway with Inference Extension |
| All regions | ||
Configure inference routing for SGLang PD-separated services | PD separation decouples the prefill and decode phases onto different GPUs, eliminating resource contention, reducing time to first token (TPOT), and increasing throughput. | All regions | ||
Securely deploy vLLM inference services in ACK heterogeneous confidential computing clusters | LLM inference involves sensitive data and model assets at risk of exposure in untrusted environments. ACK-CAI integrates hardware-based confidential computing (Intel TDX, GPU TEE) to provide end-to-end security for model inference. | All regions | Securely deploy vLLM inference services in ACK heterogeneous confidential computing clusters | |
Cloud-native AI Suite | Introducing the AI Serving Stack | The AI Serving Stack is an end-to-end ACK-based solution for cloud-native AI inference, covering deployment, intelligent routing, elastic scaling, and observability across the full LLM inference lifecycle. | All regions |
July 2025
Product | Feature | Description | Region | Related documentation |
Container Service for Kubernetes | Access ECS instance metadata in enforced mode only | ACK cluster nodes now support enforced-mode-only (IMDSv2) access to ECS instance metadata, enhancing metadata service security. | All regions | |
Subscribe to images from international registries | Use artifact subscription in ACR Enterprise Edition to automatically synchronize images from international registries such as Docker Hub, GCR, and Quay. | All regions | Obtain images from international registries through artifact subscription | |
Mount NAS by using the EFC client through CNFS | EFC improves NAS performance with distributed caching, supporting high concurrency and parallel access to large datasets. Ideal for data-intensive workloads such as big data analytics and AI training. Compared to standard NFS, EFC accelerates file access and improves read/write performance. | All regions | ||
ACK One (Distributed Cloud Container Platform) | Console-based management for GitOps | Manage GitOps capabilities from the console, including feature toggles, ACLs, ApplicationSet UI, Argo CD ConfigMap, component restarts, and monitoring. | All regions | |
Argo CD ConfigMap configuration for multi-cluster GitOps | ACK One lets you manage GitOps-related features and permissions by configuring the Argo CD ConfigMap. | All regions | ||
Inventory-aware elastic scheduling for multi-cluster fleets | ACK One provides an inventory-aware scheduler for multi-cluster fleets. When a cluster lacks resources, the scheduler deploys applications to clusters with available inventory, which then scale up nodes automatically. | All regions | Cross-region multi-cluster elastic scheduling based on inventory awareness | |
Container Service for Edge (ACK@Edge) | Configure a private connection for a leased line connection | ACK@Edge clusters can connect to the cloud over a leased line, enabling secure access to ACK and ACR while resolving network conflicts and fixed IP address issues. | All regions |
June 2025
Product | Feature name | Description | Region | Related documentation |
Container Service for Kubernetes | AI profiling | AI Profiling uses eBPF and dynamic process injection for non-invasive GPU performance analysis in Kubernetes. Attach and detach data collection on running workloads without code changes for real-time production analysis. | All regions | |
GPU node auto-healing | Node auto-healing now repairs instances affected by GPU hardware and software failures. ACK automates the full lifecycle for EGS and Lingjun node failures: detection, alerting, isolation, drain, and repair. Optional user authorization before repairs adds oversight while reducing O&M costs. | All regions | ||
CPFS for AI static volumes | CPFS for AI delivers ultra-high throughput and IOPS with end-to-end RDMA acceleration, ideal for AIGC and autonomous driving. Create static volumes in your cluster for these workloads. | All regions | ||
ACK VPD CNI component | ACK VPD CNI is the container network plugin for Lingjun nodes in ACK Pro clusters, allocating and managing network resources for Lingjun connections. | All regions | ||
ack-kms-agent-webhook-injector component | The ack-kms-agent-webhook-injector injects a KMS Agent sidecar into Pods. Applications fetch and cache credentials from KMS through a local HTTP interface, eliminating hard-coded secrets. | All regions | Import Alibaba Cloud KMS Service Credentials for Applications | |
Expanded capabilities for Gateway with Inference Extension | Gateway with Inference Extension now supports vLLM and SGLang with canary releases, inference load balancing, model name-based routing, rate limiting, and circuit breaking. | All regions | Gateway with Inference Extension Traffic Management and Inference Service Management | |
CAA solution for confidential containers on confidential virtual machines | Deploy confidential computing workloads in ACK clusters with the CAA solution. Using Intel® TDX, it protects sensitive data from external and cloud-provider threats for compliance scenarios such as financial risk control and healthcare. | All regions | Implement CAA Confidential Container Solution Based on Confidential VMs | |
Cloud-native AI Suite | Schedule Dify workflows with XXL-JOB | Dify lacks a built-in scheduler for automating tasks such as risk monitoring and data analysis. Integrate XXL-JOB to schedule and monitor Dify workflow applications. | All regions |
May 2025
Product | Feature | Description | Region | Related documentation |
Container Service for Kubernetes | Support for Kubernetes 1.33 | ACK supports Kubernetes 1.33. Create new clusters or upgrade existing ones to this version. | All regions | |
Default installation of ack-ram-authenticator component | Starting with Kubernetes 1.33, new ACK managed clusters automatically install the latest version of the managed ack-ram-authenticator component, without consuming additional cluster node resources. | All regions | ||
containerd 2.1.1 is available | containerd 2.1.1 introduces new features, such as the Node Resource Interface (NRI), Container Device Interface (CDI), and Sandbox API. | All regions | ||
Support for ossfs 2.0 | ossfs 2.0 is a FUSE-based client that mounts OSS buckets as local file systems with POSIX operations. Compared to ossfs 1.0, it improves sequential read/write performance and concurrent small-file throughput, ideal for AI training, big data, and autonomous driving. | All regions | ||
ACK One | Use ApplicationSet to coordinate multi-environment deployments and application dependencies | Combine Argo CD Progressive Syncs with ApplicationSet to build an automated deployment system that manages application dependencies between development and pre-production environments. | All regions | Use ApplicationSet to coordinate multi-environment deployments and application dependencies |
April 2025
Product | Feature | Description | Release region | Related documentation |
Container Service for Kubernetes | Create and manage Lingjun node pools | You can create and manage a Lingjun node pool in an ACK managed cluster Pro. | All regions | |
Configure a node pool by specifying instance attributes | Configure a node pool by specifying instance attributes such as vCPUs and memory. The node pool automatically selects suitable instance types during scale-out, improving scaling success rates. | All regions | ||
Real-time AI Profiling | AI Profiling uses eBPF and dynamic process injection for non-intrusive GPU diagnostics in Kubernetes. Attach and detach the profiler on live services without code changes. | All regions | ||
Enable preemption | When resources are tight, enable preemption so ACK Scheduler evicts low-priority Pods to free resources for high-priority tasks. | All regions | ||
Access services through Gateway with Inference Extension | The Gateway with Inference Extension component is built on the Envoy Gateway project. It supports all the basic capabilities of the Gateway API and the extended resources of the open-source Envoy Gateway. | All regions | ||
Generative AI service enhancements | Use Gateway with Inference Extension for intelligent routing, traffic management, canary releases, circuit breaking, and traffic mirroring for AI inference services. | All regions | ||
Back up and restore volumes from PVC to PVC | Back up and restore cloud disk data within or across ACK clusters and regions. Restore to new PVCs that can be mounted directly without modifying workload configurations. | All regions | ||
alibabacloud-privateca-issuer released | AlibabaCloud Private CA Issuer is now available. It lets you create and manage Alibaba Cloud PCA certificates in your cluster by using cert-manager. The issuer is available in the ACK App Market. | All regions | None | |
Deploy a workload and implement load balancing in an ACK managed cluster (smart managed mode) | Deploy a workload in an ACK managed cluster (smart managed mode) and expose it to the internet with an ALB Ingress for domain-based access and load balancing. | All regions | ||
Datapath V2 best practices | Optimize network configuration after enabling Datapath V2 with the Terway plugin, including Conntrack parameters and Identity resource management. | All regions | ||
Dify component upgrade guide | Upgrade ack-dify from an earlier version to v1.0.0 or later. The process includes backing up data, installing the plugin migration tool, and enabling the new plugin ecosystem. | All regions | ||
ACK One | Use PrivateLink to resolve IP conflicts in a data center network | When Serverless computing resources conflict with data center CIDR blocks in a leased-line-connected ACK One registered cluster, use PrivateLink to resolve IP conflicts. | All regions | Use PrivateLink to resolve IP conflicts in a data center network |
Schedule ACS Pods across regions | ACK One registered cluster integrates Serverless computing resources from multiple regions for dynamic cross-region GPU scheduling and unified management. | All regions | ||
Log collection | You can configure log collection by using SLS CRDs or environment variables to automatically collect container logs with Alibaba Cloud Log Service (SLS). | All regions | ||
Build a multi-cluster CD system | Combine Cloud Efficiency CD with ACK One Application Distribution to build a multi-cluster continuous delivery system. | All regions | Build a multi-cluster CD system by using ACK One and Cloud Efficiency | |
ACK Edge | Version 1.32 released | Version 1.32 is now supported. Features include optimizing requests from CoreDNS, kube-proxy, and kubelet to the kube-apiserver, reducing cloud-to-edge communication traffic, and more. | All regions | |
Network element configuration in a leased line environment | Connect on-premises servers to ACK over the internet or a leased line. For leased line connections, configure network elements first. | All regions | ||
AI Engineering Suite | HistoryServer component support | The native Ray Dashboard is only available while a cluster runs. RayCluster HistoryServer collects logs in real time and persists them to OSS for access after cluster termination. | All regions | |
KubeRay component support | Deploy the KubeRay Operator and integrate with SLS and Prometheus for enhanced log management, observability, and high availability. | All regions |
March 2025
Product | Feature | Description | Region | Related documentation |
Container Service for Kubernetes | ACK Pro managed clusters support intelligent hosting mode | When you create an ACK managed cluster, you can enable intelligent hosting mode to quickly provision a Kubernetes cluster that follows best practices. After the cluster is created, ACK automatically provisions an intelligent managed node pool. This node pool dynamically scales based on workload demand. ACK also handles all operational tasks for this node pool, including OS version upgrades, software updates, and security patching. | All regions | |
Enable tracing for control plane and data plane components | After you enable tracing for the cluster API Server or kubelet, trace data automatically flows to Managed Service for OpenTelemetry. This integration provides detailed trace visualizations, real-time topology maps, and other monitoring data. | All regions | ||
High-risk KubeConfig SMS and email notifications | Receive SMS and email alerts for high-risk KubeConfig files that still pose security risks after deletion. | All regions | None | |
Intelligent routing and traffic management with ACK Gateway with Inference Extension | Use ACK Gateway with Inference Extension for intelligent routing and traffic management for inference services. | All regions | Implement intelligent routing and traffic management with ACK Gateway with Inference Extension | |
Deploy vLLM inference applications on Knative | Traditional GPU utilization-based autoscaling doesn't reflect LLM inference load accurately. KPA adjusts resources based on QPS or RPS for more precise scaling. | All regions | ||
ACK One (Distributed Cloud Container Platform) | Unified component management for multi-cluster fleets | Define component baselines with specific versions and deploy them to multiple clusters. Supports component configuration, batch deployment, and rollbacks. | All regions | |
Dynamic distribution and rescheduling | Use PropagationPolicy to distribute replicas across clusters based on available resources. The rescheduler checks every two minutes and triggers rescheduling for Pods unschedulable for over 30 seconds. | All regions | ||
Cloud-native AI Suite | Set Slurm queue priorities | Configure Slurm queue policies for optimal scheduling and performance when jobs are submitted or change state. | All regions |
February 2025
Product | Feature | Description | Region | Related documentation |
Container Service for Kubernetes | Support for modifying control plane security groups and time zones | Modify the control plane security group and cluster time zone on the Basic Information page when initial settings no longer meet your requirements. | All | |
Node pools support custom containerd configurations | Customize containerd parameters for node pool nodes, such as configuring mirror repositories or bypassing certificate verification for specific registries. | All | Customize containerd parameter configurations for a node pool | |
Elasticity strength indicator for node pools | The elasticity strength indicator assesses node pool configuration availability and instance supply health, with recommendations to improve scale-out success rates. | All | ||
Support for batch task orchestration | Argo Workflows is a Kubernetes-native engine that orchestrates parallel tasks with | All | ||
GPU fault detection | The ack-node-problem-detector component enhances the open-source | All | ||
Accelerate pod startup in Knative services using Fluid |
| All | ||
ACK One | Schedule and distribute multi-cluster Spark jobs based on actual remaining resources | Use an ACK One fleet and ACK Koordinator to schedule multi-cluster Spark jobs based on actual remaining resources rather than requested resources, maximizing idle utilization while prioritizing online workload stability. | All | Schedule and distribute multi-cluster Spark jobs based on actual remaining resources |
Build a DeepSeek distilled model inference service in an ACK One registered cluster using ACS GPU compute power | Connect a Kubernetes cluster from an on-premises data center to an ACK One registered cluster to seamlessly scale out compute power. You can then use ACS GPU resources to deploy the DeepSeek distilled inference service efficiently. | All | ||
ACK Edge | Support for adding pod vSwitches | Add a pod vSwitch to an ACK Edge cluster using Terway Edge to expand available IP addresses when the existing vSwitch is exhausted or the pod CIDR block needs expansion. | All | |
Deploy the DeepSeek-R1 model | Use an ACK Edge cluster to manage on-premises GPUs and add cloud ACS Serverless GPUs via virtual nodes. The cluster runs tasks on on-premises GPUs first and automatically scales to cloud GPUs when needed. | All | ||
GPU resource monitoring | Integrate Prometheus monitoring with ACK Edge cluster to provide on-premises and edge GPU nodes the same observability as cloud resources. | All | Best practices for monitoring GPU resources of an ACK Edge cluster | |
Cloud Native AI Suite | Deploy a DeepSeek distilled model inference service on ACK | Deploy a production-ready DeepSeek distilled model inference service on ACK with KServe, using | All | |
Tutorial: Deploy a full-parameter DeepSeek inference service on ACK using multi-machine distributed deployment | Deploy | All |
January 2025
Product | Feature | Description | Region | Related documentation |
Container Service for Kubernetes | On-demand image acceleration for node pools | ACK supports on-demand container image loading with DADI, eliminating full downloads and decompressing data on the fly to reduce startup time. | All Regions | Accelerate container startup by using on-demand image loading |
Support for Alibaba Cloud Linux 3 Container Optimized Edition | Alibaba Cloud Linux 3 Container Optimized Edition is optimized for containerized environments, delivering higher deployment density, faster startup, and enhanced security isolation compared to the standard image. | All Regions | ||
Support for Kubernetes 1.32 | ACK supports Kubernetes 1.32. Create new clusters or upgrade existing ones to this version. | All Regions | ||
Improve resource utilization with ElasticQuotaTree and task queues | ack-kube-queue, ElasticQuotaTree, and ack-scheduler enable fair and isolated resource allocation, allowing different teams and tasks to share compute resources within a cluster. | All Regions | Improve resource utilization with ElasticQuotaTree and task queues | |
Best practice: Fine-grained resource control with resource groups | Organize ACK resources into resource groups by department, project, or environment. Combine with RAM for resource isolation and fine-grained permission management. | All Regions | ||
ACK One | Connect ACK One registered clusters to ACS compute power | An ACK One registered cluster can use container compute power from ACS. | All Regions | |
Cross-cluster service access using native service domain names | ACK One uses MultiClusterService for cross-cluster access through native service domain names, without modifying application code, Pod DNS, or CoreDNS settings. | All Regions | Access services across clusters by using native service domain names | |
Access multi-cluster resources using the Go SDK | Use the Go SDK to integrate an ACK One fleet into your platform and access member cluster resources. | All Regions | ||
ACK Edge | Cloud node scaling | When on-premises node resources are insufficient, the auto-scaling feature scales out cloud-based nodes for your ACK Edge cluster to increase scheduling capacity. | All Regions | |
Deploy elastic inference services for LLMs in a hybrid cloud | Install ack-kserve and use ACK Edge clusters cloud elasticity to deploy elastic LLM inference services in hybrid cloud, flexibly scheduling on-premises and cloud resources. | All Regions | ||
GPU sharing and scheduling | GPU sharing allows multiple pods to share the compute resources of a single GPU card. This improves GPU utilization and reduces costs.
| All Regions | ||
Centrally manage ECS resources across regions | Use an ACK Edge cluster to centrally manage compute resources across regions for full lifecycle management and efficient scheduling. | All Regions |