What is Elastic GPU Service

Updated at:

Elastic GPU Service is a member of the Alibaba Cloud elastic computing family. It provides GPU-accelerated computing resources that combine GPU and CPU computing power for artificial intelligence, high-performance computing (HPC), and professional graphics and image processing workloads.

Key benefits

Elastic GPU Service provides managed GPU infrastructure with the following advantages.

  • Broad coverage

    Elastic GPU Service is deployed at scale in multiple regions worldwide. This broad coverage, combined with delivery methods such as auto provisioning and elastic scaling, meets the burst workloads of your business.

  • High computing power

    Elastic GPU Service is equipped with high-performance GPU computing cards. Combined with a high-performance CPU platform, select instance types provide up to 1,000 TFLOPS of mixed-precision computing performance.

    GPUs have unique advantages in complex mathematical and geometric calculations. A GPU provides up to a hundred times the computing power of a CPU, particularly in floating-point operations and parallel computing, with the following features:

    • A large number of arithmetic logic units (ALUs) that handle large-scale concurrent computing.

    • Support for high-throughput operations with multi-threaded parallelism.

    • Relatively simple logic control units.

  • High network performance

    The VPC of an Elastic GPU Service instance supports up to 4.5 million PPS and 32 Gbit/s of internal bandwidth. In addition, Super Computing Cluster provides an RDMA network of up to 50 Gbit/s between nodes, which meets the low-latency and high-bandwidth requirements of data transmission between nodes.

  • Flexible billing

    Multiple resource billing modes are supported, including subscription, pay-as-you-go, preemptible instances, reserved instances, and storage capacity units. Resources can be purchased on demand to avoid waste.

Comparison with self-managed GPU servers

Elastic GPU Service provides compute servers for GPU and CPU applications. The following table compares Elastic GPU Service with self-managed GPU servers.

ItemElastic GPU ServiceSelf-managed GPU server
Flexibility
  • You can quickly create one or more Elastic GPU Service instances. - Instance types (vCPUs, memory, and GPUs) can be changed flexibly. Online upgrades and downgrades are supported. - Bandwidth can be increased or decreased as needed.

  • The server procurement cycle is long. - Server specifications are fixed and cannot be changed. - Bandwidth is purchased at one time and cannot be adjusted.

Ease of use
  • Online management through a web console. - Mainstream operating systems are preinstalled. Windows is activated with a genuine license. You can replace the operating system online. - GPU drivers can be installed at the time of purchase.

  • No online management tool is available. - You must provide the operating system and install or replace it yourself. - You must purchase and install GPU drivers yourself.

Disaster recovery and backup
  • Data is stored in triplicate. If a single copy is damaged, it is restored quickly. - Automatic recovery from hardware failures.

  • You must build the system yourself with commodity storage devices. - You must repair damaged data yourself.

Security
  • Effectively blocks MAC spoofing and ARP attacks. - Protects against DDoS attacks and scrubs traffic or applies blackhole filtering. - Provides additional services such as port intrusion scanning, trojan scanning, and vulnerability scanning.

  • MAC spoofing and ARP attacks are difficult to block. - Traffic scrubbing and blackhole filtering devices must be purchased separately. - Vulnerabilities, trojans, and port scanning are common problems.

Cost
  • Subscription and pay-as-you-go billing methods are supported. You can select the billing method that suits your business scenario. - Purchase on demand, without a large one-time investment.

  • On-demand purchase is not possible. You must fully provision for peak business loads. - The one-time investment is significant, and idle resources are wasted.

GPU-accelerated instance families

An instance is the basic computing unit that serves your business. Different instance types provide different computing capabilities. ECS instances are organized into instance families based on their usage scenarios. GPU-accelerated instances are a category of ECS instance types that provide GPU acceleration while retaining the same user experience as regular ECS instances.

When you create an ECS instance, select a GPU-accelerated instance type from the enterprise-level heterogeneous computing instance families, ECS Bare Metal Instance families, or Super Computing Cluster (SCC) instance families.

For more information about GPU-accelerated instance types, see GPU-accelerated instance families.

Billing

The billing-related features of Elastic GPU Service are the same as those of Elastic Compute Service (ECS). Billable resources include compute resources (vCPUs, memory, and GPUs), images, block storage, Internet bandwidth, and snapshots.

The following billing methods are commonly used:

  • Subscription: Purchase resources for a specified duration and pay before you use them.

  • Pay-as-you-go: Create and release resources on demand and pay after you use them.

  • Preemptible instance: Bid for compute resources that are in sufficient supply. Preemptible instances are offered at a discount compared with pay-as-you-go instances, but a reclaim mechanism applies.

  • Reserved instance: A coupon that is used together with pay-as-you-go instances. You commit to using instances of a specified configuration, including instance type, region, and zone, to offset bills for compute resources at a discounted price.

  • Savings plan: A discount plan that is used together with pay-as-you-go instances. You commit to a stable amount of resource usage, measured in USD per hour, to offset bills for resources such as compute resources and system disks at a discounted price.

  • Storage capacity unit: A resource plan that is used together with pay-as-you-go storage products. You commit to using a specified storage capacity to offset bills for resources such as block storage, NAS, and OSS at a discounted price.

    For more information about Elastic GPU Service billing, see Billing overview.