Create a GPU function

更新时间:
复制 MD 格式

You can deploy applications that require GPU-accelerated instances as functions using container images. This approach is ideal for popular AI projects, such as Stable Diffusion WebUI, ComfyUI, retrieval-augmented generation (RAG), and TensorRT. Using container images to deliver functions improves development and delivery efficiency.

Create a function

  1. Log on to the Function Compute console. In the navigation pane on the left, choose Function Management > Function List.

  2. In the top menu bar, select a region. On the Function List page, click Create Function.

  3. In the dialog box that appears, select GPU Function and then click Create GPU Function.

  4. On the Create GPU Function page, set the following parameters and then click Create.

    • Basic Configurations: Enter a Function Name. The name must be unique within the same Alibaba Cloud account and region, and must follow the naming conventions.

    • Elastic Configurations: Select an instance type. You cannot use provisioned instances and on-demand instances at the same time. After the function is created, you cannot change the instance type.

      • On-demand instances

        Configuration Item

        Description

        Example

        Instance Type

        Select On-demand Instance. Instances scale automatically based on request volume and are released when there are no requests. You are billed for what you use.

        On-demand Instance

        GPU Card Type

        Select a GPU card type. For more information about the specifications supported by different card types, see Instance types and specifications.

        Ada series

        Specifications

        Set the GPU Memory, vCPU, Memory, and Disk specifications for the function based on your business needs. After you set the specifications, the usage of each resource is calculated by multiplying the specification by the duration of use. For more information, see Billing overview.

        Note
        • All directories on the disk are writable. The disk space is shared.

        • The disk is tied to the instance lifetime of the underlying function. When the system reclaims the instance, data on the disk is lost. If you need persistent storage, you can mount a NAS file system or an OSS bucket. For more information, see Configure a NAS file system and Configure Object Storage Service.

        • GPU Memory: 48 GB

        • vCPU: 8 vCPU

        • Memory: 64 GB

        • Disk: 512 MB (not billed, Function Compute provides a free quota of 10 GB disk space)

        Minimum Instances

        If your business is latency-sensitive, after you select Elastic Instance, we recommend that you set the minimum number of instances to 1 or greater to lock resources in advance and reduce cold start latency.

        Note

        After you set Minimum Instances to 1 or more, if no elastic policy for the minimum number of instances is configured or if no elastic policy is active for a period, the current minimum number of instances is the value you set here.

        If multiple elastic policies are configured, the system calculates the Minimum Number Of Instances required when each policy is triggered. The system then uses the highest value among the active policies as the current Minimum Number Of Instances.

        For more information, see How is the current minimum number of instances calculated?.

        1

        Concurrency Per Instance

        You can configure multiple concurrent requests for a single GPU function instance. This means a single instance can process multiple requests simultaneously. For more information, see Configure concurrency per instance.

      • Provisioned instances

        Configuration Item

        Description

        Example

        Instance Type

        Select Provisioned Instance. Instances are allocated to the function from a pre-purchased provisioned resource pool.

        Provisioned instances are recommended for scenarios where predictable costs, low latency, and high resource utilization are important to ensure business stability.

        Provisioned Instance

        Provisioned Resource Pool

        A provisioned resource pool is a pool of provisioned instances that can be allocated to the target function. If your provisioned resource pool has insufficient capacity, click Scale-out in the Actions column and follow the on-screen instructions to expand it. For more information, see Provisioned resource pools (subscription).

        • Provisioned Resource Pool: fc-pool-****

        • GPU Card Type: Ada

        Specifications

        Set the GPU Memory, vCPU, Memory, and Disk specifications for the function based on your business needs. After you set the specifications, the usage of each resource is calculated by multiplying the specification by the duration of use. For more information, see Billing overview.

        Note
        • All directories on the disk are writable. The disk space is shared.

        • The disk is tied to the instance lifetime of the underlying function. When the system reclaims the instance, data on the disk is lost. If you need persistent storage, you can mount a NAS file system or an OSS bucket. For more information, see Configure a NAS file system and Configure Object Storage Service.

        GPU Memory: 48 GB

        vCPU: 8 vCPU

        Memory: 64 GB

        Disk: 512 MB (not billed, Function Compute provides a free quota of 10 GB disk space)

        Number Of Provisioned Instances

        Allocate a number of provisioned instances to the target function based on the resources available in the provisioned resource pool.

        1

        Concurrency Per Instance

        You can configure multiple concurrent requests for a single GPU function instance. This means a single instance can process multiple requests simultaneously. For more information, see Configure concurrency per instance.

        20

    • Function Code: Configure the function's runtime environment and code.

      Configuration Item

      Description

      Example

      Runtime Environment

      • Use Sample Image: Select a sample image provided by Function Compute to quickly deploy an image-based function. Select the target image from the image list under the Container Image configuration item.

      • Use Image from ACR: Under the Container Image configuration item, click Select Image From ACR. In the Select Container Image panel, select the created Container Registry instance and ACR image repository. Then, find the target image in the image area below and click Select in the Actions column. For more information, see Create a function that uses a custom image.

      Custom Image > Use Sample Image

      Container Image

      Select the target image.

      SpringBoot Web Application Sample Image

      Startup Command

      The startup command for the program. If you do not configure a startup command, the Entrypoint/CMD from the image is used by default.

      None

      Listener Port

      The port that the HTTP server in your code listens on.

      9000

      Execution Timeout

      Set the timeout period. The default Execution Timeout is 60 seconds, and the maximum is 86400 seconds.

      60

    • Instance Prefetch: In AI inference scenarios, you can configure instance prefetch to pre-warm the model. This eliminates the cold start latency for the first request.

      Configuration Item

      Description

      Example

      Instance Prefetch

      Instance Prefetch

      Configure an Initializer hook to pre-warm the instance and optimize cold starts. The hook runs a specified script or calls an interface to load the model after the function instance starts but before it processes requests.

      For more information about Initializer hooks, see Configure the instance lifecycle.

      Enabled

      Timeout

      Set the timeout period for the Initializer hook.

      60

      Prefetch Program Type

      You can configure two types of Initializer hooks to pre-warm the model: Execute Instruction and Invoke Code.

      Execute Instruction

      Instruction Content

      Configure the content of the instruction to execute. You can use custom shell implementations, such as /bin/bash, /bin/sh, /bin/csh, and /bin/zsh. Make sure the function's runtime environment supports the selected shell.

      See Callback method implementation

    • Use a role to grant Function Compute permissions to access other Alibaba Cloud services

      mytestrole

      Configure networks

      fc.auto.create.vpc.1632317****

      fc.auto.create.vswitch.vpc-bp1p8248****

      fc.auto.create.SecurityGroup.vsw-bp15ftbbbbd****

      Configure a NAS file system

      Configure an OSS file system

    • Configure the logging feature

    • UTC

      Manage tags

      key : value

      Configure resource groups

      Configure environment variables

      {
          "BUCKET_NAME": "MY_BUCKET",
          "TABLE_NAME": "MY_TABLE"
      }

Edit a function

After a function is created, you can change its image by editing the runtime on the Configuration tab of the function details page.

image

For information about other modifications, such as changing environment variables or log storage settings, see Configure a function.

  1. >

  2. image

References