Alibaba Cloud OpenLake is an all-in-one big data, search, and AI solution built on an open and controllable data lakehouse. It uses the Data Lake Formation (DLF) metadata management platform to manage structured, semi-structured, and unstructured data, providing secure access and I/O acceleration for lakehouse tables and files. The solution supports multi-engine integration and collaborative computing, unified development through DataWorks, and scheduling for large-scale tasks. We provide a one-click service to activate all related cloud products for the OpenLake solution. After your purchase, we will initialize a trial environment for the OpenLake solution, create resource instances, and import best-practice cases to help you explore the solution.
Solution overview
The Alibaba Cloud OpenLake solution, built on an open and controllable OpenLake lakehouse, provides integrated services for big data, search, and AI. Based on a public lakehouse in Object Storage Service (OSS) and integrated with the Data Lake Formation (DLF) metadata management platform, it supports the management of structured, semi-structured, and unstructured data. This ensures secure access to data tables and files, with create, read, update, and delete (CRUD) capabilities and I/O acceleration. The solution supports connections to multiple big data, search, and AI engines, enabling collaborative computing across them. By using the integrated IDE or Notebooks in DataWorks, you can perform multi-engine SQL or Python development in a unified environment, with support for visual scheduling and concurrent execution of large-scale tasks. You can easily create OpenLake lakehouse tables, perform data operations across different computing engines, and expose data for search and retrieval-augmented generation (RAG) by building multi-modal indexes. Within the same development environment, you can integrate AI feature engineering, model training, and online prediction to enhance data processing and analysis efficiency.
To help you quickly experience the full capabilities of the OpenLake solution, you can use a one-click service to activate the required products and initialize a trial environment.
Limitations
-
Only an Alibaba Cloud primary account or a RAM user with the AdministratorAccess permission can activate the OpenLake solution. For more information about permissions, see Product and console access control: RAM policy.
-
A free trial quota is available only for users who have completed enterprise real-name registration. Users who have not completed enterprise real-name registration can still activate the OpenLake solution through the Free Trial entry point, but the related products and services will be charged on a pay-as-you-go basis.
Free trial
The OpenLake free trial activates the pay-as-you-go service for the products listed below by default. It also provides a free quota for users who have completed enterprise real-name registration. You can view the specific discount amounts when you start the free trial.
-
Eligible users: Alibaba Cloud enterprise users. For details, see Enterprise real-name registration.
-
Deduction package validity: 3 months.
-
Free quota: Details are available on the Free Trial for OpenLake solution activation page.
Important-
Note the capacity and validity period of the deduction package. After the trial expires, services are charged at their original prices.
-
Because
Hologres/StarRocksinstances start consuming the corresponding resource packages from your subscription plan immediately after activation, they are automatically stopped after30 hours and 15 minutes/99 hours, respectively. You can manually start a stopped instance, but it will then be billed on a pay-as-you-go basis. -
Release resources after you finish the trial.
-
Product list
The OpenLake free trial activates the following products:
|
Category |
Product |
|
Development platform |
DataWorks (DataWorks Billing, DataWorks Basic Edition, DataWorks General-purpose Resource Group), Platform for AI (PAI) |
|
Storage services |
|
|
Computing resources |
MaxCompute, Hologres, EMR Serverless Spark, EMR Serverless StarRocks, Realtime Compute for Apache Flink, OpenSearch - Vector Search Edition |
The default credentials for some activated products are as follows:
-
OpenSearch account and password:
Username: admin,Login password: admin123. -
StarRocks username and password:
Username: admin,Password: Admin@01.
Activate the OpenLake solution
Step 1: Go to the OpenLake Solution page
You can go to the OpenLake Solution page using one of the following methods:
-
Method 1: Log on to the DataWorks console, switch to the target region, and then click in the left-side navigation pane to go to the OpenLake Solution page.
-
Method 2: Click this link OpenLake Big Data Solution to directly view the OpenLake Big Data & AI Integrated Solution.
Step 2: Start the free trial
The OpenLake solution offers a free trial that quickly deploys a trial environment. It also provides free trial resources for users who have completed enterprise real-name registration. Go to the OpenLake Solution page and click Free Trial. On the free trial purchase page, follow the instructions to deploy the OpenLake trial environment in your account with one click.
-
Confirm the product configuration.
By default, the solution activates the 10 cloud products in the product list and creates the relevant resource instances. The OpenLake free trial purchase page displays basic information about the cloud product instances to be created and shows how the free trial deduction package offsets the charges. On this page, you can review the cloud services to be activated and their billing methods.
NoteConsider taking a screenshot of this information for quick reference during OpenLake development.
-
Confirm resource deduction details.
OpenLake activates the pay-as-you-go service for the products by default and provides a free quota for 3 months. You can click View Details in the Deduction Package Plan section.
Important-
To avoid further costs, promptly release the initialized resources after you finish using the OpenLake solution.
-
Users who have not completed enterprise real-name registration can also use the Free Trial entry point to activate the OpenLake solution with one click. However, the related products and services will be charged on a pay-as-you-go basis.
-
-
Modify the VPC for the cloud products. (Optional)
The OpenLake solution requires all related cloud products to be deployed in the same VPC and vSwitch. After you confirm the purchase, the solution automatically creates a default VPC and vSwitch in your account. You can modify the deployment environment for these cloud products in the Network Configuration section.
-
Complete the purchase.
After confirming the details, click Purchase Free of Charge to confirm the payment and purchase the cloud products. Initialization may take 2 to 5 minutes. Do not close the window during this process.
ImportantAfter the purchase is successful, all related OpenLake cloud products are activated, and consumption of the deduction package begins. If an issue occurs during the initialization in Step 5, see Appendix: Manually configure the OpenLake environment.
-
Confirm that the environment is initialized.
After the free purchase is successful, the OpenLake trial environment will be initialized. You can go to the DataWorks console > to view the initialization process and results.
NoteAfter the products are created, you can also go to the console of each product to view the resource instances created by the OpenLake solution.
Experience the OpenLake solution
After the environment is initialized, a DataWorks workspace named OpenLake Solution Experience Space is automatically created in the current region. The environment required for the OpenLake solution to run is configured in this workspace, such as binding task execution resource groups and OpenLake-related engines. You can experience OpenLake in this workspace in the following ways.
To experience the features of the OpenLake solution, you must configure the DataWorks workspace to Use Data Studio (New Version). For details, see Create an OpenLake workspace.
Method 1: Use the built-in case
After the environment is initialized, the OpenLake Solution Experience Space workspace contains a built-in Notebook getting-started case. You can use this case to directly experience the capabilities of the OpenLake solution.
Step 1: View the built-in case
In the DataWorks console, on the OpenLake Solution page, click Go to Notebook. You will be redirected to the OpenLake Solution Experience Space DataStudio page to quickly experience the built-in case.
Step 2: Run the OpenLake case
The trial space automatically initializes the runtime environment for the Notebook job, such as a personal development environment. Before running the case, replace the environment variables required by the case with the actual parameters of your OpenLake environment. To obtain information about OpenLake environment variables, use the following methods:
-
To obtain information about computing resources, go to the DataWorks console. In the left-side navigation pane, click Work space to go to the Workspaces page. Click a workspace name to go to the Workspace Details page, and then click the Computing Resources tab to view the bound computing resources. You can also click Associate Computing Resources to bind a new computing resource.
-
For information about DLF, click this link.
# Enter values for global parameters 1-4 below before running the code.
# 1) Replace [dlf_region] with the ID of the region where DLF is located, such as cn-hangzhou, cn-beijing, cn-shanghai, or cn-shenzhen.
dlf_region = "[dlf_region]"
# 2) Replace [dlf_catalog_id] with your DLF data catalog ID. If you have activated the integrated OpenLake solution, we recommend that you use the DLF data catalog ID that corresponds to the name "myfirstcatalog".
dlf_catalog_id = "[dlf_catalog_id]"
# 3) Replace [accessKeyId] with the AccessKey ID of the current Alibaba Cloud account.
accessKeyId = "[accessKeyId]"
# 4) Replace [accessKeySecret] with the AccessKey secret of the current Alibaba Cloud account.
accessKeySecret = "[accessKeySecret]"
# The VPC endpoint of the DLF service and the endpoint for Hologres to access DLF are pre-configured. You do not need to modify the following configurations.
dlf_endpoint = f"dlfnext-vpc.{dlf_region}.aliyuncs.com"
dlf_endpoint_for_holo = f"dlfnext-share.{dlf_region}.aliyuncs.com"
Method 2: Import more cases
DataWorks Gallery provides classic Notebook cases. You can run these tutorials directly in DataWorks or adapt them to your business scenarios.
Step 1: Import from Gallery
-
On the OpenLake Solution page, click View Event Cases to go to the practice cases page: DataWorks Gallery.
-
On the DataWorksGallery page, search for
OpenLake solution Quick Startin All Categories.-
Method 1: After you find the case, click the Load Case button in the lower-right corner of the case card. In the dialog box that appears, configure the Work space and the personal development environment instance, and then click OK to import the case.
-
Method 2: Click the case title to view its details. On the details page, click Load Case. In the dialog box that appears, configure the Work space and the personal development environment instance, and then click OK to import the case.
-
Step 2: Go to DataStudio
When you load a case, you are automatically redirected to the DataStudio page in DataWorks. You can also go to DataStudio by using the following method.
On the DataWorks console > OpenLake Solution page, click Go to Notebook. You will be redirected to the OpenLake Solution Experience Space DataStudio page.
Step 3: Run the OpenLake case
Follow Step 2 in Method 1: Use the built-in case to replace the variables required by the case with the actual values of your OpenLake environment. Then, run the case.
Release resources
After you finish the OpenLake solution trial, release the related resources as described below to avoid further charges.
-
To release MaxCompute resources, see Release MaxCompute resources.
-
To release Hologres resources, see Release Hologres resources.
-
To release EMR Serverless Spark resources, see Release an EMR Serverless Spark workspace.
-
To release EMR Serverless StarRocks resources, see Release an EMR Serverless StarRocks instance.
-
To release Realtime Compute for Apache Flink resources, see Release resources for Realtime Compute for Apache Flink.
-
To release OpenSearch - Vector Search Edition resources, see Billing for OpenSearch - Vector Search Edition.
-
To release DataWorks resources, go to Purchased Resources and Services to release resource groups. For more information, see Stop using DataWorks services.
-
To release PAI resources, see Unsubscribe from PAI resources.
-
To release OSS resources, see How do I disable OSS or stop billing?
-
To release Data Lake Formation (DLF) resources, see Billing methods of Data Lake Formation (DLF).