Create an OpenLake workspace

Updated at:

You can use the DataWorks workspace module to quickly create an OpenLake workspace.

OpenLake workspaces

  • Product introduction:

    You can create an OpenLake workspace in DataWorks by selecting the OpenLake template. An OpenLake workspace provides an integrated solution for big data, search, and AI built on an open and controllable data lakehouse. For more information about workspace templates, see Introduction to workspace templates.

    • Provides an integrated solution for big data, search, and AI built on an open and controllable data lakehouse.

    • Uses Data Lake Formation (DLF) to manage structured, semi-structured, and unstructured data, supporting secure access and I/O acceleration for lakehouse tables and files.

    • Connects to multiple engines for collaborative computing and uses DataWorks for unified development and large-scale task scheduling.

  • Region restrictions:

    You can create and use OpenLake workspaces only in the China (Hangzhou), China (Shanghai), China (Beijing), and China (Shenzhen) regions.

Prerequisites

Create a workspace

Basic information

  1. Log on to the DataWorks console and switch to the target region in the top navigation bar.

    Important

    OpenLake workspaces are currently available only in the China (Hangzhou), China (Shanghai), China (Beijing), and China (Shenzhen) regions.

  2. In the navigation pane on the left, choose Solution > OpenLake Solution to go to the OpenLake Solution page.

  3. Click Create Workspace and configure the parameters as prompted. The following table describes the parameters:

    Parameter

    Description

    Workspace Name

    Required. A unique identifier for the workspace. You cannot modify this parameter after the workspace is created.

    Display Name

    Name the workspace based on your business to easily identify its purpose.

    Description

    Describe the main purpose and other information about the workspace.

    Isolate Development and Production Environments

    Determines whether to isolate the production and development environments. For an OpenLake workspace, set this to off.

    Use Data Studio (New Version)

    Enables the new version of Data Studio. For an OpenLake workspace, you must set this to on.

    Workspace template

    Defines the tools, resources, and features available in the DataWorks workspace. When you create an OpenLake workspace, the OpenLake template is selected by default. For more information, see Introduction to workspace templates.

    Workspace Administrator

    The administrator of the workspace.

    Create AI Workspace with Same Name

    Creates a matching AI workspace with the same name. This option is enabled by default and allows you to schedule algorithm tasks on PAI.

    Default resource group of DataWorks workspace

    The default DataWorks resource group for the workspace. You can change this setting later in the workspace configuration. For more information about resource groups, see Manage DataWorks resource groups.

    Alibaba Cloud Resource Group

    Select a resource group from Alibaba Cloud Resource Management. By default, the Default Resource Group is selected.

  4. After you configure the parameters, click Create Workspace in the lower-left corner. Click Create to confirm and proceed to the Associate Data Catalog step.

Data catalog

On the Associate Data Catalog page, add an existing or create a new DLF data catalog for the OpenLake workspace.

  1. Associate a data catalog.

    • If no DLF data catalog is available: Click Create Data Catalog to go to the Data Lake Formation console. Then, activate the service and create a DLF 2.5 data catalog.

    • If a DLF data catalog is available: Click Add Data Catalog, select the target DLF 2.5 data catalog from the search box, and then click OK.

  2. After the association is complete, click The next Step to proceed to the Associate Computing Resources step.

Computing resources

On the Associate Computing Resources page, you can associate the required computing resources with the OpenLake workspace.

  1. Switch between the Offline Computing, Real-time Query, and Multimodal Search tabs, and select the computing resources to add.

    Tab

    Resource type

    Offline Computing

    MaxCompute

    Serverless Spark

    Real-time Query

    Hologres

    Serverless StarRocks

    Flink

    Multimodal Search

    OpenSearch

  2. Click Add Computing Resource at the top of the page. Associate the corresponding computing resources and test their connectivity using the instructions in the topics below:

  3. Click Complete. You can then view the created OpenLake workspace on the OpenLake Solution page.

Workspace details

  1. On the OpenLake Solution page, find the target OpenLake workspace and click Details in the Operation column to go to the Workspace Details page.

  2. You can view information about the workspace in the following sections:

    • Focus on: View the running status of instances in the workspace.

    • Workspace Assets: View asset information, such as the number of data sources, nodes, and resources.

    • Computing Resource Usage Details: Switch the Compute Resource Type (only Hologres, Serverless Spark, and Serverless StarRocks are supported), select a specific instance, and view its resource usage.

Related tasks

To create related tasks, click a module, such as OpenLake Solution, Data Inflow into the Lake (data integration), or DataStudio, at the top of the Workspace Details page.