Data Catalog
The data catalog provides two types of directories, the default directory and the dedicated directory, to help you organize and manage the data assets visible to you in DataWorks. The default directory groups assets such as engines, databases, tables, and datasets by data source type and supports operations such as creating datasets and installing open data. The dedicated directory is based on fully managed Gravitino and supports direct connections to external metadata services.
Enter the data catalog
Log on to the DataWorks console. In the target region, click in the left-side navigation pane. On the page that appears, click Go to Data Map.
In the left-side navigation pane, click Data Map, then click Data Catalog to open the data catalog page.
After you enter the data catalog page, you can switch between the Default Directory and Dedicated Directory tabs. The page layout and common operations described in this topic are performed on the Default Directory tab. For the positioning of and differences between the two types of directories, see the following section.
Default directory and dedicated directory
The default directory and the dedicated directory serve different purposes and do not replace each other. You can choose between them based on your business scenario:
Default directory: designed for aggregation and governance. Metadata from various data sources is aggregated into Data Map through metadata collection for unified metadata retrieval and governance. The metadata is updated manually or periodically based on collection tasks and is organized by workspace and collection scope. It is suitable for centrally aggregating, retrieving, and governing collected metadata.
Dedicated directory: designed for direct access from external services. It is provided based on the fully managed Gravitino service and allows you to directly access the current metadata in external metadata services across clouds, tenants, or data lakes. The metadata is organized by metadata space, which is suitable for cross-environment direct access and isolating external metadata by space. Before you use the dedicated directory, you must allocate metadata service quotas for a subscription serverless resource group. For more information, see Dedicated Directory.
The two types of directories can be used together. To centrally retrieve and govern in Data Map the metadata managed by the dedicated directory, you can use metadata collection to synchronize the metadata to the default directory.
Page layout
The default directory uses a two-pane layout:
Left catalog tree: Shows the assets you can access, grouped by data source type. On the public cloud, the tree shows the DataSet group by default (which contains data assets under your DataWorks workspaces), and groups for each engine data source (such as MaxCompute, Hologres, DLF, Hive, MySQL). If your tenant has activated PAI, a PAI catalog group also appears; if you have activated Open Data, the DataWorks OpenData group appears at the bottom of the tree.
NoteThe data source types actually visible in the tree depend on the engines bound to your tenant and on the metadata that has already been collected. The visible scope may differ across tenants and workspaces.
Right details panel: Renders content based on the node selected in the left tree. Selecting a data source root node shows the engine overview; selecting a database or catalog shows the tables under it; selecting a table or dataset embeds the corresponding details page (the content is identical to the metadata details page).
Common operations
Browse assets: Expand a data source, database, or catalog node in the left tree to drill into its tables and datasets. Clicking a table name or dataset name switches the right panel to the corresponding details.
Search the catalog: Use the search box at the top of the catalog tree to locate a target object by node name.
Create a dataset: Under Catalogs, click DataSet. After entering the dataset list for the workspace, click Create Dataset and follow the wizard to create datasets such as OSS or NAS datasets. For more information, see Manage datasets.
Use open data: When the DataWorks OpenData group appears in the tree, you can browse and install available open datasets. For more information, see Manage open data.
Use the dedicated directory: Switch to the Dedicated Directory tab to directly connect to external metadata services across clouds, tenants, or data lakes based on the fully managed Gravitino service. For more information, see Dedicated Directory.