Shared sample datasets

更新时间:
复制 MD 格式

Shared sample datasets enable you to quickly validate data processing performance, optimize query efficiency, or verify features. They are ideal resources for development, testing, and learning.

Create a sample catalog

  1. Log on to the DLF console
  2. In the left navigation pane, click Catalogs.

  3. Click Data Sharing > Shared With Me, locate the data share named dlf_samples, and click Create Catalog.

    The catalog created from a received share is read-only.

  4. Click the Catalogs tab to view the newly created catalog.

    In the Catalogs list, the status of the new catalog appears as Running.

Sample dataset list

The shared catalog includes multiple sizes of the TPC-DS standard sample database and a search sample dataset, suitable for data testing, analysis, baseline performance evaluation at different scales, and quick validation of multimodal scenarios such as image search and vector retrieval. The available datasets include the following:

Sample database name

Sample data description

tpcds_paimon_sf1

TPC-DS 1 GB Paimon table

tpcds_paimon_sf2

TPC-DS 2 GB Paimon table

tpcds_paimon_sf10

TPC-DS 10 GB Paimon table

tpcds_paimon_sf100

TPC-DS 100 GB Paimon table

tpcds_iceberg_sf1

TPC-DS 1 GB Iceberg table

search_samples

Search sample dataset with images, vectors, and official documentation data for image retrieval, vector retrieval, and full-text search scenarios

Use sample datasets

Sample datasets are created and maintained by the DLF team and support multimodal capabilities including image search, visual exploration, and full-text search. For detailed instructions, see Quickly experience DLF multimodal retrieval.